CRYPTO

Crypto Accounting Benchmark Tests AI Models on Real‑World Journal Entries

WorldTue Sep 15 2026

Crypto Accounting Bench (CAB) is a new testing ground for AI language models. It asks models to rebuild the exact journal entry a real organization posted for a crypto‑asset deal. The benchmark holds 118 different tasks taken from seven made‑up companies. Each task packs transaction mechanics, asset amounts, base‑currency values, wallet and legal‑entity details, counterparty evidence, linked legs, recurrence patterns, tax‑lot info, and the full chart of accounts.

The goal for each attempt is a balanced entry that includes every required account, the correct side, the amount, the currency, and the precise asset quantity. Tasks are built to mirror the complexity of actual bookkeeping, so models must piece together many pieces of information. The variety of contexts forces the AI to understand both the numbers and the surrounding business rules.

Twelve models were put through the test, covering top proprietary systems and open‑weight releases. Each model got three tries per task, creating 4,248 total attempts. Researchers measured three scores: Mean Score, Best@3, and Pass@3, which counts tasks where at least one of the three tries met all rubric criteria. The best performing model earned a Mean Score of 77.43%, while the highest Pass@3 reached 56.78%. Diagnostic checks showed that base‑amount agreement was very strong at 97.8%, but deciding which account to use lagged at 56.3%. Failure analysis points to account selection and assembling a complete entry as the biggest hurdles.

The results make it clear that AI can handle the arithmetic side of crypto accounting but still struggles with choosing the right accounts and formatting the full journal entry. CAB provides a clear roadmap for future improvements, highlighting where models need more guidance and training.

actions