Your AI bookkeeper is wrong about one transaction in ten, and it will not tell you which one.
Categorization Auditor sits behind whatever tool codes your books, re-categorizes every transaction independently, and hands you the short list of the ones your bookkeeper got wrong, with the correct category and what the mistake would have cost. Not a confidence score, a worklist.
Audit your bookkeeper's coding See a sample disagreement report
THE PROBLEM
AI categorization is good, not perfect, and the last few percent is exactly where the money hides. Every serious source puts automated transaction coding somewhere between eighty-five and ninety-five percent accurate, and every one of them, in the same breath, says a human still has to review the exceptions. The trouble is that nobody hands that human the exceptions. You get a ledger that is mostly right, a confidence number the tool graded itself, and no way to know whether the row it felt sure about is the one that quietly booked a client dinner as groceries (or a personal Netflix charge as software). So the review either does not happen, or it happens the slow way, someone scrolling the whole ledger hoping to catch what the machine missed. Come audit season, or the moment the EU AI Act's auditability obligations bind on August 2, 2026, a confidence score is not the receipt anyone wants to be holding.
HOW IT WORKS
Categorization Auditor does the one thing the coder cannot do for itself, it checks the coder:
- It pulls the batch. Every transaction your bookkeeper coded, exactly as it stands. - It re-categorizes independently. AYA assigns its own category to each transaction, without looking at what the bookkeeper decided, so the two answers are genuinely independent. - It diffs and classifies. Where the two disagree, it names the kind of miss (uncategorized, too generic, a business-versus-personal flip, a tax-sensitive miscode, or a plain wrong category), and it fails closed, a malformed row is flagged for review, never silently dropped. - It prices the mistakes. Each disagreement carries the correct category as a receipt and a dollar cost, so the output is a ranked worklist a human can clear in minutes, plus a single number for what the miscodes were worth.
The diff is deterministic and model-free, so the same books produce the same audit every time, which is rather the point of an audit.
WHAT YOU GET
A disagreement report you can put in front of a reviewer, a controller, or an auditor without flinching. It has three parts that travel together, the accuracy your bookkeeper actually hit on this batch, the exception worklist (only the rows that disagree, each with the right answer), and the cost delta. The number and the proof of the number arrive as one artifact, because a percentage without the rows behind it is just another score to argue with.
WHO IT'S FOR
Anyone who signs off on books an AI helped code. A controller who owns the month-end close and cannot personally re-check every line. A fractional CFO or a bookkeeping firm carrying many clients, where one miscode repeated across a year compounds into a real misstatement. A finance team shipping AI into their ledger who will, on August 2, 2026, owe a regulator proof of what that AI actually did. If a wrong category costs you something, this is built for you. If you are budgeting your grocery spend, it is honestly overkill.
PRICING (the ladder)
- Audit. A fixed-price, one-off pass on a batch of your coded transactions. You see the disagreement report, the worklist, and the cost delta before you commit to anything else. - Metered. Per verified transaction, once the audit has proven the delta. Your bookkeeper keeps coding, we keep grading, and the worklist flows to your reviewer. - Platform. The deterministic auditor wired into your close, so no month books without a categorization receipt attached.
(Prices are hypotheses we are validating, not commitments carved in stone.)
THE PROOF (the dogfood)
We build the way this product audits. This very deliverable terminated as a verification receipt before we wrote a word of marketing, the diff compute runs deterministically on a known batch and lands the numbers we expect, the pattern passed our quality gate, and the pieces it composes were checked to resolve on disk. The auditor is producer-agnostic on purpose, it grades any coder's category against ours, which means it can grade AYA's own bookkeeping exactly the way it grades yours. We are not selling a checker we would not point at ourselves.
HONEST NOTE
Being straight with you, an honest categorization audit usually finds more mistakes than the tool's own dashboard admitted to, and that is the value, not a defect. Two things are still true today. The diff engine is real and runs deterministically right now, but a live run on your ledger waits on read-only access to your books and a booted system, so this is capability-ready, not one-click-live for a stranger yet. And the cost figures in the report start from conservative placeholders until we calibrate them against a real batch of yours, so treat the first dollar delta as a floor we sharpen, not a final invoice. The mistakes it names, though, are real rows with real correct answers, and those you can check yourself.
Audit your bookkeeper's coding
*This page is a specification. The capability it describes is not built yet, and nothing here is a claim that it runs today.*