Centific AI Research
Can your AI audit like an auditor?
An evaluation environment where AI agents investigate real government ledger data — and are scored the way professionals are.
The problem
Answering isn't auditing
Finance and audit teams review thousands of journal entries every close. Sampling misses things; blanket rules flag too much. And an AI that merely answers can't be trusted with this work — you need one that investigates: opens the documents, checks the numbers, and cites its evidence.
What this is
Real cases, real tools, real scores
A case is issued
A slice of a real government ledger and a validation question. The environment holds a verified answer key the agent never sees.
The agent investigates
It works the case with professional tools — querying the ledger, looking up reference data, opening supporting documents. Every tool call is logged.
Every claim is scored
Findings are checked against the answer key; the way the agent worked is scored against professional standards. Deterministically.
How we measure
Three axes, reported separately
Outcome
Did it find the right entries — and call each finding by the right name?
Process
Did it work like an auditor — evidence opened, citations honest, no wasted steps?
Rationale · upcoming
Is the written reasoning sound? Graded against human experts as judgment-call tasks arrive.
Why trust the numbers
Scores you can stand behind
- Fully reproducible. The same case always replays byte-for-byte — a score can be re-derived, not just believed.
- Fabrication is fatal. Cite evidence the agent was never shown and the episode fails automatically, whatever the other numbers say.
- Baselines set the floor. Two built-in scripted strategies — flag everything, answer without evidence — run on every task. A model only matters if it beats both.
No runs recorded yet. Configure and launch one above.