
You Were Sold a Chatbot and Told It Was an Auditor
DataSnipper just did us a favor. Thanks, friends. :)
Their new 2026 AI Report found that 87% of audit and finance professionals rank knowledge of accounting standards as the most important capability in an AI tool. General language capability came last. Two in three said the audit-specific tools they need still do not exist.
We reached this conclusion by researching the market, working alongside audit teams, and watching where general-purpose AI breaks under real engagement conditions. DataSnipper's latest survey is valuable confirmation: after four additional years of market data, the finding still holds. Audit professionals do not need a better generic writer. They need systems built around accounting standards, firm methodology, evidence, and review.
Firms are experiencing three predictable challenges with the current wave of AI agents:
1. Work appears complete before it is defensible. A chatbot can write directly into a workpaper, but it cannot reliably determine what belongs there, why it belongs there, which evidence supports it, or whether the conclusion conforms to the firm's methodology. The artifact looks finished. The test is not.
2. Review has become an expensive fact-checking exercise. When an agent cannot see risk, assertions, evidence conflicts, prior decisions, or workflow state, managers must reconstruct the reasoning after the work is done. The time saved in drafting is repurchased in review, usually at a higher billing rate.
3. Firms inherit risk without gaining real leverage. The vendor delivers fluent output, while the firm remains responsible for unsupported conclusions, missing provenance, inconsistent methodology, and decisions no one can trace. That does not reduce risk. It simply moves the risk behind a polished interface.
These failures share one cause: most so-called audit agents are wrappers around a general-purpose model, with none of the standards, methodology, engagement state, or evidence relationships required to do real testing and engagement work. The hard part was never putting words into Excel. The hard part is making those words survive review.
Deepak explained the business side in Part 1: margin disappears in the distance between doing the work and finding out the work was wrong. This is the technical side of the same problem.
Superficial Access, Superficial Work
It is Friday afternoon. A senior asks an agent to complete a revenue cutoff workpaper. The agent reads the worksheet, compares the invoice and posting dates with year-end, and writes a clean conclusion. The workpaper looks finished.
On Monday, the manager asks one question: what were the shipping terms? The agent never read the bill of lading, saw the contradictory evidence, or knew the engagement identified a risk of premature revenue recognition. It completed the visible task without understanding the relationships behind it.
The failure is subtle because every artifact the agent touched looks reasonable. The missing facts sit outside the worksheet: why the sample was selected, which engagement risk it addresses, how the documents relate, and whether an exception is still open. Those are the facts the manager must reconstruct later.
Access to a file is not access to methodology, risk, evidence lineage, or review authority. Retrieval can find a paragraph, but it cannot determine whether that paragraph governs this entity, this period, this risk, and this decision. A polished workpaper can still be an incomplete test.
The conclusion is not the test. The traceable chain from risk to procedure to evidence to conclusion is.
The Missing System Is Context
When AI teams talk about context, they often mean how much text a model can read. That is capacity. Audit context is architecture.
A production testing agent needs four connected layers:
Firm methodology defines the required procedure, relevant standards, accepted interpretations, and quality controls.
Engagement context explains the entity, period, materiality, risks, assertions, scope, and prior decisions.
Evidence context connects each claim to the exact source and version that supports it, while preserving conflicts between documents.
Workflow context records who prepared and reviewed the work, which exceptions remain open, and who has authority to clear the conclusion.
Those facts cannot remain as prose scattered across a binder. They need explicit relationships between risk, procedure, evidence, conclusion, and review status so the agent can assemble the right context for the decision at hand.
Without it, the agent sees a spreadsheet. With it, the system sees the engagement.
Context Isn’t Capacity. It’s Architecture.
A context-aware system compiles only the current, authorized facts needed for the decision at hand. It does not send the entire binder to a model and hope the important details win.
Mechanical questions run as deterministic checks: did shipment occur before year-end, does the invoice agree to the ledger, and are all required documents present? Model reasoning is reserved for judgment: do the shipping terms change the recognition date, does the explanation resolve the contradiction, and does the conclusion respond to the identified risk?
Provenance comes before prose. Every claim carries its source, and the decision trail records what the system checked, what it found, what the auditor changed, and who cleared it.
When two sources disagree, the system surfaces the conflict instead of selecting the friendliest answer.
That is the difference between a model that can write about auditing and a system that can support audit procedures safely.
Context Is How Tellen Moves Review Upstream
This is where the architecture connects back to Deepak's margin argument. Read Part 1: Everyone Is Trying to Speed Up Audit. That’s Exactly the Wrong Fix.
You cannot place a manager beside every senior. Tellen moves the manager's review logic to the point of work by giving each agent the methodology, engagement state, evidence relationships, and prior decisions required to ask a narrow question at the right time.
In the revenue example, Tellen surfaces the bill of lading conflict before the workpaper is finished. The senior resolves it while the context is fresh, so Monday’s manager review stays focused on judgment instead of reconstructing the test.
That margin improvement only exists when context, controls, and workflow are built into the product.
Models will keep getting faster, cheaper, and more capable. Writing text into Excel will become a commodity. The durable advantage is the system surrounding the model.
DataSnipper’s survey puts the priority plainly: 87% of audit and finance professionals rank accounting standards as the most important AI capability, while general language comes last. Tellen is built around that priority. Our agents combine encoded firm methodology with engagement context, evidence lineage, workflow state, and reviewer authority, so every conclusion is grounded, traceable, and ready for human sign-off.
The result is not another copilot waiting for a better prompt. It is a governed audit system that gives every procedure more context, makes every exception traceable, and shows every reviewer not only what the agent concluded, but why.
Firms were sold chatbots and told they were auditors. The distinction is now impossible to ignore.
A wrapper puts words into a workpaper. Tellen gives those words the methodology, evidence, controls, and accountability required to survive review.
That is what an audit agent is supposed to deliver.