Twelve questions decide whether your automation reads as a governed control or an unmanaged tool to the person signing off on your ICFR — and only one of those passes.
External auditors now systematically evaluate whether AI touching financial reporting behaves as a governed control or an unmanaged tool. Here are the twelve questions that decide which one you are.
What your SOX auditor will ask about AI in 2026
Your AP team deployed an AI approval workflow eighteen months ago. It's been quietly reconciling invoices ever since, nobody's touched it, and it's never come up in a status meeting. Then December arrives, the external SOX audit begins, and the first question the auditor asks isn't about the invoices at all — it's "explain how this control operates, in plain language, and show me the decision logic version history."
Most finance and IT teams have an answer to "does the AI work." Far fewer have an answer to "is this a control or a tool" — which is the actual question underneath every one of the twelve an auditor will ask. A tool is something a person uses and remains accountable for. A control is something the organization can prove operates consistently, is version-tracked, and produces evidence an auditor can independently verify. AI that can't answer that distinction cleanly becomes the audit finding, not the automation win it was supposed to be.
This article covers the three changes that made 2026 the turning point, the control-vs-tool framework auditors are actually applying, the twelve questions grouped by category, and the four artifacts you need ready before the walkthrough starts.
PCAOB's amended AS 2201 and AS 2101 standards, effective for fiscal years ending December 15, 2026, expand benchmarking provisions specifically for fully automated controls — meaning AI-touched controls get real scrutiny under an updated standard, not a carve-out.
Big Four firms now specifically train staff to scrutinize AI-touched controls and AI-generated evidence — this isn't a generalist auditor guessing at AI; it's a trained reviewer who knows what a real audit trail looks like.
The bar moved from point-in-time screenshots to continuous control monitoring — timestamped, attributed logs that track decision distributions and exception rates over time, not a snapshot from the week before the audit.
A control is governed, deterministic, evidenced, and version-controlled. A tool is unmanaged, probabilistic, and unexplainable. Every question your auditor asks is really testing which one your AI is.
This is where probabilistic and deterministic AI diverge sharply in audit outcomes — not because one is smarter than the other, but because only one of them produces the evidence an auditor can actually test.
| Dimension | Probabilistic AI | Deterministic AI |
|---|---|---|
| Walkthrough explanation | "The model recommended this" | Explicit English policy, cited by name |
| Decision logic | Implicit in model weights | Inspectable and versioned |
| Audit trail | Confidence score only | Full reasoning chain |
| Benchmarking eligibility | Difficult to establish | Eligible with stable logic |
| Remediation | Requires retraining | Edit the policy, redeploy |
This is not exhaustive audit or legal advice, and it reflects an auditor-side perspective rather than regulatory requirements from the SEC or under the EU AI Act. Confirm scope with your own auditors before collecting evidence, and run an internal pre-walkthrough readiness review to surface gaps early.
Finance teams that pass this cleanly don't scramble when the auditor asks. They have four things ready, deliverable within 30 minutes for each control.
Every AI touchpoint mapped to the financial assertion it affects
The decision logic, written so a non-engineer auditor can read it
Real transactions with the full reasoning chain, not a summary
Evidence the model and policy logic have stayed stable — or a logged reason why they changed
Silent third-party model upgrades reopen PCAOB AS 2201 operating-effectiveness testing at every upgrade — pin model versions per workflow and log every upgrade as an explicit, dated event, even when the vendor pushed it without asking you.
Every one of the twelve questions maps back to the same underlying requirement: certified, version-pinned logic before production; deterministic, escalating execution while it runs; and a reasoning-level audit trail after the fact. That's the pipeline AgentTrust OS is built around.
Trust Certify pins the model and policy version before a workflow reaches production and logs every subsequent change as an explicit event — the exact evidence PCAOB AS 2201 benchmarking requires to skip re-testing a stable control.
Trust Runtime enforces the deterministic, explainable execution auditors are trained to look for, escalating uncertain or high-impact decisions instead of returning a bare confidence score — the difference between "the model recommended this" and a citable policy.
Trust Audit produces the plain-language reasoning chain, tied to a specific decision ID, that turns "provide audit trails for specific decisions" from a scramble into a 30-minute artifact pull.
Confidence in every decision — pre-production certification to post-deployment audit.
Start Free →That's the most common misconception, and it undersells what's different. ITGCs test whether a system operates as designed; AI-specific testing also has to establish that the decision logic itself is inspectable and stable — something a black-box model can't demonstrate the way a documented, version-controlled policy can.
SOC 2 speaks to the vendor's security and availability controls, not to whether your specific control operates effectively over your financial reporting. An auditor will still ask for a walkthrough of your control, your decision logic, and your evidence — vendor certification is a supporting fact, not a substitute.
The underlying SOX requirement — reconstructable, attributable evidence — hasn't changed. What's changed is that AI decisions are harder to reconstruct by default, because the reasoning often lives in model weights instead of a document a human wrote. The requirement is old; the difficulty of meeting it with off-the-shelf AI is new.
Start with the AI inventory — every place AI influences a financial reporting decision, mapped to the specific assertion it affects. That single artifact usually surfaces most of the gaps before an auditor does, and it's the first thing they'll ask for anyway.
Honest answer: version pinning and reasoning capture add a small overhead to every decision that wasn't there in an ungoverned prototype. But that overhead is what turns a control into something benchmarking-eligible under AS 2201 — meaning your auditor doesn't re-test it every single year. The alternative is redoing the full walkthrough annually, which costs far more than the governance layer does.