Why Agentic AI Changes the SOX Conversation for FP&A

Agentic AI — systems that plan, call tools, and execute multi-step finance tasks with minimal human prompting — is moving from pilot to production inside FP&A functions. According to McKinsey's 2025 survey of finance leaders, more than 60% of organizations report that AI is now embedded in at least one finance process, and CFO.com reports that the majority of midsized companies have adopted AI for FP&A specifically. The implication for Sarbanes-Oxley (SOX) compliance is direct: when an AI agent can read source ledgers, draft journal entries, post accruals, and roll forecasts forward, the boundary between "preparer" and "reviewer" shifts, and the control library must shift with it.

Also worth reading: What are the definitive steps to integrate an AI finance assistant like Cleoai into existing FP&A workflows? · How do AI agentic workflows actually transform accounting and FP&A operations in 2026? · What are autonomous finance operational workflows and how do they change FP&A?

Traditional SOX controls for FP&A were designed for human workflows: segregation of duties between preparer and reviewer, manual variance thresholds, and quarterly close calendars. Agentic AI does not eliminate those controls, but it does require new ones. The PCAOB's 2024 guidance on technology in the audit, combined with the SEC's continued focus on disclosure controls following the 2023 cybersecurity disclosure rules, means that any AI system producing numbers that flow into 10-Q or 10-K filings must be inventoried, access-controlled, and tested. EY's 2025 finance AI research notes that tech companies leading in AI adoption are also the ones formalizing model governance earliest — not because regulators demanded it, but because audit committees began asking uncomfortable questions about who, exactly, "prepared" the numbers.

The Core SOX Risks Agentic AI Introduces in FP&A

The first risk is unauthorized execution. An agent that can post to the general ledger is, from a SOX perspective, a system user. If it does so without an approver in the loop, the segregation-of-duties control fails on its face. Anthropic's 2025 financial services agent documentation explicitly recommends human-in-the-loop checkpoints for any action that mutates financial state, and IBM's 2025 financial reporting AI guidance echoes this with a four-eyes principle for journal entries above a configurable materiality threshold.

The second risk is data lineage opacity. SOX Section 404 requires management to assess the effectiveness of internal controls over financial reporting (ICFR). When an AI agent pulls from five source systems, applies a transformation the model learned during training, and writes a result, the auditor must be able to trace each input to each output. McKinsey's research found that 38% of finance AI deployments in 2024 lacked a documented data lineage, which became the top finding in management letters for those companies.

The third risk is model drift. An FP&A agent trained on 2024 revenue patterns may produce materially different forecasts in 2026 if customer mix shifts. SOX does not require models to be correct — it requires controls to detect when they are not. That means drift monitoring, periodic back-testing against actuals, and a defined escalation path when variance exceeds tolerance.

A Practical Control Framework for Agentic FP&A

The framework below is built from the patterns EY, McKinsey, and Anthropic describe in their 2025 publications, adapted for a mid-market FP&A function. It assumes the agent operates inside a finance-ops SaaS platform with API access to the ERP, EPM, and a sub-ledger or two.

The first layer is identity and access. Every agent action must execute under a named service account with scoped permissions — read-only on source systems, write-only on staging tables, and zero direct access to post to the GL. Posting to the GL is reserved for a human approver or a tightly constrained "poster" agent that requires dual approval above a configurable dollar threshold. This mirrors the SOX-friendly pattern Anthropic recommends for financial services agents.

The second layer is the action log. Every tool call, every retrieval, every prompt revision must be logged with timestamp, agent version, input hash, output hash, and the user who initiated the workflow. This log becomes the audit trail. EY's research found that companies with immutable, queryable action logs cut their SOX testing hours by roughly 30% in the first year of AI deployment, because testers could sample from the log instead of interviewing preparers.

The third layer is the control owner. Every agent workflow must have a named control owner — a human controller or senior accountant — who signs off quarterly that the agent's behavior remains within the documented design. This is the single most common gap McKinsey identified: companies deploy agents, but assign no one to own the control.

The fourth layer is materiality thresholds. Define, in writing, what dollar amount and what account types require human approval before the agent can act. A common 2025 pattern is $250,000 for balance-sheet items and $500,000 for income-statement items, with lower thresholds — often $25,000 — for accounts flagged as fraud-relevant (cash, revenue, related-party).

Comparison: Traditional vs. Agentic FP&A SOX Controls

Control AreaTraditional FP&AAgentic FP&A (2026)
Preparer/Reviewer SegregationTwo named humansHuman + named agent service account; human approver required above threshold
Journal Entry ApprovalManual workflow in ERPAgent drafts in staging table; human approves via UI; dual approval above materiality
Variance InvestigationMonthly manual reviewAgent flags variances >5% or >$100K; human investigates; agent drafts commentary
Data LineageSpreadsheet lineage notesAutomated lineage graph with input/output hashing
Access ReviewsQuarterly user access reviewQuarterly review of human users + monthly review of agent service accounts
Model/Agent RiskNot applicableQuarterly back-testing, drift monitoring, control owner sign-off
Audit TestingSample of journal entriesSample of journal entries + sample of agent action logs
Change ManagementERP change ticketsERP change tickets + agent version control + prompt change log
The table makes the structural shift visible: the agent does not replace controls, it relocates them. Where a human once typed a number, the agent now proposes a number and a human approves. Where a human once traced a formula, the system now traces a hash.

Practical Steps to Implement in 90 Days

The first 30 days should be inventory and policy. List every FP&A process that will involve an agent in the next 12 months. For each, identify the existing SOX control, the proposed agent role (read, draft, post), and the residual human control. Draft a one-page AI-in-Finance policy that names the control owner, defines materiality thresholds, and references the company's existing information security and change management policies. CFO.com's 2025 reporting indicates that companies with a written AI-in-finance policy completed their first AI-related SOX walkthroughs 40% faster than those without one.

Days 31 to 60 should be pilot and instrument. Pick one workflow — typically monthly variance commentary or rolling forecast updates — and deploy the agent in shadow mode. Shadow mode means the agent produces outputs but does not write back to any system of record. During this period, instrument the action log, validate data lineage, and run the agent's outputs against human-prepared outputs for at least two close cycles. The goal is not to prove the agent is right; the goal is to characterize its error rate and failure modes.

Days 61 to 90 should be controlled production and audit prep. Move the agent to limited production with the materiality thresholds and human approval gates active. Brief internal audit on the new control design. Brief the external auditor at the planning stage of the next interim or year-end audit. EY's 2025 research found that auditors who were briefed at planning — rather than at fieldwork — raised 50% fewer scope-expansion requests during the AI-related control testing.

Common Mistakes and How to Avoid Them

The most common mistake is treating the agent as a tool rather than a user. If the agent can post to the GL, it is a user from SOX's perspective, and it must be in the user access review. Companies that skip this step typically discover the gap during a PCAOB inspection or a management letter, and remediation is expensive.

The second mistake is over-trusting the agent's commentary. FP&A agents are increasingly capable of writing narrative variance explanations, but those narratives are generated from the same data the agent used to compute the variance. If the data is wrong, the commentary will be confidently wrong. The control is not to proofread the narrative; the control is to validate the underlying data.

The third mistake is ignoring prompt and configuration changes as change management. When a finance team updates the agent's system prompt to handle a new account, that is a change to a financial reporting control. It should go through the same change advisory board as an ERP configuration change. McKinsey's 2025 survey found that 42% of finance teams had made undocumented prompt changes in the prior quarter.

The fourth mistake is assuming the vendor is responsible for SOX. A SaaS vendor can provide a SOC 1 Type II report covering its own controls, but the customer's controls over how the agent is used remain the customer's responsibility. The PCAOB has been explicit on this point in its 2024 staff practice alert.

When to Act and What It Costs

The window to act is now — specifically, before the next fiscal year-end close. Companies that deployed agentic FP&A in 2024 and 2025 are now in their first or second audit cycle with the new controls, and the lessons are public. Companies that wait until their auditor raises the issue will spend 2-3x more on remediation, based on the remediation costs EY documented in its 2025 finance AI survey.

Pricing for a B2B AI finance-ops assistant in this category typically runs from $30,000 to $250,000 per year for a mid-market company, depending on user count, ERP connectors, and the depth of agentic workflows enabled. Implementation costs — separate from subscription — usually add 50-100% of the first-year subscription, covering control design, audit liaison, and change management. The ROI case is not primarily labor savings; it is close-cycle compression (typically 2-4 days faster), reduced audit hours, and the ability to run scenario analyses that were previously too expensive.

The Honest Limits of Agentic FP&A Under SOX

Agentic AI does not eliminate the need for human judgment in FP&A. It relocates that judgment from data gathering to data validation, from number preparation to number approval, and from narrative drafting to narrative review. Companies that frame AI as a headcount reduction play tend to underinvest in the control owner role and over-invest in automation, which produces audit findings rather than audit clean opinions.

The second limit is regulatory. SOX was written in 2002 for a world of human preparers and discrete systems. The PCAOB and SEC have issued guidance, but the guidance is interpretive, not prescriptive. Companies operating in heavily regulated industries — banking, insurance — may face additional expectations from their primary regulators that go beyond SOX. Anthropic's 2025 financial services documentation and IBM's 2025 financial reporting guidance both flag this as an area where companies should expect more specificity from regulators over the next 24 months.

The third limit is model behavior. Even well-instrumented agents can fail in unexpected ways when source systems change, when prompts are revised, or when the underlying business shifts. The control framework must assume the agent will fail and design for detection, not for prevention. That is a different mindset than traditional SOX, where the goal was to prevent errors from entering the books in the first place.

Building the Business Case Without Hype

The strongest business case for agentic FP&A under SOX is not "AI replaces accountants." It is "AI compresses the close, documents the lineage, and produces an audit trail that reduces testing hours." CFO.com's 2025 reporting shows that companies framing the case this way secured budget approval roughly twice as often as those framing it as headcount reduction. The audit committee framing matters: the question is not whether AI is exciting, but whether the controls over AI are defensible.

For a mid-market FP&A team evaluating a B2B AI finance-ops assistant in 2026, the evaluation criteria should include: documented SOX control mapping, SOC 1 Type II report, configurable materiality thresholds, immutable action logs, named control owner support, and a track record with external auditors. Vendors who cannot speak fluently to these criteria are not yet ready for production FP&A under SOX, regardless of how capable their underlying models are.