A practical answer: govern AI around the decisions it changes

An FP&A team should build AI governance as a finance operating model, not as a collection of model policies copied from an IT security playbook. The unit of governance is the decision: what AI influences, the financial impact of a wrong answer, the data used to produce it, the person accountable for the outcome, and the evidence that the system remains reliable. A forecast explanation tool that merely suggests drivers for analyst review has different risk from an agent that changes the working forecast, submits journal entries, or recommends headcount reductions. Governance should therefore permit bounded autonomy rather than assuming that all AI output is either fully automated or prohibited.

Also worth reading: What is an agentic AI governance framework for finance and how do teams implement it? · How Do You Build an AI FP&A ROI Framework That Proves Financial Value? · How Should a Treasury Team Build a Forecasting System That Works in Volatile Markets?

A useful starting principle in 2026 is that every production use case receives a named business owner, a finance-control owner, a technical owner, and an approval authority. The system inventory should record the model or vendor, version, intended purpose, users, data categories, decision rights, performance measures, review frequency, and retirement date. Human review must be meaningful: reviewers need enough time, context, and authority to challenge the output. If a person can only accept or reject thousands of forecast changes in seconds, the “human in the loop” is not a control; it is a signature collector. The framework must also distinguish advisory systems from systems that initiate actions, because approving a recommendation and executing a transaction create materially different exposure.

Why FP&A needs its own governance layer

FP&A occupies an unusual position in the business because it combines uncertain information with consequential decisions. Finance teams often work with incomplete actuals, management assumptions, volatile demand, and several versions of the budget. AI can improve pattern detection and reduce manual consolidation, but it can also make an unsupported assumption appear precise. If a model forecasts revenue at $48.2 million and a planner cannot identify the customer, pricing, currency, or cannibalization assumptions behind that figure, numerical confidence can exceed the quality of the underlying decision.

FP&A governance also needs to connect AI risk to familiar finance controls: data lineage, access rights, reconciliation, segregation of duties, journal approval, management reporting, audit trails, and management representation. General AI rules may require transparency or bias testing, but finance teams must ask more specific questions. Can the tool distinguish actuals from estimates? Can it preserve the submitted budget? Does an update to a prompt alter accounting treatment, currency conversion, allocation logic, or scenario assumptions? Is management able to reproduce the reported number after a model or prompt changes? These are control questions, and they should be answered in the same evidence repository used for other high-impact finance processes.

The objective is not to prevent innovation. A rigid framework that routes every harmless drafting task through legal, security, and model-risk review will drive teams toward unauthorized shadow use. The better approach is a tiered model. Advisory analytics with no financial action and restricted data may receive lightweight review, while systems that alter forecasts, prepare entries, or support compensation or headcount decisions require stronger testing, segregation, and monitoring. Governance should be proportionate to impact, reversibility, data sensitivity, and the degree of automation.

Design risk tiers around impact and autonomy

Before deploying an AI use case, FP&A should assess it across at least four dimensions: financial materiality, data sensitivity, decision autonomy, and recoverability. A system that summarizes internal commentary for an analyst is not comparable to one that accesses bank data and initiates payments. A recommendation that management can ignore is not comparable to an agent that automatically changes the consolidated forecast after every close. Teams should not rely on vendor descriptions such as “copilot,” “agent,” or “autonomous”; they should examine the actual permissions and actions available in production.

A workable structure uses three levels. Tier 1 covers low-impact advisory use, such as drafting a narrative or identifying non-material variances, with de-identified or low-sensitivity data and mandatory human validation. Tier 2 covers decision support that enters forecasts, planning scenarios, or management reports; it requires data lineage, documented assumptions, back-testing, user access controls, and periodic finance review. Tier 3 covers action-taking or highly sensitive decisions, including journal preparation, budget updates, pricing recommendations, or workforce scenarios; it requires approved controls, segregation of duties, immutable logs, independent validation, rollback procedures, and executive acceptance of residual risk.

Governance dimensionTier 1: advisoryTier 2: decision supportTier 3: action-taking or highly sensitive
Typical FP&A useNarrative drafting, meeting summaries, variance explanationsRolling forecasts, scenario analysis, driver-based planningJournal preparation, automatic budget changes, compensation or headcount recommendations
Data accessPublic or de-identified internal dataApproved finance and operational dataRestricted, confidential, or regulated data
Human controlReview before useReview and sign-off before inclusion in reportingPre-action approval or tightly monitored delegated authority
Core evidencePurpose, data classification, user listLineage, test results, assumptions, performance trendAudit log, segregation evidence, rollback test, incident procedure, executive approval
Review frequencyAt least annually and after material changesMonthly or quarterly, based on useContinuous monitoring with formal periodic recertification
The risk tier should follow the system if its purpose, data, model, or permissions change. Downgrading a tier merely because a vendor markets the feature as “assistive” would be a control failure.

Establish ownership, accountability, and decision rights

The framework needs one accountable business owner for each use case, even when several teams use the same platform. The business owner decides whether the financial objective remains valid, who may rely on the output, and whether the system can remain in production. A finance-control owner confirms that assumptions, reconciliations, review procedures, and reporting boundaries are adequate. A technical owner monitors model availability, version changes, access, latency, and security, while an independent challenger or model-risk function tests whether controls work in practice. In smaller organizations, one person may hold more than one role, but responsibilities should still be documented and conflicts should be disclosed.

Segregation of duties deserves particular attention. The person who configures a forecast rule or approval threshold should not be the only person able to release changes. The person who maintains source data should not independently validate the resulting forecast. Teams should also decide who can override the system, what constitutes an override, and how overrides become part of the audit trail. A manual correction without a reason code can hide systematic weakness; a repeated pattern of overrides may show that the model or process design is unsuitable.

Escalation criteria should be explicit. For example, a forecast change affecting more than 1% of annual revenue or $1 million in EBITDA—whichever is lower for the business—could require finance-director approval. Thresholds must be tailored to company scale, materiality, and decision volatility. A 1% variance may be trivial for a multibillion-dollar enterprise and material for a startup. The framework should also state who can suspend a system during a close, vendor incident, data-quality failure, or unexplained performance decline. Accountability without stop authority is incomplete.

Control data from ingestion through decision

FP&A data creates governance problems because the same metric may pass through ERP exports, spreadsheets, planning platforms, data warehouses, APIs, and language-model prompts. Owners should be able to trace each important output to its source, transformation, and effective date. Teams should document whether figures are GAAP, IFRS, or management-defined; which exchange rates apply; whether actuals include late adjustments; and how eliminations and intercompany entries are handled. Metadata, timestamps, currencies, units, and organizational dimensions should travel with the data instead of being left to user memory.

Access should follow least privilege. A planning assistant may need approved revenue, cost, and headcount data, but it should not automatically gain bank credentials, unrestricted payroll records, or write access to the general ledger. Prompts, retrieved documents, model outputs, and vendor logs may contain commercially sensitive or personal information, so retention periods and deletion controls must cover the entire service chain. Contracts should identify where data is processed, whether customer inputs are used to train shared models, who can access them, how long they are retained, and what happens when the contract ends. Cross-border processing and government requests for data should be reviewed against applicable law and the company’s risk appetite.

Validation must include more than total forecast accuracy. Teams should test performance by product, region, customer segment, business unit, and time period because a satisfactory aggregate error can conceal serious subgroup weakness. A useful minimum test set should include normal periods, seasonality, recent structural changes, missing data, extreme assumptions, and known historical restatements. Finance professionals should compare AI output with existing planner judgment, statistical baselines, and a simple alternative. If AI improves forecast error by only 2% while adding material interpretability, privacy, or control costs, adoption may not be justified.

Monitor performance, changes, and business usefulness

Production governance is continuous because both AI systems and the business change. A model may remain technically available while its output becomes commercially irrelevant. FP&A should therefore monitor financial usefulness alongside uptime and security. Depending on the use case, measures might include forecast error, bias against the prior period, bias against actual results, variance-exploration speed, close-cycle time, adoption, override rate, planner correction rate, and the percentage of outputs with complete source evidence. Targets should be defined before deployment and compared with an existing process, not with an aspirational benchmark selected after results are known.

Thresholds should trigger defined interventions. A sustained 5% decline in forecast accuracy for two consecutive quarters might require root-cause analysis, while a 10% increase in unexplained overrides could trigger suspension. A drop in data freshness or a missing source lineage should stop automated action. These numbers are examples rather than universal standards, but having no numerical trigger is worse: every deterioration can then be rationalized as a temporary anomaly. The monitoring owner should record whether a threshold breach results in retuning, restricted use, human-only operation, or retirement.

Technology and business changes require reassessment. Reviews should be triggered by a new model version, prompt-template change, retrieval-source change, expanded data access, vendor acquisition, new user group, altered workflow, or use in a previously untested jurisdiction. Annual recertification is a floor, not a substitute for event-driven review. A material change to assumptions or underlying data may affect the risk even when the vendor has not changed the software. Governance should preserve versioned prompts, configurations, evaluation results, approvals, and model cards so an auditor can reproduce what users saw when a decision was made.

Treat vendors as part of the control environment

AI vendors can accelerate FP&A work, but the finance organization cannot outsource accountability merely because the system is managed as a service. Procurement and vendor-risk teams should assess financial durability, service history, security controls, model transparency, data segregation, support responsiveness, and exit capability. Contract language should address notification of model or infrastructure changes, audit rights, incident reporting, service levels, intellectual property, subcontractors, data deletion, regulatory cooperation, and the customer’s ability to export outputs and configuration.

Claims require evidence. Statements that a product is “SOC 2 compliant,” “bias-free,” or “enterprise-ready” do not establish fitness for forecast creation, accounting analysis, or management reporting. FP&A should request documentation relevant to the intended use, such as validation summaries, known limitations, performance across customer populations, logging capabilities, and administrative controls. A security certification may support access approval without serving as model validation.

Contracts should preserve a credible exit path. The company should know whether prompts, retrieval indexes, evaluation sets, and historical decisions can be exported in usable formats and how performance would be compared after migration. Multi-model architectures may reduce dependency, but adding providers only to appear resilient creates operational complexity. The better test is whether the team can disable a vendor, revert to a manual or baseline process, and preserve the audit trail within an agreed recovery period. For a 10-business-day close, an exit process that takes 90 days is inadequate.

Prevent the common governance failures

The first common mistake is confusing pilot success with production readiness. A compelling demonstration may use clean sample data, selected weeks, and an analyst who knew the business. Before production use, teams should test performance during difficult periods, document limitations, and require the same controls users will face. Another mistake is allowing a central innovation team to own AI while individual business units remain unclear about accountability. Platform approval does not mean that a finance leader has approved the financial assumptions or accepted the reporting risk.

The second failure is automating an unstable process. If actuals arrive late, definitions differ by region, or the spreadsheet contains hard-coded adjustments, an AI system will often reproduce those defects at greater speed. Process stabilization should precede automation, although governance can define where AI might help identify those defects. The third is treating model accuracy as the only performance measure. FP&A outputs are also judged by consistency, explainability, timeliness, auditability, and whether users can challenge them.

A fourth mistake is gathering extensive documentation but leaving review ceremonial. Policies that do not change permissions, approval gates, monitoring, or retirement decisions are statements of intent rather than controls. A fifth is ignoring shadow use. Employees may paste confidential forecasts into public tools even if an approved assistant exists because the approved product is slower or lacks a needed feature. Governance should provide a usable approved path and monitor unauthorized data movement rather than responding only with disciplinary language.

Decide when to act, restrict, or retire a system

Teams should act before deployment when the use case affects a material forecast, management commitment, journal, compensation proposal, or other consequential decision. They should act before granting access when data lineage, user permissions, retention, or cross-border processing cannot be established. They should act during production when monitoring identifies a threshold breach, unexplained override pattern, unauthorized user, material model change, or inability to reproduce a reported output. For lower-risk drafting and summarization, a lightweight review may be sufficient if data is appropriately restricted and every output is verified before external or financial use.

Restriction is often the correct response. A system can remain useful for exploration while losing authority to update the official forecast. If accuracy declines, planners may use it to identify questions rather than submit numbers. If a prompt update is not yet validated, teams can revert to a previously approved version or operate in read-only mode. Restriction preserves organizational learning without forcing an immediate all-or-nothing decision.

Retirement should be an explicit governance outcome. Systems accumulate cost through licenses, integration work, data access, evaluation, user support, and lingering vendor risk. Teams should set review dates and define failure conditions such as negligible usage, lack of measurable benefit, unresolved incidents, unavailable evidence, or an uneconomic cost to maintain. A 5% efficiency gain may justify one use case but not ten; evaluation effort must be counted when comparing a lightweight spreadsheet process with a production AI platform. The framework is succeeding when finance can explain not only why AI remains in use, but also why stopping it would be unnecessary or imprudent.