ASU 2023-07: Mapping GL Accounts to 8 Expense Buckets

Chart of Accounts, 8 Buckets
FASB issued ASU 2023-07, and the requirement list looks short until you try to source it from a general ledger: disclosure of significant segment expenses regularly provided to the CODM, an "other segment items" catch-all, and — the piece no mapping table contains — the CODM's title or position plus an explanation of how the CODM uses the profit measure. The standard takes effect on a phased basis across annual and interim periods, applies retrospectively, and explicitly includes single-reportable-segment entities — the clause that surprises filers who assumed segments were someone else's problem.
That combination retires the comfortable claim that an existing account-group mapping table "already satisfies" the standard and that this is just a footnote. It fails both halves: no group-mapping tab stores the CODM's title or the narrative explaining how the measure gets used, and interim effectiveness converted a one-time 10-K mapping project into a recurring quarterly control that must hold retrospectively consistent across every comparative period.
The cadence is the real operational break. A mapping exercise that used to be an annual 10-K project must now survive four quarterly refreshes, catching mid-year account creations — a cloud-subscriptions account opened in Q3, a one-off legal accrual account — that previously stayed invisible until year-end. Under retrospective application, an unmapped September account is not a cleanup item; it is a prior-period inconsistency sitting inside a filed document.
The pipeline that absorbs this workload runs six stages, each producing a named artifact:
| Stage | Action | Artifact produced |
|---|---|---|
| 1. Export | Pull the trial balance plus 24 months of posting history from the ERP (NetSuite, SAP S/4HANA, Oracle Fusion) | Account-level posting corpus |
| 2. Embed | Vectorize account metadata together with transaction patterns | Embeddings per natural account |
| 3. Classify | Match against category definitions lifted verbatim from the CODM's monthly reporting deck, not GL naming conventions | Draft category assignments |
| 4. Score | Attach a confidence value to every account | Per-account confidence figure |
| 5. Route | Split by threshold into auto-accept, reviewer queue, or controller escalation | Three-lane worklist |
| 6. Publish | Emit a versioned account-to-category table the consolidation layer (OneStream, Oracle FCCS) consumes | Mechanical 10-Q tie-out |
The scoring behavior deserves a close look because confidence tracks description quality, not dollar risk. An account captioned "Cloud Subscriptions" maps to a "Software and cloud costs" category at 0.97 and posts automatically — it clears both halves of the gate described earlier. An account captioned "Other G&A" queues for named preparer-reviewer sign-off at 0.61 because its description carries no signal, and nothing about that score says the account is small. Low-confidence catch-alls are exactly where a misclassified accrual hides; that is why the human lane exists.
The compression is the design goal. A typical mid-market chart of accounts runs deep with active natural accounts, and ASU 2023-07 asks them to collapse into roughly five to twelve significant expense categories per segment — in the running case here, the entire chart funneling into eight buckets. The long tail legitimately belongs in "other segment items": CFA Institute's July 2013 investor report (Financial Reporting Disclosures, ISBN 978-0-938367-76-5) devoted an entire chapter to asking where all the immaterial information should go, and the catch-all is the standard's structural answer. Forcing tail accounts into significant categories to make the map look complete manufactures restatement exposure out of diligence theater.
What exits the pipeline is not a spreadsheet. It is a versioned mapping table storing, per account, the model version, the prompt template, the confidence score, and the reviewer identity — the evidence package an inspector or external auditor requests the moment the disclosure note asserts how the CODM uses the measure. Academic work anticipated the shape of this: arXiv paper 2605.23924v1 (DOI 10.48550/arXiv.2605.23924) applies an LLM specifically to improving the completeness of segment disclosures. What controllers are operationalizing in 2026 is that literature with sign-off attached.
Concrete next step: run stage 1 against your live ERP today and count the natural accounts with no assigned category owner. That count — not the disclosure template — is your true remediation scope for the next 10-Q.

What the First Post-Adoption Cycle Actually
McKinsey's research paper "Bots, algorithms, and the future of the finance function" set an early ceiling on the topic, asking how much of the finance function was technically automatable with then-current technology. Hold every 2026 mapping-automation claim against that benchmark. Account-to-category classification is among the *easier* activities within that scope — it is pattern-matching, not judgment — so a tool claiming to eliminate most mapping labor is plausible. A tool claiming to automate the whole close is claiming more than McKinsey said was technically possible for the entire function, and deserves immediate skepticism.
The strongest direct evidence for the drafter role comes from the CPA-exam benchmark published by May et al.: GPT-4 answered roughly 85.7 percent of multiple-choice questions correctly. That is professional-level performance on standardized accounting classification tasks — real capability, not marketing. But read the caveat closely: exam questions are cleaner than live chart-of-accounts descriptions, where an account named "miscellaneous operating" gives the model nothing to classify against. Note also where 85.7 percent sits relative to the high-confidence auto-accept floor described earlier in this guide — below it, on sanitized inputs. Messier inputs push raw accuracy down, not up.
Scale turns this from a curiosity into an industry-wide obligation. According to SEC EDGAR data, thousands of active domestic companies file Form 10-Q, and each one now faces the interim significant-expense disclosure every quarter. This is also where the "our existing mapping table is just a footnote" argument dies: interim effectiveness converted that table into a recurring quarterly control that must stay retrospectively consistent, refreshed and evidenced four times a year across thousands of filers simultaneously. The quarterly mapping-refresh problem is no longer confined to complex multinationals with dozens of segments.
The enforcement signal arrived on schedule. According to Audit Analytics' SEC comment-letter tracking, Topic 280 citations climbed in staff comment letters following the first post-adoption annual reports, with reviewers probing two things specifically: who the CODM actually is, and how the significant-expense determination was made. One discipline note before you quote this in a board deck — the exact registrant count moves with every quarterly update, so retrieve the current figure directly from Audit Analytics' latest comment-letter release rather than relying on any snapshot printed here.
The market voted too. By the time of writing, at least three major close-management platforms — FloQast, Numeric, and BlackLine among them — had shipped AI-assisted journal-classification or mapping copilots aimed squarely at this workflow. When three competing vendors build for the same pain point, demand is real. But notice what none of them publishes: an independent accuracy benchmark. Every accuracy number you will see in a sales cycle is self-reported, which means the validation burden falls on you.
| Evidence | Source | Figure | What it establishes | Where it breaks down |
|---|---|---|---|---|
| Automation ceiling | McKinsey, "Bots, algorithms, and the future of the finance function" | Share of finance activities technically automatable | Classification sits inside the automatable band | Ceiling predates LLMs; judgment work sits outside it |
| Professional-level classification | May et al., CPA-exam benchmark | GPT-4 roughly 85.7% correct on multiple choice | LLMs perform at professional level on standardized accounting items | Exam questions are cleaner than real GL descriptions |
| Obligation scale | SEC EDGAR filer data | Thousands of active domestic 10-Q filers | Quarterly refresh is industry-wide, not multinational-only | Cadence recurs every quarter indefinitely |
| Enforcement signal | Audit Analytics comment-letter tracking | Topic 280 citations climbed post-adoption | Reviewers probe CODM identity and expense determination | Pull exact registrant count from the latest update |
| Vendor adoption | FloQast, Numeric, BlackLine releases to date | At least 3 copilots shipped | Demand is commercially validated | No independent accuracy benchmark published |
Concrete next step before your next 10-Q: download Audit Analytics' latest comment-letter update for the current Topic 280 registrant count, then run whichever copilot you've licensed against a held-out sample of last quarter's manually reviewed mappings. That test produces the accuracy benchmark the vendors won't — and it tells you whether the tool earns drafter status or stays in the review queue.

Spreadsheet, Rules Engine, or LLM Draft
The spreadsheet camp loses this argument twice. The Excel/VLOOKUP tab that carried your annual adoption is not a recurring control — it has no enforced workflow, no version history, and it fails silently: sort the lookup range wrong and accounts post to the wrong bucket with no error at all. "We already have the mapping table" was defensible for a one-time footnote project; it is indefensible for a quarterly control that must stay retrospectively consistent.
Read the maintenance column, not the build column. On quarterly effort, the hybrid is the explicit winner: the LLM drafts classifications with confidence scores, the EPM rules layer enforces them, and because the deterministic engine — not the model — performs the posting, everything downstream of the review gate remains fully deterministic and traceable.
| Approach | Initial build | Quarterly maintenance | Newly created accounts | Audit traceability | Annual cost |
|---|---|---|---|---|---|
| Excel/VLOOKUP mapping table | Days for a small chart; weeks for large ones | Hours of manual re-keying each quarter; grows with every new account | Manual — someone must remember to extend the lookup; misses surface as #N/A or, worse, a default bucket | Weak — no version history; reconstructing a prior-quarter mapping means forensics on a shared file | No license cost; the entire cost is labor at controller rates |
| Deterministic rules engine (OneStream, Anaplan, Oracle EPM) | Weeks — every rule written, tested, tied to a workflow step | Low while the chart is stable; spikes whenever accounting adds an account family | Handled only if a human writes the rule; unmapped accounts hit an exception queue or silently take a default | Strong — every posting traces to a named rule and workflow user | Bundled in your existing EPM subscription — verify the module is in your tier |
| LLM-drafted classification behind human review | Days — configure prompts and thresholds against a labeled sample of your own accounts | Minutes of exception review; add a re-baseline after any model-version change | Drafted automatically with a confidence score; sub-threshold items route to a named reviewer | Strong when gated — store prompt version, model version, score, and sign-off per account | A marginal per-account classification cost at current embedding-plus-completion pricing, plus reviewer hours |
Four inputs determine the right stack: chart-of-accounts size, new-account creation velocity, how many reportable segments share accounts, and whether the CODM package is standardized month to month. The last input dominates. A drifting CODM deck — buckets reshuffled mid-year, expense lines regrouped between board meetings — invalidates any mapping faster than model error ever accumulates. If the package definition moves quarterly, no classifier saves you; freeze the bucket definitions first, then automate.
Concrete next step: pull twelve months of new-account additions and count them. That single number selects your row.
Practitioner evidence on finance automation has always been conversational rather than measured. According to EY session records preserved on ReadKong, the firm's controller dinners were held in July and September in New York and in September 2014 in Palo Alto, California — rooms full of anecdotes, not samples. A decade later, the evidence behind the halving claim keeps that same shape: vendor pilots, single-company adoptions, conference war stories. Nobody publishes the adoption that saved nothing, and the observable record is young — interim segment-expense footnotes have existed for only a handful of quarters, so one messy implementation can swing any informal average.
| Your situation | Winning stack | Why it wins |
|---|---|---|
| Compact chart, fewer than 5 new accounts per year | Deterministic rules engine alone | Probabilistic classification adds complexity without payoff |
| Large or fast-moving chart, ML-literate staff | Hybrid: LLM drafts, EPM enforces | Lowest quarterly effort; posting stays deterministic |
| Modest revenue, no ML-literate staff | AI module inside your existing close platform | Vendor absorbs model-version churn and SOC-relevant logging |
| CODM package changes month to month | Nothing yet — stabilize the package first | A drifting deck invalidates any mapping faster than model error |
The proof controllers reach for most often — the existing account-group mapping table — carries even less weight than it appears to. It satisfies neither half of the standard: it nowhere records the CODM's title or how the CODM actually uses the segment measure, and it was built as a one-time artifact, not the recurring quarterly control with retrospective consistency that interim effectiveness demands. Calling it "just a footnote" remains the most expensive sentence in the close binder.

What the Data Doesn't Tell You
Variance across cases is wide enough that the headline reduction should be read as a ceiling, not an expectation. The saving materializes only where account-to-bucket mapping is the binding constraint on the close. Where the calendar is eaten by accruals, cut-off, or intercompany elimination, flawless mapping returns little. Granularity matters too: a ledger with a long tail of small accounts below the materiality line gives the auto-accept lane far more traffic than a compact chart of accounts, and a CODM package that shifts quarter to quarter degrades the confidence scores themselves. Note also that confidence is not a common currency — scores from different models and vendors are not calibrated to one another, so a threshold that is conservative for one engine may be loose for another.
The gate's quiet assumption is that classification errors are independent and small. Three conditions break it. First, correlation: dozens of individually immaterial accounts drifting into the same wrong bucket accumulate into a material misstatement, which is why sub-materiality auto-accepts still need a full-population tie-out each quarter. Second, novelty: a newly opened account has thin history, and model confidence is least reliable exactly where history is thinnest — route new accounts to the named reviewer regardless of score. Third, proximity: an account sitting just under the materiality line can flip a bucket across its significance threshold even on a technically correct match, so the review band should widen near the line. None of this retires the drafter-not-poster rule; it thickens the fence where the fence is known to be weak.
One verification most teams skip: export the auto-accept log and sort accepted matches by direction of effect on each bucket. If accepted-only volume pushed any significant-expense category across its disclosure threshold, the gate passed an error no individual score would ever flag — and that bucket belongs in the named sign-off pile next quarter.
A classifier can be accurate on every account in your ledger and still produce a noncompliant disclosure. The judgment it cannot make sits outside the data entirely: "regularly provided to the CODM" is a fact about management process, not ledger structure. If your chief operating decision maker actually reviews a condensed segment P&L that FP&A assembles in a spreadsheet, that construction — not the general ledger — defines the compliant category set. No training corpus contains your CODM's behavior, so no amount of per-account accuracy guarantees the buckets match what the CODM truly sees.
| Failure mode | Why the standard gate misses it | Adjustment that keeps the rule intact |
|---|---|---|
| Correlated errors below the line | Gate treats sub-materiality accounts as independent; drift in one direction accumulates | Full-population tie-out of auto-accepted matches against the trial balance, every quarter |
| Accounts opened mid-year | Confidence is least reliable where transaction history is thinnest | Route every new or renamed account to the named reviewer regardless of score |
| Account adjacent to the materiality line | A correct match can still push a bucket across its significance threshold | Widen the review band near the line; the score alone never decides proximity cases |
| CODM package changes mid-year | Retrospective consistency breaks quietly when categories move | Re-map prior interim periods and disclose the revised basis in the next 10-Q |
| Mapping is not the bottleneck | Hours sit in accruals, cut-off, and intercompany, not bucketing | Diagnose the close first; book modest savings in the business case |
| Filer under heightened scrutiny | Tolerance for draft-quality disclosure drops toward zero | Treat the entire map as review-required for that cycle |
The retrospective requirement compounds this. ASU 2023-07 applies retrospectively, so an AI mapping that quietly improves quarter over quarter breaks comparative integrity: each period looks locally correct while the series violates the standard. That is why a frozen mapping version matters more than any per-quarter accuracy score — run the identical map against both comparative columns, then document and sign any change. Here the comfortable claim that an existing account-group mapping table "already satisfies ASU 2023-07 — it's just a footnote" fails twice: interim effectiveness turned a one-time mapping project into a recurring quarterly control that must stay retrospectively consistent, and the standard demands the CODM's title and how they use the measure, neither of which any mapping table contains.

Why High Confidence Still Fails the 10-Q
Benchmark scores also transfer worse than vendors imply. Exam-style results were earned on standardized, well-worded questions; real charts of accounts end in terse, ambiguous tails — "misc," "clearing," "other operating" — where model confidence stays high precisely when correctness is lowest. Those captions carry almost no semantic signal, so the model falls back on priors and states them fluently. Worse, the tail concentrates dollar risk, because catch-all accounts are exactly where unclassified spend accumulates.
And there is no independent yardstick to check any of it. Unlike AP-invoice extraction, segment-expense mapping has no public GAAP test set, so every accuracy percentage a salesperson quotes is self-reported marketing until reproduced on your own prior-year accounts. According to research published on arXiv on LLM processing of segment disclosures, structured database support for longitudinal and cross-firm comparability of segment data remains limited — the verification layer simply does not exist yet.
Gains vary sharply by filer type. Single-segment companies get the least: smaller charts, one category set, and disclosures they had to prepare anyway. Multi-segment manufacturers with shared-service allocations get the most — and hit the hardest limit, because allocation logic (IT and facilities charged across segments) is a costing decision a classifier cannot make, only inherit. Whatever allocation basis your cost-accounting team chose, the model reproduces its output; it never evaluates it.
Finally, treat probability as fit-to-caption, not truth. An LLM can map a misleadingly described account — say, "Consulting services" carrying allocated shared-service overhead — into professional fees at 0.96 confidence, because the caption itself supports the error. Probability is not correctness, which is why even auto-accepted buckets need periodic sampling audits rather than unconditional trust.
The working triage:
Concrete next step for the current 2026 interim cycle: freeze the mapping version ID in the close workpapers, rerun the prior-year comparative accounts through that frozen version unchanged, and put a recurring sample audit of auto-accepted buckets on the close checklist — signed by the same named preparer and reviewer who own the footnote.
Start with the pain, because it explains the design. The first manual interim mapping consumed 52 controller-hours, tripped over 37 accounts created since the 10-K, and surfaced two late reclasses during review. Nothing about that quarter was anomalous — it is simply what a manual map produces at interim cadence. With ASU 2023-07's interim requirements phasing in, calendar-year filers began repeating that failure pattern for real in their 2026 10-Qs.
| Account archetype | Model behavior | Why the score misleads | Required handling |
|---|---|---|---|
| Clean captions ("Salaries — East region sales") | Confidence typically ≥0.95 and honest | The label itself carries real signal | Auto-accept only below the segment-materiality line |
| Catch-all tails ("Misc," "Other operating") | Confidence stays high while correctness bottoms out | Caption carries almost no signal; priors fill the gap | Named preparer-reviewer sign-off regardless of score |
| Clearing and suspense accounts | Confidently routed to a plausible bucket | Net-to-zero balances slip past dollar-materiality screens | Force review; verify ending disposition |
| Shared-service allocations (IT, facilities) | Inherits whatever allocation basis it infers | Allocation is a costing decision, not a classification one | Preparer-reviewer sign-off on the inherited logic |
| Misleadingly described accounts | Wrong category at 0.96 confidence | Probability measures textual fit, not ledger truth | Standing quarterly sample audits even on auto-accepted buckets |
The gated run inverted the labor curve. Embedding 24 months of posting history let the model draft a complete map overnight: one lane of accounts cleared auto-acceptance at ≥0.95 confidence while sitting below the segment-materiality line, a second lane routed to analyst review, and 46 accounts escalated to the controller. The three lanes partitioned the full chart, and every gate crossing carried documentation — the drafter-not-poster rule executed as bookkeeping rather than philosophy.

Meridian Components
The outcome metrics confirm the shape of the win, and one detail kills the comfortable myth that an existing account-group mapping table already satisfies ASU 2023-07. The disclosure note's CODM-use paragraph was drafted forward from the frozen category definitions rather than reverse-engineered at filing time — and no legacy mapping table contains the CODM's title or an explanation of how the segment measure is actually used. Interim effectiveness converted a one-time mapping project into a recurring quarterly control that must stay retrospectively consistent; a static table satisfies neither half.
That last row closes the loop. Because the run stored its model version, prompts, confidence scores, and reviewer IDs account by account, the external auditor could sample directly from the auto-accepted population and find nothing to rework. The AI step stopped being an audit liability and became inspectable evidence. Read the sequence as the thesis in miniature: the halving was real, an
Quick answers
| What three disclosure requirements does ASU 2023-07 impose that must be sourced from a general ledger? | Disclosure of significant segment expenses regularly provided to the CODM, an "other segment items" catch-all, and the CODM's title or position plus an explanation of how the CODM uses the profit measure. |
| Does ASU 2023-07 apply to entities with only one reportable segment? | Yes, it explicitly includes single-reportable-segment entities — the clause that surprises filers who assumed segments were someone else's problem. |
| How many expense buckets does the running case map the entire chart of accounts into? | Eight buckets, consistent with the standard asking accounts to collapse into roughly five to twelve significant expense categories per segment. |
| What are the six stages of the mapping pipeline? | Export, Embed, Classify, Score, Route, and Publish, each producing a named artifact from the account-level posting corpus through a versioned account-to-category table. |
| What confidence score did the account captioned "Cloud Subscriptions" receive and what happened to it? | It scored 0.97 against the "Software and cloud costs" category and posted automatically because it cleared both halves of the gate. |