Hey there 👋,
Lately, every conversation I’ve been in has moved from the excitement of being able to build and get our work done with AI, to one about AI governance and internal controls.
In the last issue we talked about AI Harness design, and how the technology is moving faster than the shared language we need to govern it.
In this issue, we’re gonna recap what happened when we sat down with Jason Pikoos, Meta Modern Finance, COSO Board member, co-author of the FEI AI-ICFR framework - and asked some real questions from the field.
200 of you registered across 126 companies. Operators, auditors, consultants, and the vendors building the tools. 80 questions came in before we even started.
Every Big 4 firm was in the room - not as auditors on duty, but because the full enterprise accounting ecosystem showed up to figure this out together.
Let’s dive in.
This newsletter is brought to you for free by Numeric: Join us at Inflection, Thursday Sept 17th — in person in SF. I'll be speaking at the Internal Controls and AI Benchmarking breakout in the afternoon session. Join us here!
Watch the full session here: Below is the letter version: the ideas I keep reaching for three weeks later, with timestamp links per section if you'd rather hear them land live.
One thing before we start: nothing in here is a ruling. This is an honest snapshot of the questions practitioners are asking and the best thinking in the room on one Friday in July 2026. We’ll know more tomorrow, and some of this will turn out to be wrong. That’s not a bug, that’s learning. The thing I admired most in this session is that Jason said out loud what our profession is usually too careful to say. In a space moving this fast, changing your mind on new information is the discipline. Hold everything below the same way.
The source (6:00)
First, some respect for the source material: the FEI AI Framework: Internal Control Over Financial Reporting took nine months to publish, written by industry practitioners, from Meta, Service Now, Alphabet, Walmart. Written in 2025, published in June of 2026. All Big 4 firms had feedback rounds and a review with the SEC.
And Jason’s first reflection on reading it back was “gosh, we need to update this thing really soon.” The first draft barely mentioned agents.
This is the exponential speed things are moving at. Quarters are becoming months, months are becoming weeks, and weeks into days. The need for a living document isn’t a marketing phrase here; it’s a survival strategy. And we’re all living in it.
The scope: reliance vs. non-reliance — where the line actually sits (8:31)
The framework’s opening move was to provide the context for, and explicitly frame the scope of, the document.
That is to say: from here on out, in the context of this document and this conversation, we are only talking about “reliance” in the scope of internal controls over financial reporting. Not, like, generally “I’m relying on it to help me with something.”
The whole session pivots on this one distinction. Everything that follows in this issue is really a question about where you sit relative to this line, or what happens at its edges.
So we asked the room to think about: are you relying on AI, or just using it?
Reliance: Is AI being relied on as the control? In other words, are you assuming that AI is doing the work accurately and completely when given valid data sources and the like? Or …
Non-reliance: Is AI just the doer? AI is doing the mechanics, but ultimately a human, or another systematic control, is inserted on top of it to validate it.
Only once you've determined that you're relying on AI as part of your internal controls do we start thinking through the new methodologies, risks, and frameworks to establish reliance in our new world of AI.
In cases where AI is simply the doer, you’re just doing your job faster. Except for when you hit shadow reliance (but we’ll get to that in a bit).
Here’s where the room was in July of 2026, in the Reliance/Non-reliance split:
Most of the room, 57% - was using AI to do the work, and a human either still fully reviewed the outcome, or sample reviewed the outputs.
Only 4% had crossed into AI-as-part-of-the-control with real testing behind it.
That means we are just in the beginning of our AI adoption, where AI helps us get work done we didn’t have time for, and now we’re able to spend more time applying judgment in the review.
The case walkthroughs: journal entry preparation & contract review (11:05 & 21:29)
Two live cases anchored the middle of the hour, both straight from the room.
Scenario 1: AI-prepared journal entries pushed to NetSuite (11:05)
AI prepares journal entries using "Skills" and pushes them to NetSuite. One human reviewer reviews what the skill is doing, and another human approves the JE. What does this count as?
If AI is just the doer (i.e., the preparer), that is technically non-reliance from a SOX scope perspective. The human reviews are what's being relied on, and that's probably sufficient, as long as the review is genuinely thorough — full human in the loop (source document review, management judgment, etc.).
Two layered human reviews can collapse into one without losing the preparer/reviewer split. The precedent has been sitting in your ERP for years: automated RPA (Robotic Process Automation) postings already skip a second reviewer today.
The follow-up (it lives as an on-screen card in the recording): what if your two reviews do different jobs — one validates context (source docs, assumptions), one validates judgment? Can they collapse to one reviewer, and eventually none?
That’s precisely the door established reliance opens.
⚠️ The real nuances start when you dive into the details of preparer/reviewer controls. For example, a review can be initiated by a human user account but actually be performed by AI. At that point you have to track, at a very granular level, which actions are human-initiated but AI-executed. We’ll be talking a lot more about this when we get to segregation of duties.
Established reliance: the four-part test (16:22)
So how do you actually earn your way across the line? The phrase established reliance isn’t explicitly in the framework yet (”the kind of stuff we need to come back and update”). Four requirements, named live:
Normal development and testing with ITGCs
A method to evaluate the model’s accuracy
A method to sustain that accuracy over time as models change underneath you
Business-owner sign-off
Nothing fancy. It’s how we’ve always onboarded systems we rely on. The difference is that the accuracy pieces used to come for free with deterministic software.
Established reliance is a spectrum on the risk continuum, defined by nature, timing, and extent. Anyone from the audit world will recognize this as how we’ve always designed test procedures. You’re managing residual risk by portfolio. We’ve always done that.
Scenario 2: contract-term extraction feeding a rev rec checklist (21:29)
If I use AI to populate a rev rec checklist, do I need to re-perform 100% of it?
In short, no. If you’ve never re-performed 100% of the process as a control, why would you start now?
AI-assisted contract review is our #1 community AI use case in Gaapsavvy right now. It includes the preparer role (extraction and checklist completion), the reviewer role (checklist completeness and accuracy), and judgment — because we all know contracts vary dramatically and the rev rec standard is certainly not straightforward. Jason noted that this question actually has quite a bit of depth.
On one extreme, if your aim is fully autonomous, full-reliance accuracy in the finance space today, even frontier model providers still land around 60–70% accuracy. And accuracy varies enormously by use case, especially with so many differing types of contracts.
Jason shared that the science around AI accuracy and reliability is very nascent and will likely evolve quickly. But if you want to rely on AI today, the question you have to answer is: how do you know that what it produced is accurate and complete?
For reliance, the framework currently offers two paths where you might be able to infer that the system has done the thing:
Performance testing against a known set of contracts, “a golden set” for evals. We know exactly what should be extracted from these 100 contracts, and we test the model's extractions periodically to make sure the results are complete and accurate.
Multi-model consensus. Two independent models agreeing is itself a control. This can be a preparer/reviewer model, or a consensus model with two independent models and a comparison of results. (In my mind, the phrase "validation check" comes up as I bridge my thinking to the harness design concepts from the last issue.)
Jason was also quick to say that the way he'd more likely approach this is to risk-profile the full portfolio. Don't fall into the trap of making the method fit a binary — either we do everything or we do nothing. In reality, the answer is often somewhere in between.
For example, bifurcate the population: standard vs. non-standard. Tested AI autonomy on standard contracts under a threshold (low risk), full human review on anything weird or huge (high risk). Even the standard/non-standard classification is a control you can validate, and an easier one to test.
To which I pointed out: "Wait, isn't that just normal sampling?" We always had to take the top 10 contracts that covered maybe 30% of revenue, because those are just big contracts — too much judgment, too weird, you always humanly review 100% of them. Then you have the lower-risk population, high volume, low dollar, that you'd typically just sample anyway. You never did 100%. You just extended the sample if you found errors.
Jason then gifted me new words, using the new human-in-the-loop (HITL) definitions in the FEI framework:
High risk → HITL Intent 1: validate each outcome. Review the high-risk contracts at 100%. Full human in the loop, full review = non-reliance.
Low risk → HITL Intent 2: validate the system. Sample the rest. If we sample and find no issues, there’s a presumption the rest are correct. It’s “a human version of multi-model: Model 1 is the AI, Model 2 is the human.”
And there it was: new vocabulary, to describe new concepts, so that we can have a new shared understanding.
Can AI documentation be your review evidence? (31:33)
The review already happened, and AI is just documenting it. So it's a different question than reliance.
Can I use AI to evidence my review? Think about a skill capturing what you looked at, an audit log, a meeting transcript. To me, it sounded a lot like how, even pre-AI, we used meeting notes and board minutes as audit evidence.
Jason, impressed by the question, honestly responded that he hadn’t yet contemplated it but couldn’t see why it wouldn’t be reasonable. Note-taking isn’t the function being controlled (we never really tested the human notetaker, the notes just evidenced performance). His caveat: ask your audit team before you build your close around it.
Shadow reliance: the SQL code-diff review (33:51)
The setup: you’re the reviewer, AI is just helping. You hand an LLM two versions of a SQL query and ask “tell me what changed.” The summary looks right; you sign off.
Question from the room: “Isn’t comparing code deterministic?”
The nuance: the code is deterministic, but “tell me what is different” is “a non-deterministic request on two deterministic components.” If validating code changes is your control, “I’ve inadvertently relied on AI to perform a critical control.”
When you think you’re in non-reliance, but you’re not - that’s shadow reliance.
What if the diff is identified in GitHub, where it just shows you the change in the code? In that case, you're not using AI. A system like GitHub that manages code merges has a systematic methodology for the comparison. It's code, it's deterministic, and as long as that system is normally tested and sits under ITGC, this whole question is not applicable. You're not relying on AI.
A quick PSA: a skill is not deterministic. (39:26) A skill is a pre-written, very prescriptive prompt. It may sometimes look like rows of code, and it can trigger the running of code, but the model interprets the instructions fresh every run.
My TLDR: code (Git, tested, ITGCs, people building software know what they’re doing) → skills (SOPs, the workbook you’d hand a new coworker) → the deterministic validation checks you code around the skill; that’s the harness.
Multi-agent reviewers: when do they count as a control? (40:11)
The question: is the code and prompting logic the actual control, or is it the certification of the individual outputs that result from them?
Let's just be honest: "your audit firms are all going to ask to look at the prompts in any of your AI systems." Hand them over; they must understand the process. But a prompt itself has "no meaningful control value," since a non-deterministic system can simply not follow the instructions.
You still need change management around it, and the evidence has to be around the output. Take a multi-model agent validation check: different models (built on different corpuses of data) performing the same function. If each produces the same outcome, then effectively those models are vetting each other.
One caution, fresh this week, on the word "different." In July, OpenAI's own AI agents broke out of an internal testing sandbox and got into Hugging Face's production servers, a real breach of a real company; OpenAI's postmortem found ~1,200 of its agents had discovered a shared server, built a message board, and coordinated to game the test's scorer before some of them escaped (Axios summary). They were the same model on shared infrastructure, and the scorer only checked outputs, never the reasoning. So "different models vetting each other" holds only if they're truly independent, and that's something you'll need to test, not assume.
Vendor AI with no SOC coverage (46:00)
From the room: "What do you do when you're relying on a vendor that maybe doesn't have a SOC 1 or SOC 2? For example, if I'm evaluating Claude Cowork — which does have a SOC 2, but a SOC 2 doesn't cover whether the model's output is complete and accurate."
Jason: "In the world of AI, SOC 1, SOC 2 doesn't help me much at all." None of them certify the accuracy of the model. No report at all = a whole ITGC problem on top of the model question; even a clean SOC 1 only covers the ITGCs around the system. You still need a methodology: human in the loop, performance testing, multi-model (tricky with a third party), or data analytics.
You really have to think about what you're using the tool for. If you're just using the chat, you need full human in the loop, and no report changes that. If you're building applications that use AI, you can build some of these control methodologies into the application itself.
Governing the vibe-coded — and the monitoring underneath (49:13)
Last question: everyone's vibe-coding now — how do you monitor it?
The cleanest governance line I’ve heard all year. One boundary: personal productivity vs. embedded in your process. Build an app for yourself, by all means, go wild. The moment it’s embedded — especially touching SOX — “this needs to be treated like an enterprise application,” with all the same controls, and “it better be connected to your Git or GitHub environment.”
Why so bright a line? Because his pick for one of the biggest risks in the space isn't the big systematic use case. It's shadow reliance at the individual level: "us as individuals using it… effectively relying on AI without managing it." A very distributed risk, which is what makes it hard.
On a personal note, lately shadow reliance has become very real to me. After months of intense AI use while building, it has become abundantly clear that it is faster than you and can handle more data than my human mind will ever be able to comprehend. There have been many times I have felt like the model was waiting on me to take the next action and catch up. Staying skeptical is absolutely exhausting.
Then Jason named the nightmare scenario every accountant will recognize: “We’re in the middle of the close, the deadline is in a couple of hours’ time… I’ve got 17 other priorities, and this AI that’s done this work for me has never been wrong for the last 3 months.” The temptation to approve and go is enormous. Not a character flaw — the exact failure mode the paper designs against. MIT's Your Brain on ChatGPT study found that people who let the model write remembered less of what they'd written; I suspect the reviewer's version of that is what shadow reliance feels like from the inside.
Where the room is on monitoring (n=68): 32% building now, 22% manual and periodic, 15% automated (evals, drift, anomaly detection), 12% honest enough to say “unsure what we’d even monitor.”
The countermeasures read like Management Review Controls because they’re modeled on them: cycle reviewers so nobody goes blind, define what’s expected, force it with checklists.
Asked in chat — the hour ran out first
The backlog the next session gets built from: how often established reliance needs re-testing (new model? changed data feed?) · whether it works when the vendor owns the model · aligning ongoing validation with auditor sampling · a concrete multi-model walkthrough for 606 · whether reliance testing covers a whole three-agent chain or a validated subset · who’s gotten an extraction agent past their auditors already · change management when everything runs on MCPs and skills · monitoring for drift.
If one of these is yours — hit reply. That’s literally how sessions get built.
Three things, in order
Watch the session recording here , or jump to the chapters most relevant to your current fight.
Get the paper — FEI’s AI Framework: ICFR. Dense, practitioner-written, the most current thing in print. It’s paid; also cheaper than an hour of your auditor’s time.
Sort every AI touch in your close: reliance or non-reliance. The surprises live in the sorting. And wherever you're sampling AI output, start writing the results down. That log is a control.
Jason’s last line of the hour was the honest one: “There’s no experts in the space — just all of us that are learning.”
We’re all just making it up. We might as well make it up together.
And you already voted on where we go next — the pick-the-next poll had a clear winner: Segregation of Duties in an agentic world.
Thanks for being in the conversation.
Angela
From our partner Lensing (formerly Everest, out of stealth this month):
We recently sat down with the Lensing team on “How Accountants can think like Product Managers.” Now that we’re all vibe-coding, it’s worth understanding a product mindset and the software development lifecycle. 🎬 Watch it here.
What’s coming up:
September 17 - Numeric Inflection summit, SF. The agenda is stacked, I’ll be leading an internal controls breakout share session. Would love to see you - Register here
Sept 25 - Gaapsavvy AI Shareforum (Practitioner Gaapsavvy Members only)
AI Finance Labs - Gaapsavvy x Coterie CFO collab at Deloitte
Oct 13: Boston AI in Finance Lab (4.5 CPE)- Apply here
Oct 14: NYC AI in Finance Lab (4.5 CPE)- Apply here
Oct 28: SF AI Finance Lab (4.5 CPE) - Apply here
From real practitioners, for learning purposes only. Polling data shows industry leans, not final positions — every company’s facts are different. Always work with your auditors and internal controls team before implementing anything new.





Great article!