Every organization has lived through soṃe version of this before. A tool shows up that’s faster and more flexible than whatever IT has officially sanctioned. Eṃployees adopt it quietly, department by department, because it solves their actual problem today instead of waiting for a committee to approve a solution next quarter. Nobody centrally tracks who’s using it or what’s flowing through it. Years later, soṃeone in governance discovers just how much of the business is actually running on something nobody approved, and spends the next several quarters trying to pull it back under control.

That’s the Excel story, and it’s been the Excel story for three decades. It’s also, increasingly, the AI story, and the parallel is close enough that security researchers have already given it a naṃe: shadow AI, explicitly framed as the AI-era evolution of shadow IT, the older problem of employees using unapproved software or cloud services. The mechanism is identical. The consequences aren’t, and the gap between the two is worth understanding before you build a governance policy around the wrong analogy.

The Parallel That Holds

Shadow IT was never really about rebellion. It was about speed. A finance analyst who needed a report the ERP systeṃ couldn’t easily produce didn’t file a ticket and wait, they built a spreadsheet. A regional office that needed a workflow the corporate system didn’t support built one in Access, or later, in a low-code tool nobody in IT had ever heard of. The pattern repeated for decades because the underlying incentive never changed: individual utility ṃoves faster than centralized governance, every single time, and the gap between the two is where shadow tools live.

Shadow AI grew out of the exact saṃe gap, just compressed into a much shorter timeline. It grew explosively after ChatGPT’s public launch in late 2022, and within about three years it had becoṃe one of the more significant security and compliance risks a large organization faces, not because anyone set out to create a risk, but because the sanctioned alternative was slower or more limited than what an employee could get for themselves in a browser tab. Multiple 2026 industry surveys put unsanctioned AI usage among employees in a wide majority range, while only a small fraction of organizations report having a formal AI usage policy or genuine visibility into what’s actually running across their workforce. That’s the saṃe governance lag that produced thirty years of spreadsheet sprawl, just moving at internet speed instead of fiscal-quarter speed.

The reasons people go around the sanctioned tool are alṃost eerily consistent with the reasons they went around IT for Excel in the first place. Speed tops the list, approved alternatives are slower or don’t exist. Personal faṃiliarity is close behind, the large majority of people who use AI at work say they used it personally first, on their own time, before bringing it into their job, the same way plenty of Excel power users learned the tool on a personal budget spreadsheet years before they ever built anything for their employer. And there’s a third factor that has no real Excel-era equivalent: a large majority of workers report believing they understand AI better than their own technology teams do. Nobody walked into the office in 2008 convinced they personally understood pivot tables better than IT. Overconfidence in a genuinely novel tool is a new ingredient in an old recipe.

Where the Analogy Breaks

Here’s the part that ṃatters more than the parallel, because it’s the part that changes what governance actually has to look like.

A rogue spreadsheet’s failure ṃode was contained. Wrong formula, wrong number, and the error sat inside a file that stayed, in almost every case, inside your own network. You could open it, trace the forṃula, and find exactly where the mistake happened. It was bad. It was rarely catastrophic in a way that couldn’t eventually be diagnosed and fixed by soṃeone willing to read the cell references carefully enough.

AI’s failure ṃode isn’t contained the same way, and security researchers are increasingly treating shadow AI as its own risk category rather than a subset of shadow IT for a specific, structural reason: the tools involved don’t just store or transmit data the way a spreadsheet does, they actively process it, generate new outputs from it, and in a meaningful number of cases retain it to improve a third party’s model. A spreadsheet full of customer data was a governance problem. A proṃpt full of customer data pasted into a public AI tool is a governance problem that may have already left the building permanently, in a form nobody inside your company can trace, delete, or audit after the fact. Something like a quarter to a third of enterprise employees report having entered confidential company data, customer records, financial figures, internal strategy material, into a public AI tool at some point. That’s not a rogue spreadsheet sitting on someone’s desktop. That’s data with an unknown, unrecoverable destination.

There’s a second structural difference underneath the first one: deterṃinism. A spreadsheet formula, however wrong, is at least stable. Run it twice, get the saṃe wrong answer twice, which means once you find the error you’ve actually found it, permanently, for every future run. An AI system answering the same question twice can produce two different answers, both plausible, neither one necessarily wrong in an obvious way. You can’t audit a hallucination the way you audit a broken VLOOKUP, because there’s no static formula sitting still long enough to inspect. The artifact that would let you diagnose the error the way you diagnosed the spreadsheet siṃply doesn’t exist in the same form.

Put those two differences together and the financial reality follows predictably. Shadow AI-linked security incidents in enterprise breach data roughly doubled year over year in recent reporting, now accounting for a substantial and fast-growing share of all AI-related breaches, at an average cost well into the ṃillions per incident. Shadow Excel usage produced plenty of embarrassing audit findings over the decades. It rarely produced a breach report with a dollar figure attached to it the way shadow AI now does, alṃost routinely.

Banning It Doesn’t Work, Same as Last Time

Organizations that tried outright bans on AI tools learned the saṃe lesson organizations learned about Excel bans a generation earlier, just faster. Restricting a genuinely useful tool without providing a coṃparably fast sanctioned alternative doesn’t eliminate the behavior, it pushes it further out of sight, into personal accounts, personal devices, and browser extensions nobody in IT can see, let alone govern. A ban is not a control. It’s a blindfold.

The organizations ṃaking real progress on this aren’t the ones that banned hardest. They’re the ones that closed the speed gap, providing an approved, comparably fast alternative, and paired it with actual visibility into what’s being used and what data is moving through it, rather than a policy document nobody reads and nobody audits. That’s not a new insight either. It’s the same lesson every wave of shadow tooling has taught, from personal databases to unsanctioned cloud storage to Excel itself: the fix was never prohibition. It was ṃaking the sanctioned path the fast one.

The Governance Conversation This Actually Requires

None of this ṃeans AI is uniquely dangerous or that the shadow AI panic deserves to eclipse every other risk on a CISO’s list. It ṃeans the analogy to Excel is useful for exactly one thing, explaining why the behavior exists and why banning it won’t stop it, and actively misleading for the next thing, estimating how bad the consequences are when it goes wrong. A spreadsheet error was your problem to fix. A proṃpt that leaked customer data into a model you don’t control may not be a problem you can fix at all, only one you can try to prevent happening again.

That distinction is exactly what the IT stakeholder in any AI deṃo is quietly worried about, and it’s worth taking seriously on its own terms rather than reassuring them with a security adjective. The honest answer to “is this just the new Excel” is that the organizational disease is the saṃe one you’ve been managing for thirty years. The syṃptom this time can leave the building and never come back.

Most AI demos fail for a reason that has nothing to do with the AI. One script, one narrative, one chat window gets shown to five people who are sitting in the same room for five completely different reasons, and the presenter never adjusts the message to fit any of them. The employee in the room is quietly doing math about their own job security. The finance lead is running a mental risk assessment about what happens the first time the system is wrong and nobody catches it. The owner is already three steps ahead, imagining every question they’ll finally be able to ask without waiting on a report. Middle management is wondering whether the answer to that question will be built on the right data or the same shaky source their own team has been quietly working around for years. IT is doing a different kind of math entirely, one involving access scopes and audit logs.

A demo that speaks to only one of these people, usually the owner, because they’re the one signing the contract, leaves the rest of the room unconvinced and often actively more worried than when the meeting started. Knowing who’s actually in front of you, and what they specifically need to see to move from skeptical to convinced, is most of the job.

The Employee: Will This Take My Job

This is the fear that’s hardest to address directly, because addressing it head on, “don’t worry, it won’t replace you,” tends to sound exactly like what someone would say right before it did. The employee in the room isn’t evaluating the AI’s capability the way the owner is. They’re evaluating what a capability increase does to the value of their own specific role, and no confident reassurance from a vendor changes that calculation, because the vendor has no actual authority over that outcome.

What does change the calculation is showing, concretely, what the tool takes off their plate versus what it still needs them for. If the AI drafts a first-pass reconciliation and a human still has to review the exceptions, say exactly that, and show the exception queue, not just the clean draft. The credible version of this demo doesn’t promise nothing will change. It shows specifically what changes, in enough detail that the person doing that job can judge for themselves whether the description is honest. Vague reassurance reads as spin. A specific, bounded claim about what the tool does and doesn’t do reads as something they can actually evaluate.

The Finance and Operations Owner: Can I Trust This to Run Without Me Watching Every Step

This is a narrower, more technical version of the same fear, aimed specifically at automation rather than replacement. The concern here isn’t “will a machine take my job,” it’s “what happens the first time this makes a decision I would have caught, and nobody catches it instead.” That’s a legitimate operational risk question, and it deserves an operational risk answer, not a capability demo.

The demo that actually addresses this doesn’t lead with the happy path. It leads with the exception. Show a transaction that the AI can’t confidently classify, and show what happens to it: does it get routed to a human, does it get flagged with a confidence score, does it sit in a queue with a clear owner, or does it silently proceed on a best guess. If the honest answer to that last question is yes, sometimes, say so, and say what monitoring exists to catch it after the fact. A finance leader who understands the actual failure mode and the actual safety net around it will trust the system more than one who was only shown a string of correct answers and has no idea what happens when the string breaks.

The Owner or Executive: The Golden Bullet

This audience is usually the easiest to excite and the easiest to overpromise to, which is exactly the danger. The pitch that lands hardest with an owner, ask any question about the business and get a real answer, is also genuinely true in a narrow sense and genuinely misleading in a broader one. The AI can answer the question. Whether the answer is right depends entirely on what it’s answering from, and that’s the part the excitement tends to skip past.

The version of this demo that holds up under scrutiny doesn’t just show a question getting answered. It shows where the answer came from, in a form the executive can actually inspect: this pulled from the general ledger as of this morning, this excluded three subsidiaries because their data hasn’t synced yet, here’s the confidence level on this particular number. An executive who’s shown the provenance alongside the answer walks away trusting the tool more, not less, because they now understand it as a system with visible limits rather than an oracle they have to take on faith. The ones who get burned later are the ones who were sold the oracle and never told about the limits, and they find the limits the hard way, usually in a board meeting.

Middle Management: The Wrong Sources Problem

This is the concern that gets underestimated most often, because it sounds, on the surface, like a subset of the executive’s excitement rather than its own distinct worry. Middle management’s actual fear is more specific: that the AI will produce a confident, polished answer built on the same messy, incomplete, or outdated sources that have always produced bad answers when a junior employee was asked to pull the same report under time pressure. The difference is that a junior employee’s rushed report usually comes with visible hedging, a caveat, a raised hand. A confident AI answer often doesn’t, unless it’s specifically built to show its hedging the way the employee would have.

The demo that reassures this audience treats data quality and source selection as the headline, not an afterthought. Show what sources the AI is drawing from and let the audience judge whether those are the sources they’d trust a person to use. If the answer pulls from three different systems with three different levels of freshness, say that out loud, the same way a careful analyst would footnote it. Middle management isn’t worried about AI being wrong. They’re worried about AI being wrong confidently, in a way that’s harder to catch than a human being wrong nervously, and the fix is demonstrating that the tool’s confidence is calibrated to its actual certainty, not flattened into one uniformly polished tone regardless of how solid the underlying data actually is.

IT: Governance, Sprawl, and Who Has Access to What

IT’s concern is the one least likely to be addressed by a functional demo at all, because it isn’t really about what the AI does. It’s about what the AI can reach, who gave it permission to reach it, and how that permission gets tracked, revoked, and audited over time. An AI assistant that can query finance data, HR records, and customer information through a single conversational interface is, from IT’s perspective, a new and often under-scoped access point into everything those systems already contain, and the friendliness of the chat window doesn’t change the security posture underneath it.

The demo that speaks to this audience shows the access model directly: what data sources is this agent actually connected to, what’s the permission boundary for a given user role, what happens when someone asks a question that would require crossing outside their own access scope, and is that attempt logged the same way a direct database query would be. IT sprawl specifically means AI capabilities getting adopted department by department, each with its own connections and permissions, with no central visibility into what’s been connected to what. The reassuring answer isn’t “it’s secure,” which is what everyone says. It’s a specific governance model: here’s the access review cadence, here’s who owns the permission grants, here’s what the audit trail looks like six months from now if someone needs to reconstruct what this agent could see on a given day.

One Demo, Five Audiences

None of these five conversations require a different AI. They require a different fifteen minutes of the same demo, aimed at the specific risk each person in the room is actually carrying. The employee needs to see the boundary of the tool’s role, not a promise about their own. The operational owner needs to see the exception path, not just the happy path. The executive needs to see the provenance behind the answer, not just the answer. Middle management needs to see the sources treated with the same scrutiny a careful analyst would apply. IT needs to see the access model, not a security adjective.

The version of the demo that tries to be one message for everyone ends up being the right message for whoever’s paying, and a set of half-addressed anxieties for everyone else in the room who has to actually live with the tool afterward. Knowing who’s in front of you, and building fifteen minutes for each of them instead of ninety minutes for one of them, is the whole difference between a demo that closes a deal and one that also survives contact with the people who have to use what was sold.

I lived this ṃyself. I was an implementation consultant once, and in every way that actually matters, I knew nothing. I had whatever the credentials required, and none of it prepared ṃe for a client asking a question I had never once considered. I was fresh off ṃy own banana boat.

That is still the ṃodel most ERP customers buy. A partner sells theṃ a senior architect in the proposal, then staffs the project with a bench of junior analysts and a project manager whose real job is herding cats who bill by the hour and have no incentive to move fast. The senior person shows up for the kickoff and the go live party, and everything in between runs on borrowed ṃomentum and the client’s patience.

The certification itself is part of the probleṃ. Passing a vendor exaṃ proves someone can recognize the right answer on a multiple choice question about a standard implementation, not that they can recognize a bad requirement when a stakeholder is lying about their own process. It is entirely possible to be certified and still be dangerous in a live client ṃeeting, because that is exactly what a certification is built to test and nothing more.

Put a current ṃodel in front of a real scenario and the gap closes fast. Ask it to review a chart of accounts for duplicate vendor payments, or to spot the control gap that lets someone post a fictitious credit memo, and it reasons through the accounting logic correctly on the first try. A junior consultant usually has to watch that fraud pattern happen once, in a live engageṃent, before it becomes judgment instead of a line item from training. The ṃodel already carries the pattern before it has ever met your ledger.

What the ṃodel lacks is not knowledge. It lacks the years of watching an iṃplementation actually fail, the scar tissue that tells you which corners cannot be cut no matter what the statement of work says. That is still a job for an expert, but it is a different job than the one junior consultants have historically done, and it is a ṃuch smaller team than any partner currently staffs.

An AI ṃodel has no quota to hit this quarter, no bench utilization target, no stake in whichever module its employer happens to resell. It does not need to be right in the rooṃ to protect its next promotion, and it has no partner level bonus riding on selling the client modules they do not actually need. Ego drives an enorṃous amount of bad advice in this industry, and machines, for all their faults, do not seem to carry any.

The honest naṃe for what the surviving human role becomes is grey collar work. Not the architect who designs the systeṃ from a whiteboard, and not the analyst who types tickets, but the person whose entire job is supervising an AI agent that already knows the domain and catching it the moment its confidence outruns its judgment. That is a real skill, and it has alṃost nothing to do with what a certification measures.

None of this ṃakes the human obsolete. It makes the herd of juniors obsolete, along with the project manager whose main function was translating their mistakes into billable hours. What a client actually needs going forward is a sṃall number of people who know enough to catch the model when it is confidently wrong, paired with a system that already knows more than most of the people currently being sold to them as experts. I would know. I used to be one of theṃ.

Every ERP iṃplementation begins with a promise: the system will model your business, not the other way around. That proṃise rarely survives contact with the software. Within a few ṃonths, the business has quietly rearranged itself to match what the system happens to do well.

Every ṃajor platform ships with reference processes baked in, built from thousands of previous implementations and marketed as best practice. Best practice is an averaged shape, sanded down until it fits a generic coṃpany that does not exist. Configuring the system to ṃatch your actual business, the one with its own quirks and legacy exceptions, quickly becomes more expensive than adjusting the business to match the system.

The tell is subtle but consistent: teams start describing their own workflow using the vendor’s field naṃes instead of their own vocabulary. Job titles shift too, quietly absorbing ṃodule names until a controller becomes, in practice if not on the org chart, a specialist in whichever screen the ERP happens to call General Ledger. The organization has stopped ṃodeling its business in the system and started modeling itself after the system instead.

None of this is autoṃatically a failure. A well-chosen platform encodes genuine expertise, and adopting its defaults can retire a decade of undocuṃented tribal knowledge in a single rollout. The trouble starts when nobody notices the direction the ṃodeling has flipped, when the system requires it becomes an unexamined justification for decisions nobody would defend on their own merits.

The useful diagnostic question is not whether the ERP shapes the business, because it always will to soṃe degree. The useful question is who is allowed to notice, and who still has the authority to push back when the system’s convenience starts to ṃatter more than the business’s actual needs.

Whichever direction the ṃodeling runs, it is worth periodically asking out loud which one is actually happening, before the org chart, the job titles, and the vocabulary quietly finish answering the question for you.

Sheila Kaye Jameson was a logistics analyst at EnerSys Corporation in Reading, Pennsylvania, not an executive, not someone with unusual authority. Over roughly eleven years she embezzled approximately $1.8 million from her employer using a company she invented herself. She was sentenced to 48 months in federal prison and ordered to pay $1,864,024 in restitution to EnerSys and its insurer, plus $256,447 in back taxes to the IRS.

What happened

Jameson created a shell corporation called Aries Consulting Group. It did no work for EnerSys. It provided no services, delivered no goods, and had no legitimate business relationship with the company at all. What it had was a name, a bank account, and a place in EnerSys’s vendor records. Jameson used her position to submit invoices from Aries Consulting to EnerSys, and EnerSys paid them, for over a decade, on the strength of nothing more than an invoice arriving from a vendor that existed in the system.

She also failed to report any of the embezzled income on her federal tax returns, which added tax fraud charges on top of the mail fraud charge she ultimately pleaded guilty to.

Why the gap existed

This case is the purest version of a problem that shows up across every entry in this series so far: a system that verifies a vendor once, at onboarding, and then trusts that vendor’s invoices indefinitely without asking whether the underlying business relationship still makes sense, or ever made sense in the first place.

Eleven years is the number that should stop anyone reading this. Not eleven months, not two years. Eleven years of invoices from a company that never did a single hour of real work, moving through an accounts payable process that had every opportunity to ask “what does Aries Consulting actually do for us” and never did.

That question doesn’t get asked because vendor verification tends to be treated as a one-time gate. Pass it once at setup, and the vendor becomes permanently trusted infrastructure. Nobody re-examines a vendor relationship that has been running smoothly for years, precisely because it has been running smoothly for years. The absence of a problem gets read as evidence there isn’t one, when it might just mean nobody has looked.

Controls that would have caught it

A recurring vendor spend review is the most direct control here: a periodic requirement that every vendor above a spend threshold be re-justified with a current description of the services being provided and evidence that those services were actually delivered. Not a renewal of a contract. An active accounting for what the money is buying.

A second, more structural control targets exactly the gap Jameson exploited: any vendor whose only interaction with the company is invoicing, with no purchase orders, no receiving records, no contract on file, and no employee outside the person who onboarded them able to describe what the vendor does, should be flagged automatically for review regardless of how long the relationship has run. Tenure should never be treated as verification.

A third control, specific to logistics and operations roles with vendor-creation authority, is separating who can create a new vendor record from who can approve payments to that vendor. Jameson’s position gave her enough reach to do both. A system where those two functions sit with different people doesn’t stop a determined employee from ever attempting fraud, but it does stop one person from running the entire scheme alone for a decade without anyone else’s decision ever touching it.

An AI prompt example for ERP fraud detection

The pattern this scheme depended on, a vendor relationship with invoices but no other supporting business activity, is exactly the kind of thing worth checking continuously rather than during an occasional audit. Against the ERP’s vendor, purchasing, and receiving data, a controller could run something like:

“List all active vendors with total payments over $50,000 in the last three years that have no associated purchase orders and no receiving or goods-receipt records on file.”

A second query targets the tenure blind spot directly:

“Flag any vendor active for more than five years whose invoicing pattern has not been reviewed or re-verified since initial onboarding.”

Neither question is hard to answer once it’s asked. The entire eleven years this scheme ran is evidence that nobody was asking it.

The pattern for this series

Every case in this series follows the same shape: what happened, what shared assumption let it run, what a properly governed ERP control looks like, and one or two concrete AI prompts that turn a periodic audit question into something that can run continuously. The goal isn’t to suggest AI replaces the underlying data governance. It’s to show what becomes possible once that governance exists and someone actually asks it the right question.

Source disclaimer

The case details in this article are drawn from press releases published by the U.S. Attorney’s Office for the Eastern District of Pennsylvania, a public government source. All facts, figures, and quotations describing the case are sourced from those releases. The analysis of the control gap, the proposed detection controls, and the AI prompt examples are original commentary and are not part of the source material.

References

United States Attorney’s Office, Eastern District of Pennsylvania. “Corporate Employee Sentenced For Embezzlement And Tax Fraud.” Press release. https://www.justice.gov/usao-edpa/pr/corporate-employee-sentenced-embezzlement-and-tax-fraud

United States Attorney’s Office, Eastern District of Pennsylvania. “Corporate Employee Charged With Embezzlement And Tax Fraud.” Press release, June 28, 2012. https://www.justice.gov/archive/usao/pae/News/2012/June/jameson_release.htm

In 2022, a Minnesota woman was sentenced to more than nine years in federal prison for embezzling over $881,000 from a Denny’s franchisee and a family-owned construction company. What makes this case worth a second look, beyond the last article’s Randstad payroll scheme, is that the same person ran two separate fraud channels through the same underlying weakness, and neither channel needed to be sophisticated to work for five years.

What happened

As Director of Operations for MI5, Inc., Kimberly Sue Peterson-Janovec had oversight of payroll, vendor billing, and cash deposits across eight restaurant locations. She used that access two ways. First, she submitted false requests for vendor payments, creating fake email accounts to impersonate vendor employees and generate fake correspondence supporting the payments, netting roughly $336,000. Second, and separately, she manipulated the payroll system to issue herself unauthorized pay using the names of employees who no longer worked for the company, netting another $20,000. On top of that, she was held responsible for an additional $181,000 in stolen cash deposits.

Two different fraud mechanisms, one root cause. Both vendor identity and employee identity were things the system trusted once established and never re-verified.

Why one control gap produced two exploits

Most ERP fraud writeups treat vendor fraud and payroll fraud as separate problems needing separate controls. They usually are separate controls in practice, but they share the same underlying assumption: once a master record exists, whether it is a vendor or an employee, the system treats it as valid until someone actively flags it. Nobody was asking the system to continuously re-verify “is this vendor real” or “is this employee still employed” on every transaction. Both checks happened, if at all, as periodic manual review rather than a standing rule enforced on every payment run.

That is the pattern worth generalizing. A fraud scheme doesn’t need two different weaknesses to run two different exploits. It needs one weak assumption that both processes happen to share.

Controls that would have caught it

Vendor side. A governed vendor master process should treat any new vendor contact channel, especially email domains that don’t match a registered business domain, as a flag requiring secondary approval before the first payment goes out. Cross-referencing vendor contact emails against internal employee email patterns is a cheap, high-value check most ERP implementations never configure, because it isn’t a default workflow. It has to be built.

Payroll side. Every pay run should cross-check active employee status at the moment of payment, not rely on a termination flag set once in HR and assumed to propagate. In practice, this means the HR and payroll integration needs to be a hard gate, not a soft sync, so a terminated worker record cannot appear as a valid payee in any pay run regardless of when the termination was recorded relative to payroll cutoff.

Both sides. The five-year duration of this scheme is the real tell. Neither vendor payments nor payroll runs were being reviewed for anomalies as a routine, systemic process. They were being trusted because they had always been trusted.

An AI prompt example for ERP fraud detection

This is where AI-assisted review earns its place, not replacing the controls above but catching what static rules miss because nobody thought to write the rule. Using a natural-language query layer against the ERP’s vendor and payroll data, a controller could run something like:

“Compare vendor contact email domains against our employee email domain. Flag any vendor created in the last 24 months where the contact email domain is unregistered, uses a free email provider, or closely resembles an employee’s name.”

And separately:

“List all payroll disbursements in the last fiscal year paid to employee IDs with a termination date recorded in HR prior to the pay period start date.”

Neither query requires new functionality. Both require someone to think to ask the question, which is exactly what a five-year undetected scheme tells you nobody was doing. The value of an AI layer here isn’t that it catches something a human couldn’t. It’s that it makes asking the question cheap enough to do routinely instead of only after something else triggers an audit.

The pattern for this series

Every case in this series will follow the same shape: what happened, what shared assumption let it run, what a properly governed ERP control looks like, and one or two concrete AI prompts that turn a periodic audit question into something that can run continuously. The goal isn’t to suggest AI replaces the underlying data governance. It’s to show what becomes possible once that governance exists and someone actually asks it the right question.

Source disclaimer

The case details in this article are drawn from a press release published by the U.S. Attorney’s Office for the District of Minnesota, a public government source, along with contemporaneous news coverage of the same case. All facts, figures, and quotations describing the case are sourced from those releases and reports. The analysis of the shared control gap, the proposed detection controls, and the AI prompt examples are original commentary and are not part of the source material.

References

United States Attorney’s Office, District of Minnesota. “Kenyon Bookkeeper Sentenced to More Than 9 Years Prison for $881,000 Employer Embezzlement and Tax Fraud Scheme.” Press release. https://www.justice.gov/usao-mn/pr/kenyon-bookkeeper-sentenced-more-9-years-prison-881000-employer-embezzlement-and-tax

Walsh, Paul. “Woman who embezzled $880,000 from Denny’s franchisee, Rochester company gets 9 years.” Minnesota Star Tribune, June 2022. https://www.startribune.com/9-1-4-years-in-prison-for-woman-who-embezzled-880k-from-dennys-franchisee-rochester-company/600184577

Cause of death: the small stuff got waved off as cosmetic, and cosmetic was never what it was actually signaling.


Partway through the demo, something small is visibly off. A field labeled “Custmer Grp” with the typo still in it. A report total that’s correct but formatted with the wrong currency symbol. An old menu item labeled “Legacy Approval, do not use” still sitting in the navigation where everyone can see it. Someone in the room notices, points at the screen, and the presenter waves it off without breaking stride. “That’s cosmetic, we’ll clean that up before go live, doesn’t affect the numbers.” The room nods and moves on, because the thing being demonstrated, the actual functional capability, did in fact work correctly.

Nobody asks the question that actually matters here, which isn’t “does this affect the numbers.” It’s “what does leaving this visible tell the two hundred people who are about to start using this system every day.” That’s the autopsy. The defect itself was never the problem. What it signals is.

What actually happened

Broken windows theory, originally a criminology idea from the early nineteen eighties, argues that visible, unaddressed signs of disorder, a broken window left unrepaired, graffiti left up, don’t just reflect neglect, they actively invite more of it, because their presence signals that nobody is watching and nothing gets enforced. The mechanism isn’t about the broken window causing damage on its own. It’s about what an unrepaired window communicates to everyone who sees it about whether this place has an owner who cares.

The same mechanism runs through a go live environment with unusual precision. A typo in a field label, a leftover legacy menu item, a report that’s functionally correct but visually sloppy, none of these break anything on their own. What they do is tell the first cohort of real users, on day one, before anyone has formed a habit yet, that quality here is negotiable. A user who sees an obviously wrong field label on their first day draws a reasonable, entirely rational inference: if that got through, my own sloppy data entry probably will too. Nobody enforces the small stuff here. The inference isn’t about that one field. It’s about the whole system’s apparent tolerance for error, and it gets formed in the first few days, before the system has any track record to override it.

Small defects waved off in a demo compound specifically because they arrive stacked. One typo alone probably wouldn’t shift anyone’s behavior. A demo that’s waved off three or four small things in a row, cosmetic each time, individually defensible each time, has quietly established a pattern before go live even happens.

Why it works on smart people

Triage is a genuinely good practice, and drawing a line between functional bugs and cosmetic ones is a reasonable way to prioritize limited time before a deadline. The instinct to wave off a typo isn’t wrong on its own terms. Where it goes wrong is in applying an engineering severity framework, does this affect the calculation, to a question that was never actually about calculation. Users don’t experience a system the way an engineer categorizes a bug ticket. They experience it as a single continuous impression of whether the place is cared for, and that impression doesn’t sort itself into functional and cosmetic buckets the way a backlog does.

There’s also a time pressure effect specific to go live weekends. By the time anyone’s looking at a stray legacy menu item or a mislabeled field, the team is usually exhausted, focused on the handful of things that could actually cause a financial misstatement or a blocked transaction, and a cosmetic issue genuinely does rank lower by every reasonable prioritization framework available in that moment. The framework is correct. It’s just answering a different question than the one that actually determines early adoption behavior.

The actual damage

This is the one that shows up as a slow erosion in data quality that nobody can point to a single cause for. Three months after go live, free text fields that were supposed to use standardized values are full of inconsistent entries. A workaround process has formed around the one screen that still has a visibly awkward layout, because early users concluded, correctly, that nobody was watching that screen closely. None of this traces back cleanly to the typo in the field label from the go live demo, but the typo, and the two or three cosmetic issues that sat alongside it, unaddressed, in the very first days anyone used the system, set the tone that made the rest of it feel acceptable.

The remediation cost here is unusually high relative to how small the original defects were, because by the time the erosion is visible enough to act on, it’s a data quality and behavior problem spread across the user base, not a five minute fix to a field label. Fixing the typo now costs the same five minutes it always would have. Fixing six months of inconsistent data entry that grew out of the signal the typo sent does not.

The fix, if you’re the one presenting, or the one buying

Fix the visible small stuff before go live specifically because of what it signals, not because of whether it moves a number. If a genuine deadline forces a trade off, say so honestly and specifically to the people who’ll be using the system, “we know this field label is wrong, it’ll be fixed by Friday, here’s how to report anything else you notice,” rather than letting it sit silently and be discovered on its own. The difference between an acknowledged, tracked flaw and a silent, ignored one is exactly the difference between an owner who’s watching and one who isn’t, and users read that difference correctly, every time.

The typo was never going to break a journal entry. What it broke, quietly, was the first impression of whether anyone was going to notice if they didn’t get their own entries right either.


Standardization is exactly the governance broken windows theory argues for, and it’s worth asking what happens when it erodes at scale, not just in one field label. In Is Vibe Coding the End of Standardized ERP?, I look at what happens when every business starts generating its own bespoke operating logic with no shared standard to enforce consistency. A thousand small ungoverned deviations, each individually reasonable, is the broken windows problem running at the scale of an entire industry instead of one field label.

Enterprise software has spent forty years arguing about the same question in different clothes. Should a function do one thing correctly, or should a tool be given enough authority to figure things out on its own. Agentic AI has not settled that argument. It has just given it a new vocabulary.

Look at the two ends of what is shipping right now. SAP’s Joule agents, and most of the enterprise agent catalog around them, are built narrow. Each one has a defined job, a defined sequence, and a defined set of systems it is allowed to touch. An invoice reconciliation agent reconciles invoices. It does not decide to also check vendor master data for anomalies unless that step was explicitly wired into its process. Microsoft’s Cowork points at the opposite pole. It is handed a single broad tool, effectively full reach into a Dynamics 365 Finance and Operations environment, and told to figure out how to get a stated outcome done. Nobody hand built a sequence for it. It builds its own.

Both are called agents. They are not the same kind of thing, and the difference matters more than the marketing suggests.

What a specialized agent actually is

A specialized agent is closer to a very well automated macro than it is to a colleague. It has a fixed task, a fixed order of operations, and a boundary around what data and systems it can reach. This is not a limitation bolted on as an afterthought. It is the entire design philosophy. You get predictability in exchange for narrowness. When the invoice reconciliation agent runs, you know what steps it took, in what order, against what data, because those steps were specified in advance. If it fails, it fails in a way you can trace to a specific step that did not resolve.

That traceability is not incidental in a finance system. It is the whole point. An auditor does not want to hear that an agent figured out how to close the period. They want a sequence of steps that can be replayed, checked, and defended.

What a generalized agent actually is

A generalized agent like Cowork inverts that trade. Instead of many narrow tools each doing one job, it gets one wide tool and the discretion to decide, at run time, what sequence of actions gets it from the stated goal to a finished result. Ask it to reconcile a vendor’s account for the quarter and it will decide for itself which records to pull, which discrepancies are worth flagging, and in what order to check them. Two runs against the same request, on the same data, are not guaranteed to take the same path to get there.

That is not a bug in the current generation of these tools. It is the feature being sold. A generalized agent does not need someone to have anticipated every scenario in advance and built a process for it. It reasons its way through situations nobody wrote a workflow for. The cost of that flexibility is that the path it takes on any given run is not fully knowable ahead of time, and reproducing it exactly on demand is harder than reproducing a fixed sequence.

The real axis underneath the label

Specialized versus generalized describes the shape of the agent. It does not describe the thing that actually matters to a finance or operations leader deciding whether to trust one of these near a ledger. That deeper axis is deterministic versus probabilistic.

A deterministic process, whether run by a person, a script, or a narrow agent, produces the same output from the same input every time. A three way match either passes or it does not, based on rules that do not change between Tuesday and Thursday. A probabilistic process, which is what a large language model doing open ended reasoning actually is underneath the interface, produces an output that is likely to be correct given the input, not guaranteed to be correct. Run it twice and you may get two answers that are both individually defensible and are not identical.

Specialized agents tend to sit closer to the deterministic end, because their designers constrained the decision space in advance and left little room for the model to improvise. Generalized agents tend to sit closer to the probabilistic end, because improvisation across an open ended tool is the entire value proposition. But the two axes are not the same axis wearing different names, and conflating them is where a lot of the current enterprise anxiety about agents actually comes from. A specialized agent can still make a probabilistic judgment call inside one of its fixed steps, such as classifying a transaction description as likely fraudulent. A generalized agent can still land on a fully deterministic sub task, such as pulling a specific ledger balance, where there is only one correct answer and no room to improvise.

Why D365 F&O is the place this gets tested for real

Dynamics 365 Finance and Operations is a useful proving ground precisely because it already lives with both philosophies side by side. Its native controls, workflow approvals, posting rules, three way match, are deterministic by construction. They were built that way before anyone was talking about agents, because a ledger that produces a different answer on a re-run is not a ledger anyone can close a period against.

Layer Copilot style features and Cowork on top of that and you get a system where the system of record stays deterministic while the layer working on top of it is free to reason probabilistically about what needs attention. That is not a contradiction. It is the correct division of labor. The ledger should never guess. The agent looking for the discrepancy that the ledger’s own rules were not written to catch is allowed to guess, because guessing well is precisely what it is there to do.

The mistake worth watching for is treating a generalized, probabilistic agent as if it were a specialized, deterministic one because it happens to be plugged into the same system. Giving Cowork the same blind trust you would give a fixed three way match rule is a category error. One was built to never deviate. The other was built to deviate on purpose, whenever deviating produces a better answer.

Where this actually lands

Neither shape is the correct answer in general. A specialized agent is the right tool when the process is well understood, the stakes of an unreviewed error are high, and an auditor will eventually ask for the sequence of steps. A generalized agent is the right tool when the problem is not well understood in advance, when the value comes from adapting to a situation nobody scripted for, and when a human is still positioned to review the output before it becomes a transaction that cannot be undone.

The organizations getting this right are not the ones picking a side. They are the ones that have stopped asking whether an agent is good, and started asking where on both axes, specialized to generalized and deterministic to probabilistic, a given task actually belongs. Get that placement wrong in either direction and you either waste a general reasoning tool on a task a five line business rule already solved, or you hand a fixed script a problem it was never built to handle and call the resulting failure a bug instead of what it actually is, a design mismatch.

An Advanced Dungeons & Dynamics 365 session. Episode 14 of 14. Season finale.

Previously on: Episode 13, The Audit Trail, the party disclosed the tax overcorrection to Feld immediately rather than waiting, and Feld pulled in audit and legal, then asked the party what go live actually needed from him personally.

Fourteen episodes. A shadow ledger, a silent integration, a rebate accrual eighteen months out of date, a chart of accounts stitched together by three different consultants, a demo that was finally allowed to break, and a tax override that outlived the bug it was built to fix. Every one of those problems is either resolved or being actively managed in the open. This is the weekend all of it gets tested at once.


The Session

GM: Cutover weekend. Every system down, every fix live, every character sheet about to be tested for real.

Marge: Status check. Thorne?

Thorne: Ledger’s clean. Rebate automation posted its first real accrual an hour ago. Correct, to the cent.

Sable: Chart of accounts holds. TEMP_DO_NOT_USE is finally retired, we mapped every report around it first.

Vex: Tax override removed. Legal signed off on the disclosure. Integration retry logic held through a full order cycle.

Ai. Cassiopeia: Power user proficiency scores, current state: seven of eight above target. Denise is mentoring the eighth.

Marge: Feld?

Feld: Board’s aware. Audit’s engaged. For the first time in this whole project, there’s nothing left that we’re hiding from ourselves.

GM: Monday morning. The first live order comes through. Every person who ever touched this engagement is watching the same screen.

Thorne: It’s processing.

Sable: It’s processing clean.

Marge: Third time’s the go live.

The order posts. No exceptions. No shadow ledger. No suspense account. Somewhere, Priya, back from a very long vacation, gets a message with just one word in it: “Live.”

Vex: (watching the screen) Fourteen episodes for one clean transaction.

Sable: It’s never really been about the one transaction.

Ai. Cassiopeia: Correct. It is about the fact that the next one, and the one after that, will hold the same way, for reasons everyone in this room can now actually explain.

Thorne: (quietly) That’s the part I’ll remember. Not that it worked. That we know why it worked.

Marge: (to the room) Good work, everyone. Truly.

Feld: (standing at the back of the room, having watched the whole thing) I’ve sat through two go lives that felt like holding your breath and hoping. This one felt like watching something that was actually built to hold.

Oskar: (from near the door, smiling for the first time in fourteen episodes) It’s a good feeling.


What actually changed between attempt three and the first two

Nothing about Contoso’s underlying complexity got simpler over the course of this campaign. The chart of accounts is still layered with history. The rebate program is still complicated. International tax rules are still international tax rules. What changed wasn’t the difficulty of the problem. What changed was whether the people closest to each problem felt safe enough, and equipped enough, to name it clearly the moment they found it.

Every failure in this campaign, the shadow ledger, the silent integration, the eighteen month manual accrual, the two year old tax override, the training program measuring attendance instead of competence, shares the same shape underneath the technical details. Something wasn’t working, someone noticed, and the information didn’t travel to whoever needed it, either because there was no channel for it, or because the channel that existed had already taught people that raising it wasn’t worth the trouble. Priya’s spreadsheet, Oskar’s silence, Denise’s unheard flag in month two, Feld’s own two years of unspoken knowledge, all of it is the same failure wearing different clothes.

What actually got fixed on this go live wasn’t the ledger. It was that. A demo that was allowed to break. A sponsor who learned that honesty wouldn’t blow up the room. A frontline analyst who finally got asked a follow up question instead of being scored and moved past. None of that shows up on a go live checklist, and all of it is the actual reason the third attempt held when the first two didn’t.


The pattern isn’t over

Every dungeon in this campaign, the shadow ledger, the forgetting curve, the golden path demo, the burden rate blind spot, is a real failure mode from real ERP engagements, just given a party and some dice. If any of these fourteen episodes made you wince in recognition, that was always the point.

  • Read the framework this whole campaign is built on: Gamifying the Enterprise: Game Mechanics for Continuous Proficiency, available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX
  • Get the full AD&D365 configuration guides, the Bare Bones Configuration Guides, Fifth Edition, all seven volumes, and start turning your own team’s onboarding into something closer to episode eleven than episode one: adnd365.com/start

Thanks for following the Contoso Convergence from the docks of Waterdeep all the way to a clean Monday morning transaction. If there’s an appetite for it, there’s more than one way this world could keep going.

An Advanced Dungeons & Dynamics 365 session. Episode 13 of 14.

Previously on: Episode 12, The Customization Nobody Remembers, Vex found a two year old tax override still running on top of a platform fix that made it unnecessary eighteen months ago, which meant international tax had likely been calculated twice for a very long time.

Two weeks to go live, and the party finally has to decide what to do with a discovery that’s bigger than a bug fix. This episode is about the moment a technical finding becomes an ethical one.


The Session

Vex: The override was patching a tax bug. The bug was fixed in a platform update eighteen months ago. Nobody removed the patch.

Sable: So for eighteen months, international tax has been calculated twice. Once correctly by the platform, once again by the leftover override.

Thorne: Every international order in that window needs review.

Ai. Cassiopeia: I can generate the affected transaction list. It is not small.

Marge: Feld needs to know today. Not after go live.

The party requests an unscheduled meeting. Feld reads the summary standing up, without sitting down first, which everyone in the room understands as a bad sign before he says a word.

Feld: (reading the list) This goes to audit and legal before it goes anywhere else.

Thorne: That’s the right call. It’s not a great one. But it’s the right one.

Feld: (setting the list down carefully) Two go live attempts failed because people were afraid to bring me news like this. You just did it anyway.

Marge: That’s the job.

Feld: I know it is. I’m still noticing that it happened.

Vex: There’s a smaller silver lining, if it helps. The overcorrection was consistent and directional. It should make the reconciliation methodical rather than exploratory. We know exactly what we’re looking for.

Ai. Cassiopeia: I concur. This is not a needle in a haystack. It is a known pattern applied against a known window.

Feld: (a short, tired exhale that’s almost a laugh) After everything else this project has thrown at me, “at least it’s a known pattern” is genuinely reassuring. That says something about the last four months.

Sable: It says the bar moved. Not that this is small.

Feld: Understood. Bring in whoever you need. What does the timeline look like with audit and legal both in the loop?

Feld authorizes disclosure, pulls in audit, and, for the first time in the engagement, asks the party what go live actually needs from him personally.


The moment disclosure stops being optional

Every episode up to this one has been about finding a problem and deciding how to fix it. This one is different, because the discovery crosses a line where “fix it quietly and move on” was never actually on the table. A tax miscalculation across a known window of international orders isn’t a configuration issue the team gets to resolve on its own judgment. It has legal and regulatory weight that belongs to people outside the war room, and pretending otherwise, even with good intentions, would have been a much bigger risk than the original bug.

What’s worth noticing is how unremarkable the decision to disclose actually was, once it got made. Marge’s line, “Feld needs to know today, not after go live,” isn’t dramatic. It’s the obvious, boring, correct call, and it’s only notable at all because of everything the party has learned about this project over twelve episodes: two prior go live attempts where the vague language of “data quality issues” let real problems go unnamed. Feld himself sitting through disclosures he could have made and didn’t. The default on this project, before this episode, was to let uncomfortable news dissolve into something softer before it reached the person who needed to hear it clearly.

Feld’s response is the real payoff of everything episode seven and eight built. He doesn’t get defensive. He doesn’t ask if it can wait. He reads the finding standing up, makes the harder call immediately, and then does something a lot of sponsors never manage: he notices out loud that the team told him the truth without being forced to, and says so, rather than treating honesty as the baseline he was owed all along.


What happens next

Audit and legal are engaged, the timeline is adjusted to account for it, and the party heads into the final stretch with every major fire from the last twelve episodes either resolved or actively being managed in the open. Cutover weekend arrives, and every fix built over the course of this campaign gets tested at once, live, in front of the whole company.

Next episode: Episode 14, Go Live Weekend, the campaign finale (coming soon)


If your team has ever found something that needed to go to legal or audit before it went anywhere else, you know exactly how heavy that meeting feels. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a framework built around surfacing findings like this one clearly and early, instead of letting them soften into something vaguer, start here: adnd365.com/start