AI assistants are not a new idea. They are the first real attempt to fix a problem management research documented seventy years ago and never actually solved for most people.

Every pitch for an AI assistant makes some version of the same claim: it will take the administrative weight off your day so you can focus on the work that actually matters. That claim is usually presented as new. It is not. It is the exact same claim that got made about executive secretaries, and it is backed by the same research, run three separate times across seven decades, always finding the same number.

The short version: a large and remarkably stable share of professional work, something close to 40 percent, is low-judgment administrative overhead rather than the work someone was actually hired to do. For most of the twentieth century, the only fix on offer was a personal secretary, and that fix was rationed almost entirely by seniority. AI assistants are the first attempt to offer that same relief to everyone else. Here is where that number comes from, and why the history matters for judging whether the new fix is real.

The first study

In 1951, a Swedish economist named Sune Carlson did something nobody had done before. He asked a group of managing directors to keep detailed diaries of their working days, logged in real time rather than reconstructed from memory afterward. The result, published as Executive Behaviour, is generally regarded as the first systematic empirical study of what managers actually do with their time, as opposed to what people assumed they did.

The picture that emerged was not flattering. Executive days turned out to be fragmented, reactive, and dominated by short bursts of low-level activity. Phone calls. Correspondence. Brief conversations. Interruptions. The romantic image of the executive locked away making weighty strategic decisions bore little resemblance to the diaries. Most of the day was consumed by administrative traffic that did not require an executive’s judgment at all, it just required someone competent to handle it.

Carlson’s book was not widely read at the time. It was sparse on conclusions and thin on theory. But the method survived. Two decades later, Henry Mintzberg ran a similar observational study and arrived at the same finding, dressed in more modern language. Managers do not spend their days thinking. They spend their days responding.

The finding keeps getting rediscovered

What is striking is not that this was found once. It is that it has been found repeatedly, in different decades, using different methods, on different populations of workers, and the number keeps coming back roughly the same.

In 2013, Julian Birkinshaw and Jordan Cohen ran a three year study of knowledge workers and published the results in Harvard Business Review under the title “Make Time for the Work That Matters.” Their headline finding: knowledge workers spend an average of 41 percent of their time on discretionary activities that offer little personal satisfaction and could be handled competently by someone else. Participants who went through a structured process to identify and shed that work cut desk work by roughly six hours a week and meeting time by two more.

Five years later, Harvard Business School’s Michael Porter and Nitin Nohria published the results of tracking 27 Fortune 500 CEOs for a full year, more than 60,000 hours of coded time-use data. It remains the most detailed public look at how chief executives actually spend their days. The finding was not that these people were lazy or undisciplined. It was that even at the very top of an organization, with every resource available to protect their time, a large share of it still got eaten by things that did not need a CEO to do them.

Three studies, seven decades apart, using diaries, surveys, and direct observation, converging on the same basic fact: a large and remarkably stable percentage of professional work is low-judgment administrative overhead. Not incompetence. Not poor discipline. Just the physics of running an organization, where information has to move, calendars have to align, and someone has to read the email before anyone can decide whether it matters.

Why the fix was always rationed

Here is the part of the story that gets skipped over. The research kept finding the same problem, but the solution it kept prescribing, dedicated administrative support, was never distributed according to who had the problem. It was distributed according to seniority.

A junior employee drowning in the same 41 percent of low-value work as a CEO simply absorbed it. There was no equivalent relief further down the org chart. And as email, shared calendars, and self-service scheduling tools spread through the 1990s and 2000s, even the people who had once had dedicated secretarial support increasingly lost it, on the theory that the tools themselves had closed the gap. The secretary and executive assistant roles contracted sharply. The underlying problem the research had documented did not contract with them. It just became something everyone quietly managed alone.

This is worth sitting with, because it means the historical relationship between “having administrative support” and “being more productive” was never really tested at scale. It was tested at the top of organizations, on people who already had every other advantage, and the results were treated as proof of a general principle that was never actually given the chance to apply generally.

What changes when the support is not rationed by rank

The research gives a fairly precise description of what kind of work is worth taking off someone’s plate: correspondence, scheduling, screening, routine drafting, information retrieval, status tracking. Not judgment work. Not relationship work. The mechanical layer underneath both of those things.

That is also, not coincidentally, close to the exact list of tasks that current AI tools are best at and are being adopted for first. Which suggests the honest way to think about AI assistants is not as a replacement for a human secretary, doing the same job with different hardware. It is closer to the first real attempt to deliver the same category of relief that Carlson’s executives had, to people who were never senior enough to be given it.

That reframe matters because it changes what the interesting question is. The old question, does an executive with a secretary outperform one without, was really a proxy for a different question that took seventy years to ask properly: how much of anyone’s professional capacity is being spent on work that has nothing to do with why they were hired, and what happens when that overhead gets pushed down toward zero for everyone, not just the people at the top of the org chart.

There is a genuine limit to the analogy, and it is worth naming rather than skating past. A human assistant carries judgment and institutional memory that took Carlson’s own subjects years to build with the people supporting them, knowing which call to interrupt a meeting for and which one to let go to voicemail. Whether that layer of judgment gets replicated, approximated, or simply left undone is not a question the old research can answer, because the old research was never testing for it. What it can tell us, with unusual consistency across seventy years of data, is the size and shape of the problem being solved. That part, at least, is no longer a mystery.

An Advanced Dungeons & Dynamics 365 session. Episode 5 of 14.

Previously on: Episode 4, Page Seventeen, the party found the missing integration from the missing SOW page, an undocumented patch quietly feeding the shadow ledger, failing silently for three weeks with nobody watching.

Every implementation eventually produces a moment like this one: a real decision, a hard deadline, and not enough information to make the decision comfortably. This episode is that moment.


The Session

CFO: One hour. Fix it or freeze it. I have a board call.

Marge: Vex, can you patch the retry logic in an hour?

Vex: I can patch it. I can’t test it in an hour. Not against three weeks of backlog.

Thorne: Then we freeze it and post the backlog manually. Clean, auditable, slow.

Sable: Slow is fine. Wrong is not fine. I vote freeze.

Marge: I vote fix. If we freeze it now, it becomes tribal knowledge that “the integration is off” and nobody turns it back on. I’ve seen that before.

Thorne: Two and two.

Ai. Cassiopeia: I was not asked to vote. I will offer the information anyway. The backlog is not three weeks of data. It is three weeks of data plus a duplicate batch from a retry storm on day one. The real backlog is roughly half what Vex estimated.

Vex: …that changes the math. I can test that in an hour.

Marge: Then we fix it. Vex, go. Sable, you’re validating every posting before it touches the real ledger.

The room splits into motion. Vex pulls up the retry logic, hands moving fast. Sable stands over his shoulder, watching every line scroll past like she’s guarding a gate. Thorne quietly starts drafting a manual reconciliation plan anyway, just in case the fix doesn’t hold.

Thorne: (not looking up) Belt and suspenders.

Marge: Good instinct. Keep going.

Twenty minutes in, Vex hits something.

Vex: Found the actual bug. It’s not the retry logic itself. It’s a timeout value that’s too short for the freight audit feed on high volume days. The retries were never the problem. They were a symptom of a timeout nobody ever tuned.

Sable: So the real fix is smaller than we thought.

Vex: Much smaller. One config value.

Ai. Cassiopeia: I recommend documenting this timeout setting explicitly going forward. It has apparently been a default value since implementation, never revisited.

Forty one minutes later, Vex’s hands are still on the keyboard when the CFO walks back in early.

CFO: Well?


Why Cassiopeia’s vote mattered more than a vote

The most interesting moment in this scene isn’t the fix. It’s the tie. Two votes for freeze, two for fix, and the deciding information came from the one party member who doesn’t get a formal vote at all. That’s not a coincidence, and it’s not there to make a point about AI having opinions. It’s there because Ai. Cassiopeia had access to something the humans in the room didn’t have time to check under pressure: the actual shape of the backlog, not the estimated shape of it.

Under a real deadline, teams default to their gut. Freeze feels safer because it’s reversible. Fix feels riskier because it isn’t. Both instincts are reasonable, and both were wrong in the same way, because both were built on an estimate nobody had time to verify. The team wasn’t voting on values. They were voting on incomplete information, and neither side knew it.

This is the actual case for a grey collar workforce member in a room like this one. Not to replace judgment, Marge and Thorne still made the call, but to remove the guesswork underneath the judgment before the decision gets made instead of after. And notice what Vex found once the pressure came off enough to actually look: the retry logic everyone assumed was broken wasn’t the real problem at all. It was a five minute config fix hiding behind three weeks of assumed complexity.

That’s usually how it goes. The scary problem and the real problem are rarely the same problem.


What happens next

The fix holds, and landed cost starts posting clean for the first time in weeks. But validating every posting before it touches the real ledger means someone has to look closely at everything already in the system, and that closer look surfaces something nobody in the room was looking for: a second, older shadow process that’s been quietly running in the background of Contoso’s books for a lot longer than three weeks.

Next episode: Episode 6, The Ghost in AP (coming soon)


If your team has ever voted on a fix under a deadline without actually knowing the real size of the problem, you’re in good company. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a framework for surfacing the real problem before the deadline forces a guess, start here: adnd365.com/start

Cause of death: three years of production data quietly became four hundred clean sample records for the day of the pitch.


Somewhere in the middle of the demo, the presenter opens a grid. Customers, items, transactions, whatever the domain calls for. It scrolls smoothly. Every row has every field populated. Names are properly capitalized. Addresses have all their parts. There are no duplicate customer records for “Acme Corp,” “ACME Corp,” and “Acme Corp.” with a trailing space that your actual system has accumulated over a decade of different people typing the same name slightly differently.

The presenter doesn’t say “this is sample data.” They don’t need to. The grid looks so much like a real company’s data that the distinction quietly stops mattering to the room, and everyone leaves the meeting having watched a migration that never happened, of data that was never really yours.

What actually happened

Four hundred rows of clean, plausible-looking data is not a migration. It’s a mockup wearing a migration’s clothes. Somebody built that dataset specifically to demonstrate the target system’s data model, which means it was constructed backward from what the target system wants to receive, rather than forward from what your source system actually contains. It has never been through a real extract. It has never hit a field length limit, a character encoding mismatch, a required field that’s been null in your source system since 2019 because nobody enforced it, or a foreign key that points to a parent record that got deleted three reorganizations ago.

Real migrations die on exactly these details, and none of them are visible in a four hundred row demo grid, because the demo grid was never subjected to the process that would surface them. The sample data is a hypothesis about what your data looks like. It has not yet met your data.

There’s a second layer under this. Even when a vendor does an actual proof of concept against a real extract of your data, that extract is usually a snapshot, cleaned once, run through a mapping exercise once, and shown once. It demonstrates that a migration is possible for that slice, on that day, with that much attention paid to it. It does not demonstrate that the full historical dataset, run through the same process without the benefit of a team hand-tuning exceptions in real time, will produce the same result.

Why it works on smart people

Data problems are boring in a way that makes them easy to underestimate from the outside. Nobody gets excited describing thirty thousand customer records with inconsistent capitalization, or a decade of transactions where the currency field was optional for the first four years, and that lack of drama works against the diligence the problem deserves. A demo that skips the data reality skips the part of the story that was never going to be compelling to watch anyway, and audiences let it go for the same reason they let the seamlessness of the Golden Path Demo go: friction is a strange thing to ask someone to add back in.

There’s also a scale-blindness effect. A grid of four hundred rows and a database of four million rows look identical in a screen share, because you’re only ever looking at the same twenty rows on screen at once. The demo cannot visually communicate that the four hundred clean rows are a curated sliver, not a representative sample, so the brain does what it usually does with limited visual information: it extrapolates, and assumes the part it can see generalizes to the whole.

The actual damage

This is the one that blows up the project timeline more reliably than almost anything else, and it does it quietly, in the data cleansing and reconciliation phase that was budgeted as a two-week task because the demo made data migration look like a solved problem. Then someone runs the real extract, and it turns out eight percent of vendor records have no valid tax ID, eleven percent of item records reference a unit of measure that was deprecated four years ago, and there are nineteen thousand duplicate customer records that need to be identified and merged before go-live, none of which showed up in four hundred rows of hand-picked sample data.

The two-week task becomes a two-month task, the go-live date moves, and the business case that assumed a smooth data conversion now has to absorb a delay that nobody priced in, because the thing that actually determines a migration’s difficulty, the messiness of the real data, was the one thing the demo was specifically built not to show.

The fix, if you’re the one presenting

Run the demo against a real, ugly extract, even a small one, and don’t clean it first. Show the duplicate detection running against actual duplicates. Show what happens when a required field is null. Show the exception queue, and how many records land in it, and what the resolution workflow actually looks like for the person who has to work through that queue by hand. It’s a less polished five minutes. It’s also the only five minutes that tells the prospect anything real about what their conversion will cost.

A migration demo that never encounters bad data hasn’t demonstrated a migration. It’s demonstrated the destination.


The mapping exercise itself looks tidy in a lab too. In Familiar Ground: Mapping CRM to ERP, every concept has a clean twin on the other side, names changed, forms wider, same underlying logic. Real source data is rarely that cooperative, which is exactly the gap this autopsy is about.

An Advanced Dungeons & Dynamics 365 session. Episode 4 of 14.

Previously on: Episode 3, The Labyrinth of Financial Dimensions, the party mapped a chart of accounts that turned out to be three different structures stapled together, and found a dimension called TEMP_DO_NOT_USE quietly carrying eleven thousand postings nobody remembered approving.

Every SOW has an appendix nobody reads until they have to. For Contoso, that appendix was page seventeen, the integration list, and it went missing before the party ever set foot in the building. This episode, they finally find it.


The Session

GM: Vex finds page seventeen. It was stapled into the wrong appendix.

Vex: Six integrations. EDI for two vendors, a freight audit feed, a payroll export, a legacy AR aging tool, and… this one.

Thorne: Which one?

Vex: No name. No owner. Just an endpoint.

Ai. Cassiopeia: I have log access. It has been failing silently since three weeks ago. Retry logic swallows the error.

Sable: Failing into what?

Vex: (long pause) It posts landed cost adjustments. Into the shadow ledger.

Marge: So the spreadsheet Priya built wasn’t tracking a gap in the system. The system was quietly feeding it.

Thorne: Then the six figure variance from episode two…

Ai. Cassiopeia: Is not a variance. It is three weeks of unposted freight true ups sitting in a dead queue.

The room goes quiet. Somewhere down the hall, a fax machine that shouldn’t still exist starts printing.

Vex: (staring at the endpoint) This integration doesn’t have a name because whoever built it never expected anyone to find it. Look at the naming convention. It doesn’t match anything else in the environment.

Sable: It’s not part of the original build.

Vex: No. It’s a patch. Somebody built this after the fact, quietly, to route around a problem they didn’t have time to fix properly. And then never told anyone it existed.

Marge: Which means it was never in scope for testing. Never in scope for documentation. Never in scope for anything.

Ai. Cassiopeia: Correct. It has been running in the dark since before this engagement began.

Thorne: (quietly) We didn’t inherit a shadow ledger problem. We inherited a shadow integration that created one.


The workaround nobody admits to building

Every party in this scenario has been telling the truth as they understood it. The CFO didn’t know the integration existed. Priya didn’t know her shadow ledger was being fed by something outside her control. Even the original consultants probably didn’t know, since this integration was never part of the documented build. Somebody, at some point, hit a wall, needed the numbers to move, and built a bridge nobody else was told about.

This is the quiet failure mode that doesn’t show up in a status report. Not a missed deadline. Not a scope change. A small, undocumented patch, built with good intentions under real pressure, that outlives the person who built it and the reason they built it. Nobody decided to create three weeks of silent failure. Somebody just wrote retry logic that swallowed errors instead of surfacing them, because surfacing them would have meant admitting the patch existed in the first place.

The lesson here isn’t “don’t build workarounds.” Sometimes a workaround is the only thing standing between a business and a missed close. The lesson is that a workaround without an owner, without documentation, and without visibility into its own failures isn’t a bridge. It’s a debt that compounds silently until someone like Vex has to go digging for an endpoint with no name.


What happens next

The party now knows exactly where the six figure variance came from, and it isn’t a variance at all, just three weeks of backlog stuck behind a silent failure. But knowing the cause and fixing it in time are two different problems. The CFO wants an answer before his board call, and he’s giving the party one hour to decide whether to patch the integration live or freeze it and clean up the backlog by hand.

Next episode: Episode 5, One Hour (coming soon)


If there’s an integration running quietly in your environment that nobody currently owns, you’re not the only one. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want the configuration discipline that keeps workarounds like this one from going dark in the first place, start here: adnd365.com/start

Cause of death: two systems that had never spoken before were shown having a conversation written for them.


The presenter switches windows. A record gets created in System A. A few seconds later, as if by magic, the corresponding record appears in System B, fully formed, correctly mapped, no errors. “And that’s it,” the presenter says, “they just talk to each other.” The room relaxes. Integration, historically the single most reliable way for an implementation to go over budget and past deadline, has apparently been solved by two systems having a friendly chat while everyone watched.

Nobody in the room asks what was actually watching that conversation, or who taught it what to say. That’s the autopsy. The two systems didn’t learn to talk to each other. Someone wrote both sides of the script, tested it exactly once, against exactly one scenario, and ran it live in front of you.

What actually happened

An integration demo almost never shows the integration. It shows the happy path of the integration, which is a different and much smaller thing. The record that got created in System A was built to contain precisely the fields System B expects, in precisely the format System B expects them, with no null values in the fields that would trigger a mapping error, no duplicate keys, no encoding mismatch, none of the thousand small inconsistencies that live in a company’s actual data the moment more than one person or one legacy system has touched it.

The script connecting the two systems, whether it’s a middleware platform, a custom connector, or a scheduled job, was very likely written specifically for this demo, tuned against this one scenario, and has never been asked to handle a partial failure, a duplicate record, a field that arrives populated in one system and empty in the other, or a timeout on either end. It works, in the same sense that a bridge works if you only ever drive one specific car across it at one specific speed.

The part that never gets demoed, because it can’t be demoed in three minutes, is everything that happens when the sync fails halfway through. Does the transaction roll back cleanly on both sides, or does System A now believe the record synced while System B never received it. Is there a retry, and if so, does the retry create a duplicate. Is there an alert, and does it go to a person who is actually watching for it, or does it silently populate an error log nobody has looked at since the demo environment was built.

Why it works on smart people

Integration failures are, structurally, invisible until they aren’t. A sync that fails silently doesn’t announce itself. It just produces a slowly widening gap between what System A believes is true and what System B believes is true, and that gap is usually discovered by someone downstream, weeks or months later, reconciling numbers that don’t match and trying to figure out why.

Because the failure mode is invisible, “the demo showed it working” carries more weight than it should, simply because there’s no immediately visible counter-evidence in the room. A broken UI is obvious the moment you see it. A broken integration is obvious only in the reconciliation report nobody runs until month-end close, by which point the demo is a distant memory and the sales team has moved on to the next opportunity.

There’s also a vocabulary problem working in the vendor’s favor. “They just talk to each other” is a satisfying sentence, and it papers over an enormous amount of engineering that either exists, robustly, behind that sentence, or doesn’t exist yet and was built specifically to survive one scripted run.

The actual damage

This is the one that shows up as a reconciliation nightmare rather than a single dramatic failure. Two systems that were sold as integrated drift slowly apart in the weeks after go-live, each one silently correct according to its own records, disagreeing with the other in ways nobody notices until an audit, a customer complaint, or a finance close turns up numbers that don’t tie out. By then the question isn’t “does the integration work,” it’s “how long has it not been working, and what decisions got made on bad data in the meantime.”

The remediation is almost always more expensive than building the integration correctly the first time would have been, because now it includes both the engineering fix and a data cleanup project to reconcile however many weeks or months of silent drift accumulated before anyone noticed.

The fix, if you’re the one presenting

Show a failure on purpose. Send a record with a missing required field, or a duplicate key, and show what happens: does it error visibly, does it queue for retry, does someone get notified, does the other system stay in a known, correct state while the problem gets resolved. If the honest answer is “we haven’t built that handling yet,” say that, and say what the plan is. A prospect who sees a deliberate, controlled failure and a sane recovery path trusts the integration more than one who only ever saw the happy path, because they now know what happens on the day, and there will be a day, when the happy path isn’t what shows up.

Two systems that have never disagreed in front of you haven’t been integrated. They’ve been introduced.


This is exactly the failure mode a ledger-first architecture is built to make impossible. In Is Headless ERP Enough, or Just a Step in the Right Direction?, I walk through a prototype where two disconnected nodes post independent transactions and converge without conflicts, with no consensus protocol and no room for one system to quietly believe something the other doesn’t.

An Advanced Dungeons & Dynamics 365 session. Episode 3 of 14.

Previously on: Episode 2, The Shadow Ledger, the party discovered that Priya had been running a shadow spreadsheet for fourteen months, revealing a six figure variance nobody could explain. Priya left for vacation before anyone got an answer.

Every ERP has a chart of accounts. Not every ERP has a chart of accounts that actually makes sense. Sometimes what looks like a financial structure is really three different structures, built at three different points in the project’s history, quietly stapled together and never reconciled.

That’s the dungeon the party walks into this episode.


The Session

GM: You descend into the chart of accounts. It is, structurally, three different charts of accounts stapled together.

Sable: There are four segments for cost center. Four. Two of them mean the same thing in different departments.

Vex: (picking the lock on a legacy dimension table) Found something. A dimension called TEMP_DO_NOT_USE.

Thorne: How many postings?

Vex: Eleven thousand, two hundred and six.

Marge: Whose decision was that?

Vex: Nobody’s. That’s the problem. Somebody created it during UAT for testing and it just never got decommissioned. It’s load bearing now.

Ai. Cassiopeia: I flag that removing it will break forty three saved reports, none of which are documented.

Sable: We’re not removing it today. We’re mapping around it. Nobody touch the dimension.

A junior analyst walks past the war room door, glances in at the whiteboard now covered in a hand drawn diagram of overlapping cost center segments, and quietly closes the door again without saying anything. Nobody in the party blames him.

Thorne: (studying the diagram) This isn’t a chart of accounts. It’s an archaeology site.

Sable: Every layer’s a different consultant’s fingerprints. I can tell you which segment was added during which go live attempt just by how badly it fits with the one next to it.

Vex: (still elbow deep in the dimension table) There’s a comment field on this one. Someone left a note.

Marge: What’s it say?

Vex: “Will fix in phase two.”

Ai. Cassiopeia: I checked the project timeline. There was no phase two.


Why nobody decommissions the temp table

Every implementation accumulates a TEMP_DO_NOT_USE somewhere. It starts as exactly what the name says, a placeholder for testing, something built under deadline pressure that was always going to get cleaned up later. Then UAT ends, go live happens, and “later” quietly becomes never, because by the time anyone has bandwidth to revisit it, eleven thousand transactions are already posted against it and forty three reports are already built on top of it.

This is how technical debt actually accumulates in an ERP. Not through one bad decision, but through a hundred reasonable ones that never got revisited. Nobody sat down and decided Contoso should have four overlapping cost center segments. Somebody added one during the original build, somebody else added a second during a scope change nobody documented, and a third arrived during the failed second go live attempt when a different consultant didn’t know the first two existed.

The dangerous part isn’t that the structure is messy. Messy is survivable. The dangerous part is that nobody currently alive on the project understands the whole thing well enough to safely change it. That’s not a data model problem anymore. That’s an institutional memory problem, and it’s the same failure mode as the shadow ledger from episode two, just wearing a different disguise.

Sable’s instinct here is the right one: map around it before you touch it. You don’t fix an archaeology site by bulldozing it. You document what’s actually there first, then decide what’s safe to change.


What happens next

The party has a working map of the chart of accounts and a firm rule not to touch TEMP_DO_NOT_USE. But mapping the dimension structure surfaces the missing piece from episode one: the integration list from page seventeen of the SOW. In the next episode, the party finally gets their hands on it, and finds one integration on the list with no name, no owner, and no documentation at all.

Next episode: Episode 4, Page Seventeen (coming soon)


If your own chart of accounts has a TEMP_DO_NOT_USE hiding in it somewhere, you’re not alone, and you’re not imagining the risk. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a real framework for documenting and untangling structures like this one before they become load bearing, start here: adnd365.com/start

Cause of death: the case study was true, and that’s exactly the problem.


Two-thirds of the way through the deck, a new logo appears. A real one, a company you’ve heard of, sometimes a competitor’s supplier or a name from your own industry vertical. The slide has a number on it, usually a big one: forty percent reduction in close time, three million recovered in duplicate payments, six months to positive ROI. Underneath the number is a quote, attributed, sometimes even video, from a real person who really said those words.

Nothing on that slide is fabricated. That’s what makes this one the hardest autopsy in the series. The other demos in this blog die from omission, pacing, or seamlessness hiding a seam. This one dies from something subtler: a true statement about one company, presented in a context engineered to make you believe it’s a claim about yours.

What actually happened

The reference customer on the slide is not a random sample. It is, almost by definition, the single best outcome the vendor has produced across their entire installed base, selected specifically because the number is large and the customer is willing to say it out loud. Somewhere behind that slide are dozens or hundreds of other implementations that landed closer to the median, plus a smaller number that struggled or stalled, none of which get a logo or a quote, because nobody puts “we got most of the way to the business case, eventually, after two scope changes” on a slide.

There’s also a matching problem the case study never surfaces. The reference customer’s forty percent reduction in close time happened inside a specific starting condition: a particular level of process maturity, a particular data quality baseline, a particular willingness internally to change how work got done. The case study tells you the outcome. It almost never tells you the starting line, and the outcome without the starting line is not a number you can subtract your own situation from.

The quote does real work here too. A specific named person saying a specific thing on camera reads as harder evidence than an aggregate statistic, even though a single testimonial is a sample size of one, hand-selected from a population the vendor controls entirely.

Why it works on smart people

Humans are wired to trust specific, named, social proof more than abstract statistics, and this isn’t a flaw, it’s usually a reasonable heuristic. A named person willing to put their reputation behind a claim on camera is, in most contexts, more credible than an anonymous number. The problem is that the heuristic evolved for a world where the sample in front of you was roughly representative of the population, and a vendor-selected reference customer is the opposite of representative by construction.

There’s a second effect working alongside the first. By the time the reference slide appears, you’ve usually already sat through thirty or forty minutes of a demo that felt competent, so the case study isn’t landing on a skeptical audience, it’s landing on an audience that has already been primed to trust what they’re being shown. The reference customer isn’t doing the persuading alone. It’s the closing argument after the room has already been warmed up.

The actual damage

This is the one that turns into an internal expectations problem before it turns into a vendor problem. Someone in the room, often not maliciously, repeats the number in an internal steering committee deck as though it were a forecast rather than someone else’s outcome. “Similar companies have seen a forty percent reduction” quietly becomes “we’re targeting a forty percent reduction,” and by the time the project charter gets written, a single best-case data point from a different company, with different starting conditions, has become your project’s success criteria.

When your actual results land closer to the median, which is where most results land by definition, the project doesn’t get judged against a realistic baseline. It gets judged against the reference customer’s outcome, which nobody on your team ever should have agreed to as the target in the first place.

The fix, if you’re the one presenting

Show the range, not just the peak. If you have a reference customer at forty percent, say what the twenty-fifth and seventy-fifth percentile outcomes look like too, and say why the reference customer landed where they did, what was true about their starting point that might or might not be true about the prospect’s. A specific, named case study is still worth showing. It’s worth showing better, with its context attached, instead of as a number floating free of the conditions that produced it.

The honest version of that slide is less dramatic. It’s also the only version that survives contact with a steering committee eighteen months later.


This is the same shift I wrote about in The New Expert Isn’t the One With the Answers. Having the number was never the hard part. Knowing whether that number applies to your situation is.

An Advanced Dungeons & Dynamics 365 session. Episode 2 of 14.

Previously on: Episode 1, The Summons at Waterdeep Docks, the party arrived at Contoso Coffee Roasterie to find a missing scope document, a go live date that had quietly moved up by six weeks, and a whiteboard asking a question nobody could answer: where is the margin going.

Every implementation has an official system of record. And every implementation, if it’s been running long enough, has an unofficial one too. Usually a spreadsheet. Usually built by someone who got tired of waiting for the real numbers to be right.

That’s where episode two picks up.


The Session

GM: In the war room you meet three Contoso stakeholders: the CFO, the head of Ops, and a woman named Priya who nobody introduces by title.

Sable: (casting Data Mapping) I’m getting margin data. It’s coming from… this isn’t the ERP. This is a workbook.

Priya: That’s mine. I’ve been tracking real margin since the system doesn’t roll up freight variance correctly.

Marge: How long has this shadow ledger existed?

Priya: Since the first go live attempt. Fourteen months.

Thorne: (rolling a Perception check) The workbook total and the GL total are six figures apart.

Vex: Which one’s right?

Priya: I don’t know anymore. I leave for vacation tomorrow.

Ai. Cassiopeia: I would like it noted that I offered to reconcile this in October. The ticket was closed as “won’t fix, not urgent.”

The CFO leans over Thorne’s shoulder to look at the variance number. He goes pale in a way that has nothing to do with the lighting.

CFO: That number is bigger than our reported profit.

Nobody says anything for a moment. Priya starts quietly packing her laptop bag, the way someone packs when they’ve decided this is no longer their problem to solve, whether or not that’s actually true.

Priya: (standing) My flight’s at six tomorrow morning. I really am sorry.


Why the spreadsheet always wins

Every consultant who has done this work long enough has met a Priya. Somebody who didn’t set out to build a shadow system, they just needed the numbers to be right for a meeting, and the ERP wasn’t giving them right numbers fast enough. So they built a workaround. Then the workaround became a habit. Then the habit became fourteen months of institutional memory that lives entirely on one person’s laptop, reconciled by hand, understood by nobody else in the building.

This isn’t a story about a bad employee cutting corners. It’s what happens when a system doesn’t earn trust fast enough, and a business still has to make decisions in the meantime. Priya’s spreadsheet isn’t the failure. It’s evidence of a failure that happened somewhere upstream, probably long before she ever opened Excel.

The dangerous part isn’t that the spreadsheet exists. It’s that nobody in the room can currently say which number, the workbook or the GL, is actually correct. Six figures of ambiguity, and the person who understands the discrepancy best is getting on a plane in twelve hours.

That’s not a data problem anymore. That’s a continuity problem. And it’s the kind of gap that a proficiency system is supposed to catch long before it turns into a war room moment.


What happens next

Priya is gone by morning, and the shadow ledger goes with her, at least in terms of anyone who can explain it. The party has six figures of unexplained variance, a CFO who now can’t unsee the number, and a chart of accounts they haven’t even looked at yet. In the next episode, they go looking for where the discrepancy actually lives, and find a financial dimension that’s been quietly absorbing thousands of postings that were never supposed to exist.

Next episode: Episode 3, The Labyrinth of Financial Dimensions (coming soon)


A quick gut check for anyone running their own ERP: is there a shadow spreadsheet keeping your business honest right now? Gamifying the Enterprise: Game Mechanics for Continuous Proficiency digs into exactly why that happens, and how to design proficiency systems so it stops. Available on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want the actual configuration playbook behind scenes like this one, start here: adnd365.com/start

Cause of death: the feature that closed the deal was never actually in the room.


Somewhere around minute forty of the demo, the presenter hits a gap. The thing you actually asked about, the reason you took the meeting, doesn’t quite exist yet. What happens next is the tell. The slide doesn’t say “we don’t do that.” It says “coming in the next release,” said in exactly the same tone of voice as everything that already works, with exactly the same confident click-through pacing, so that by the time the meeting ends, the feature that doesn’t exist has fully merged in your memory with the fifteen features that do.

Nobody lied. That’s what makes this one interesting to cut open. The roadmap slide was real. The quarter listed on it might even be accurate, as of the day the deck was built. And yet the effect on the room is functionally identical to a lie, because a promise wearing a product demo’s clothing gets evaluated with a product demo’s scrutiny, which is to say, almost none.

What actually happened

Every roadmap item in a sales deck starts life as an engineering estimate, gets filtered through a product manager’s optimism, gets filtered again through a sales engineer who needs this quarter’s number, and arrives in front of you as a single, confident bullet point that has shed every unit of uncertainty it was born with. “Q3” meant “Q3, if the two prerequisite features land on time and nothing gets reprioritized” back at the whiteboard where it was written. By the time it’s read aloud in your conference room, it just means Q3.

The demo compounds this by never distinguishing, in pacing or tone, between the click that shows something real and the click that shows a mockup of something planned. Both get the same enthusiasm. Both get the same “and here’s where you’d.” The interface doing the showing doesn’t have a font for “this is a Figma file with a database connection painted on.”

You are, in effect, being shown two different products stitched into one seamless walkthrough: the one that ships today, and the one that exists only as a commitment on a slide, and you’re being asked to make one buying decision that covers both.

Why it works on smart people

Buyers are trained, correctly, to evaluate a vendor’s direction and not just their current state. Nobody wants to buy a system that solves today’s problem and ignores next year’s. So a roadmap conversation is a legitimate, necessary part of due diligence. The trick isn’t the existence of the roadmap. It’s the demo borrowing the roadmap’s credibility and lending it back to itself.

There’s also a timing problem working against you. The roadmap feature is almost always introduced as the answer to the exact gap you just identified in the product, which means it lands at the precise moment you’re feeling a little disappointed and looking for a reason not to be. “Coming in Q3” isn’t just information at that point. It’s relief, and relief is a bad state to be evaluating claims in.

The actual damage

This is the one that shows up on a signed contract with a footnote nobody reads until it matters. Somewhere a business case got built with the roadmap item load-bearing in it, sized as though it were a current-state capability, because in the meeting it felt like one. The actual purchase decision, the one with budget and a signature attached, priced in a feature that was, at signing, a Jira ticket with a target quarter next to it.

Q3 arrives. The feature either doesn’t ship, ships in a reduced form that solves half the original problem, or ships correctly but a year later, after a reprioritization nobody outside the engineering org heard about. Your business case, however, was built on the version of the feature that existed only in the demo room, and now someone has to explain to their own leadership why the thing everyone signed off on isn’t the thing they got.

The vendor isn’t necessarily acting in bad faith here. Roadmaps genuinely slip, for genuinely defensible reasons. But “the vendor wasn’t lying” is cold comfort to the person holding a business case that assumed a delivery date as fact.

The fix, if you’re the one presenting

Change the font, literally or figuratively, the instant you cross from shipped to planned. A different slide background, a verbal flag, a pause, anything that makes the seam audible. Say the confidence level out loud: “this is committed and in QA,” versus “this is prioritized but not yet started,” versus “this is directionally where we’re headed and I wouldn’t bet a contract on the date.” Those are three different products. Let the buyer evaluate them as three different products.

It costs you a little bit of momentum in the room. It buys you a customer who signs with accurate expectations, which is the only kind of customer who’s still happy with you eighteen months later.

The roadmap wasn’t the lie. The seamlessness was.


I’ve argued elsewhere that no self-respecting architect leaves the scaffolding up once the building is done. A roadmap slide is the one place I’d argue for the opposite: leave the scaffolding very visible, since half of what’s on screen hasn’t been built yet.

Cause of death: nobody asked what the prompt actually did.


The setup is always the same. The presenter opens a chat panel next to the application. Types a sentence, something like “create a purchase order for our top vendor and route it for approval.” Hits enter. Three seconds later, a purchase order exists, fully populated, correctly routed, and the room exhales the specific kind of gasp that sales decks are built around.

Nobody in the room asks the only question that matters: what happened between the sentence and the purchase order.

That’s the autopsy. Not whether the AI worked. It worked. The patient died of the room’s collective decision not to look inside the box.

What actually happened

There are, broadly, three things that could have occurred in those three seconds, and they have wildly different implications for what you’re buying.

The prompt could have triggered a genuinely flexible model reasoning over your live data and your actual configuration, in which case, remarkable, and worth every follow-up question you can throw at it. Or it could have matched against a narrow, pre-built intent, one of a finite list the vendor trained and tested specifically for this demo, in which case the “AI” is doing roughly what a well-labeled button would do, dressed in a chat window. Or, and this is the quietly common one, it could have worked because the demo environment was built so there was exactly one top vendor, exactly one approval path, and exactly one plausible interpretation of the sentence, which means the model didn’t have to be smart. It had nothing to be confused by.

You cannot tell these three apart from the audience seat. That’s not an accident. It’s the format working as intended.

Why it works on smart people

Chat interfaces borrow credibility from every other chat interface you’ve ever used. You already know how to type a sentence and get a reasonable response, because you do it daily with a general-purpose assistant that genuinely is flexible. The demo imports that trust wholesale and applies it to a narrow, scripted intent recognizer that shares none of the underlying capability.

There’s a second layer, too. Asking “what exactly did the model do” mid-demo feels like a strange thing to interrupt with, technical in a way that makes you look like you’re missing the point of the show. So the room lets it go, same as it lets the Golden Path Demo’s pristine data go. The social cost of asking is higher than the informational value anyone expects to get from asking, so nobody does, and the vendor never has to answer a question nobody asked.

The actual damage

This is the one that costs real money, because “AI does it for you” gets sized into the business case. Someone in procurement multiplies the demo by the number of purchase orders your company cuts in a year and arrives at a headcount reduction that made it into a slide before a single production prompt had been run against a real vendor master with your actual data quality problems in it.

Then go-live happens, and the model that flawlessly handled “our top vendor” in the demo turns out to need a disambiguation step for the fourteen vendors named some variant of “Smith Industries” in your real vendor table, and the “flexible AI agent” turns out to be a narrow intent match against six pre-built scenarios, and scenario seven is where your actual business lives.

The gap between what was promised and what shipped doesn’t show up as a bug. It shows up as a quiet, expensive redefinition of what “handled by AI” meant, discovered by whoever now has to explain the missed headcount number.

The fix, if you’re the one presenting

Show your work, on purpose, before anyone has to ask. After the prompt returns its clean result, pull back the curtain for ten seconds: here’s the intent it matched, here’s the data it pulled from, here’s what happens if the vendor name is ambiguous. If the honest answer is “the model reasoned over live data,” say that, and let it be more impressive for being specific. If the honest answer is “this is a pre-built scenario for the categories we support today,” say that too. A buyer who knows exactly what they bought doesn’t come back angry in month four. A buyer who assumed general intelligence and got a narrow, well-executed lookup does, and they remember whose demo sold it to them.

The trick was never the model. It was the silence right after it worked.


If the “narrow intent versus genuine reasoning” distinction sounds familiar, it’s the same fault line I dug into in Is Headless ERP Enough, or Just a Step in the Right Direction?, on what it would actually take for AI to be the architecture instead of a coat of paint on top of it.