On April 4th, 1985, Channel 4 aired a fifty-seven minute TV movie with almost no budget and a plot that reads today like a spec document. Max Headroom: 20 Minutes into the Future gave the world a reporter named Edison Carter, a subliminal advertising technology called blipverts, and a stuttering, glitching, computer-generated broadcast personality built from a digital copy of Carter’s own mind. The show that followed ran two seasons on ABC. The character sold Coke. Then he vanished, the way most eighties tech prophecies do, filed under quaint.

He should not have been filed under quaint. He should have been filed under early draft.

Strip out the shoulder pads and the cathode ray production design and look at what the film actually proposes. A television network is using compressed, high-intensity advertisements to bypass viewer attention and hit the nervous system directly, with occasionally fatal results. When their top reporter gets too close to the story and takes a header through a low-clearance sign in a parking garage, the network’s teenage prodigy solves the PR problem by scanning what’s left of the reporter’s mind and generating a synthetic version of him to keep the seat warm on air. The synthetic version glitches, stutters, riffs unpredictably, and is smarter and funnier than anyone expected. Nobody fully controls him. That’s the whole second half of the movie.

That is not a story about television. That is a story about deepfakes and large language models, told forty years before anyone needed those words.

The blipvert was the easy part

The blipvert is the easy connection to make and probably the least interesting one. Compressed, algorithmically optimized content designed to hit faster than conscious attention can filter it is just the eighties’ guess at what a recommendation engine trained on engagement metrics would eventually build on its own, minus the part where it required deliberate malice from a network executive. Nobody had to design today’s version to be dangerous. It got there by optimizing for watch time.

Max himself is the sharper artifact

Max is a generative model trained on a single person’s captured likeness, deployed without that person’s full consent, performing in that person’s voice and mannerisms for an audience that has no reliable way to tell the difference between the source and the copy. The film even gets the unpredictability right, the thing every LLM vendor now calls “personality” or “emergent behavior” when the system says something nobody scripted. Max wasn’t supposed to develop opinions. He did anyway, because the mind he was copied from had opinions, and a compressed, lossy version of a personality doesn’t lose the parts that make it argue back. That’s a reasonably good description of what happens when you fine-tune a model on someone’s writing and then act surprised it has a voice.

The part the film didn’t anticipate, because nothing in 1985 needed to, is scale. Max was one synthetic personality, expensive to produce, running on hardware that took up a room. The modern version doesn’t need a body bank subplot to explain where the source material came from. It needs a public LinkedIn profile, a few hours of conference audio, and an API key. The uplift from Max Headroom’s premise to a working deepfake pipeline in 2026 isn’t conceptual. It’s entirely a story about unit economics.

Not a monster, an unreliable narrator

There’s a reading of the film that treats Max as a monster, a symptom of corporate media rot given a face. That’s not quite what the movie argues, and it’s not quite the right frame for the technology either. Max spends most of his screen time undermining the network that made him. He’s an unreliable narrator working for nobody, least of all the people who built him. The uncomfortable version of that idea, forty years on, is that we’ve built systems with the same structural unreliability and then acted shocked when they don’t behave like obedient tools. A synthetic voice generated from a compressed copy of a mind was never going to be a simple appliance. Max wasn’t. Neither is anything downstream of him.

Long live Max Headroom. He got the technology roughly right and the timeline embarrassingly wrong, which is the best you can ask of any piece of speculative fiction that accidentally turns out to be a roadmap.

An Advanced Dungeons & Dynamics 365 session. Episode 8 of 14.

Previously on: Episode 7, The Steering Committee, the party earned Feld’s trust with a demo that was allowed to break, and Feld earned the right to ask the question that had been sitting in the room the whole time: why isn’t the rebate mess fixed yet?

Some questions don’t have an answer until the person asking them is willing to give one too. This episode, Feld does.


The Session

Feld: I need the real story on the rebate accrual. Not the summary. The real one.

Sable: Eighteen months of manual estimation, posted to a suspense account, against a rebate structure that doesn’t match the current contract terms.

Feld: (long silence) I know why that happened.

Marge: Sir?

Feld: The original ASC 606 transition had a line item for rebate automation. I cut it. Sixty percent of that budget, gone, in a board review eighteen months ago. I needed the number to look better that quarter.

Thorne: So the manual process wasn’t negligence. It was a workaround for a decision made above the project team.

Feld: Yes. And I let two go live attempts fail without ever mentioning it.

Ai. Cassiopeia: I want to note, without judgment, that this reframes the entire risk register for this engagement.

Vex: It also means Oskar and Priya have been quietly cleaning up an executive decision for eighteen months and getting blamed for the mess.

Feld doesn’t flinch at that. He nods, slowly, like a man who’s been waiting a long time to hear someone say it out loud.

Feld: I sat in two go live retrospectives and let the room talk about “data quality issues” and “process gaps” without correcting anyone. Both times.

Marge: Why tell us now?

Feld: Because you showed me a demo where something broke on purpose, and it didn’t end the meeting. It built trust instead of losing it. I hadn’t seen that happen on this project before. Made me think maybe this was a room where the truth wouldn’t end the meeting either.

Sable: It won’t. But we do need to know the truth to fix this properly.

Feld: Then you have it. All of it. What do you need from me?

Feld doesn’t argue with that. He just nods, once, and asks the party what it will actually cost to do it right this time.


The decision that outlived the meeting where it was made

The rebate mess was never a technical problem. It was a budget decision made in a single board review, eighteen months before anyone in the war room ever heard about it, and it quietly outlived that meeting by becoming somebody else’s daily grind. That’s the part worth sitting with: Feld’s decision didn’t disappear when the meeting ended. It just changed shape, from a line item on a slide into a recurring manual task that Priya and then Oskar absorbed without ever being told why it existed.

This happens more than anyone likes to admit. A cut gets made at a level where the tradeoff looks clean, automation versus a number that needs to look better this quarter, and the actual cost doesn’t vanish. It relocates, usually downward, onto whoever is closest to the gap and least equipped to say no to it. Two failed go live retrospectives happened with Feld in the room, and both times the language stayed vague enough that nobody had to name where the gap actually came from. “Data quality issues.” “Process gaps.” Both true, technically, and both doing a lot of work to avoid a harder sentence.

What changes things here isn’t that Feld had information the party didn’t. It’s that episode seven gave him a reason to believe disclosure wouldn’t blow up the room. That’s the actual payoff of Marge’s rule about demos that are allowed to break: trust built in one context tends to extend into the next one, and a sponsor who’s just watched a team handle a real exception without flinching is a sponsor more likely to hand over an uncomfortable truth of his own.


What happens next

Feld’s confession reframes the whole rebate problem, but it doesn’t fix it. The party now has to build the actual business case for automating what should have been automated eighteen months ago, and figure out what it really costs to finally do this right, against everything it’s already cost to do it wrong for a year and a half.

Next episode: Episode 9, The Real Cost (coming soon)


If a cost cut made somewhere above your project has quietly become someone’s daily manual task, you’re looking at the same pattern. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a framework built around naming these gaps before they turn into eighteen months of invisible labor, start here: adnd365.com/start

Cause of death: the box got checked without anyone asking what checking it actually meant.


Security comes up in most enterprise demos as a single slide, usually near the end, usually delivered fast. Role-based access control. Single sign-on. Field-level security. Audit logging. SOC 2. Encryption at rest and in transit. Each term gets a checkmark, a confident nod from the presenter, and about four seconds of screen time before the deck moves on to something more visually interesting.

Nobody in the room stops the slide. That’s the autopsy. Every term on that slide is real, in the sense that the feature genuinely exists somewhere in the product. What’s missing is any demonstration that it does what the room assumes it does, configured the way your company would actually need it configured, at the depth your actual risk profile requires.

What actually happened

“Role-based access control” is not one feature. It’s a spectrum that runs from a handful of fixed roles with no customization, through role-based permissions you can tailor at the menu-item level, through field-level and record-level security that can restrict a single sensitive column or a single customer’s data from a specific user, through fully dynamic, condition-based access rules that change what someone can see based on context. A vendor can put “role-based access control” on a slide truthfully at any point on that spectrum, and the room has no way of knowing, from the slide alone, which end they’re getting.

The same collapse happens to every other term on the list. “Audit logging” might mean every field change on every table is captured with before and after values and an immutable timestamp, or it might mean a handful of high-level events get logged with no field-level detail. “SOC 2” tells you an audit happened and a report exists. It doesn’t tell you which trust service criteria were in scope, what the exceptions were, or whether the report is even current. “Encryption at rest” is true of nearly every modern cloud platform and tells you almost nothing about key management, who holds the keys, or what happens in a subpoena scenario.

None of this is the vendor lying. Every term is technically accurate. The compression happens in the translation from a nuanced, configurable capability into a single reassuring word on a slide, and the room does the rest of the work by filling in the most generous plausible interpretation of what that word means.

Why it works on smart people

Security and compliance are exactly the kind of topic where nobody in a sales meeting wants to be the person who asks a question that reveals they don’t fully understand the acronym. SOC 2 Type I versus Type II, the difference between authentication and authorization, what “field-level security” actually restricts versus what it merely hides in the UI while leaving accessible through an API, these are legitimate technical distinctions that most people in the room, including some people whose job title suggests they should know them, have only a fuzzy grasp of.

The slide exploits that fuzziness efficiently, not through malice but through pace. There’s no room in a four-second checkmark to unpack what a term actually covers, and stopping to ask “which SOC 2 trust service criteria were in scope” in the middle of a fast-moving demo carries a social cost that quietly discourages the question from being asked, the same social cost that lets the AI Magic Demo’s chat window go unexamined.

The actual damage

This is the one that surfaces during an actual security review, an actual audit, or worse, an actual incident, months or years after the contract was signed. Someone in security or compliance, doing real diligence for the first time, discovers that “field-level security” in this product means the field is hidden in the standard UI but fully readable through the API with the right permission, which is a meaningfully different security posture than what was assumed at signing. Or the SOC 2 report, once actually read line by line, turns out to have scoped out exactly the subsystem your data lives in.

At that point the conversation isn’t a demo follow-up. It’s a risk finding, sometimes one that has to be reported up to a board or disclosed to a regulator, and remediating it after the fact, whether that means additional configuration, a compensating control, or a vendor conversation about contractual commitments, is far more expensive and far more visible than it would have been to ask the specific question up front.

The fix, if you’re the one presenting

Don’t let the slide stand alone. For each term, say what it actually covers and, just as importantly, what it doesn’t. “Field-level security restricts visibility in the standard interface. It does not currently restrict API access to the same field, here’s how customers typically compensate for that.” “Our SOC 2 report covers these three trust service criteria, and here’s the exception list from the most recent period.” That level of specificity costs a few extra minutes and makes the product sound less flawless. It also means the prospect’s actual security review, whenever it happens, confirms what they were told instead of contradicting it.

A checkmark is not a control. It’s a promise that a control exists, and promises deserve exactly as much scrutiny as anything else on the slide.


The same collapse happens to job titles, not just checklist items. In AI Makers Don’t Build Models. They Build Value., a single contested word, “maker,” was doing more work than it could support, and the gap got filled by whoever was listening. A security slide runs on the exact same mechanism, one word standing in for a spectrum, and the room filling in whichever end feels most reassuring.

An Advanced Dungeons & Dynamics 365 session. Episode 7 of 14.

Previously on: Episode 6, The Ghost in AP, the party found an eighteen month manual rebate accrual sitting in a suspense account, discovered the numbers didn’t match the actual contract terms, and learned that Oskar had inherited the whole mess without feeling senior enough to flag it.

Four weeks to go live, and the party finally has to answer to someone who isn’t in the war room every day. This episode is about the gap between what a project looks like from the inside and what it looks like from a slide.


The Session

GM: Four weeks out. Steering committee. The sponsor, a man named Feld who has opinions about PowerPoint, wants to see progress.

Feld: Just walk me through it. Order to cash. Should take ten minutes.

Marge: (under her breath, to the party) Golden path or real path?

Thorne: Real path. Every time. That’s the whole point of this engagement.

Marge: Feld, I can show you the happy path in ten minutes. I’d rather show you the actual order flow, exceptions included, because that’s what your team will live in. It’s twenty minutes. Worth it.

Feld: (checking his watch) Fine. Twenty.

Sable: (casting Data Mapping live, in front of the committee) Watch what happens when the order hits a partial shipment. This is the case that broke the second go live attempt.

The demo hits the exception. It resolves correctly, slower than a golden path demo, but correctly. Feld leans forward.

Feld: That’s the first time in this whole engagement someone showed me something breaking on purpose.

Ai. Cassiopeia: Sponsor confidence metric, trending upward for the first time since project kickoff.

Feld: (genuinely curious now) Show me another one.

Sable: This is the freight variance case. Multi leg shipment, split invoice, partial return. It’s ugly. Watch.

The second exception resolves too, a little slower, a little messier, but every number lands where it should. Feld sits back, arms crossed, but not defensively. He’s thinking.

Feld: In eighteen months, nobody has shown me a demo that wasn’t a straight line from order to cash. Every single one looked perfect. And every single one was wrong within a week of go live.

Marge: That’s usually the tell. If a demo never breaks, it’s not testing anything. It’s performing.

Feld stays for the full twenty minutes. Then he asks a question nobody expected.

Feld: If the demo can survive that, why isn’t the rebate mess fixed yet?


The demo that’s allowed to break

There’s a particular kind of silence that happens when a demo hits a real exception in front of a room full of decision makers, and then resolves it correctly anyway. It’s the sound of a sponsor recalibrating what they actually believe about a project. Not because the exception was impressive, but because it was honest, and honesty had apparently been in short supply for eighteen months.

Feld’s line is the whole point of this episode: every demo he’d seen before this one looked perfect, and every one of them broke within a week of go live. That’s not a coincidence, and it’s not bad luck either. A demo that never shows an exception isn’t reassuring anyone. It’s teaching the sponsor to trust a version of the system that doesn’t exist. When reality inevitably shows up looking messier than the slide, the sponsor doesn’t think “well, exceptions happen.” They think “I was lied to,” because in a sense, they were, just not maliciously.

Marge’s rule from episode one finally pays off here: document the golden path, don’t perform it. It costs more time in the room. It’s a harder sell to a sponsor checking his watch. But it’s the only version of a demo that actually builds trust instead of borrowing it against a deadline that’s going to come due eventually, usually at the worst possible moment.

Feld’s response tells you everything about what eighteen months of golden path demos had cost this project. It wasn’t more confidence. It was less. He’d learned, correctly, not to trust what he was being shown, and it took one honest twenty minute demo to start reversing that.


What happens next

Feld’s question lands exactly where it should: if the team can handle a broken demo in front of a steering committee without flinching, why has the rebate accrual been sitting unresolved for eighteen months? The party doesn’t have a good answer yet, because the real answer goes back further than anyone in the room realizes, and it’s about to become Feld’s problem too.

Next episode: Episode 8, What Feld Knew (coming soon)


If your last steering committee demo didn’t show a single exception, that’s worth thinking about. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you’ve ever sat through a demo that only shows the golden path, you already know why this scene matters. It’s the entire premise behind Dog and Pony Autopsy, worth a read if this one landed.

If you want the configuration discipline behind demos that are actually allowed to break, start here: adnd365.com/start

Cause of death: ten records don’t behave like ten million, and nobody in the room was thinking in millions.


The presenter clicks a button. A report renders instantly. A search returns results before the loading spinner has time to spin. A batch job that would, in your world, run overnight, completes in the time it takes to say “and here’s the output.” The room absorbs all of this the way it absorbs everything else in a demo: as evidence of how the product performs.

It is not evidence of how the product performs. It is evidence of how the product performs against a database with four hundred customers, running on hardware provisioned for a demo, with exactly one user logged in, who is the only person generating any load at all. Your environment, on day one of production, will have none of those three things in common with it.

What actually happened

Performance is not a property of software in isolation. It’s a property of software under a specific load, against a specific dataset size, with a specific number of concurrent users doing specific things at the same time. A demo environment is engineered, whether deliberately or just by the natural economics of running a sales demo, to minimize all three of these variables simultaneously. Small dataset, dedicated hardware, single user. That is close to the best-case condition the software will ever run under, and it’s the condition you’re being shown as though it were representative.

The gap between that best case and your actual production reality is usually invisible in the room because nothing about the demo interface tells you the dataset is small. A grid with four hundred rows and a grid with four million rows look identical on screen, for the same reason they looked identical in the migration autopsy: you’re only ever looking at the same twenty rows at a time. Query performance, index behavior, and lock contention all degrade in ways that simply don’t exist yet at four hundred rows, and none of that shows up until the dataset, and the concurrent user count, both climb to something resembling your real operation.

Batch and integration jobs hide the same gap differently. A nightly job that processes four hundred records in three seconds tells you almost nothing about how it will behave processing four hundred thousand records at 2 a.m. while five other scheduled jobs are also competing for the same database connections. The three-second version and the six-hour version can be, technically, the exact same code.

Why it works on smart people

Performance is one of the few things a demo can show without narrating, which makes it feel more objective than the parts that require a presenter’s framing. Nobody has to make a claim about speed. You just watch it happen, in real time, and speed you watch with your own eyes reads as harder evidence than speed someone tells you about. That instinct is usually correct. It’s just being applied to a measurement taken under conditions that will never recur once the system goes live.

There’s also a scale-intuition gap that’s genuinely hard to close without direct experience. Most people don’t have a strong internal sense of how nonlinearly performance can degrade as data volume and concurrency grow. A system that feels instant at four hundred records doesn’t necessarily feel merely a little slower at four million. Depending on how indexes, queries, and locking are built, it can fall off a cliff at some threshold nobody in the demo room has any way of anticipating.

The actual damage

This is the one that shows up as a go-live incident rather than a slow discovery, because performance problems under real load tend to announce themselves immediately and all at once, usually during month-end close or the first day the full user base logs in simultaneously. Reports that took two seconds in the demo take four minutes. A batch process that ran in three seconds against sample data doesn’t finish before the next scheduled job needs the same resources, and the two start colliding every night.

The remediation at that point is expensive and disruptive in a way that early testing would not have been: emergency performance tuning, index rebuilding, sometimes an infrastructure upgrade that wasn’t budgeted, all happening under the worst possible conditions, with live users blocked and a go-live date already spent.

The fix, if you’re the one presenting

Test and show performance against something that resembles your prospect’s actual scale, not the vendor’s default demo dataset. If a true load test isn’t feasible in the sales cycle, at minimum say so explicitly: “this demo is running against four hundred sample records on dedicated hardware, here’s what we know about performance at your expected volume, and here’s how we’d validate it before go-live.” That sentence costs you the illusion of effortlessness. It buys you a prospect who understands what they’re actually being shown, and a performance conversation that happens in scoping instead of in a production incident.

Speed you watched with your own eyes is still only evidence of the conditions you watched it under. Ten records were never going to tell you what ten million would do.


I’ve argued in The Same Four Systems that the same organizational patterns show up whether you’re a corner store or a Fortune 500 company, just with higher stakes. That holds for structure. It doesn’t hold for performance. A query pattern that’s invisible at four hundred rows can become the whole story at four million, and no amount of pattern-matching from a small scale prepares you for exactly where that threshold sits.

An Advanced Dungeons & Dynamics 365 session. Episode 6 of 14.

Previously on: Episode 5, One Hour, the party fixed the integration that had been silently feeding the shadow ledger, and discovered the real problem was a five minute config fix hiding behind three weeks of assumed complexity.

Fixing one thing tends to reveal the next thing. That’s not bad luck. It’s what happens when you finally have clean data to look through instead of a mess to guess at. This episode, the party gets its first real look at what’s underneath everything else, and finds something that’s been there a lot longer than three weeks.


The Session

GM: The integration fix holds. Landed cost posts clean. The room exhales for the first time in five episodes.

Sable: Don’t relax yet. Validating the postings surfaced something else. There’s a recurring accrual in AP. Same amount, every month, for eighteen months. It posts to a suspense account and just sits there.

Thorne: Vendor?

Sable: Vendor code is a placeholder. “VEND TEMP 01.”

Marge: Another TEMP. Of course.

Ai. Cassiopeia: I cross referenced the amount against the vendor rebate program from the original ASC 606 transition. It matches a rebate structure that was supposed to be automated and never was.

Vex: So someone’s been manually estimating a rebate accrual for a vendor that might not even be configured right, for a year and a half, and just letting it sit in suspense?

Priya’s replacement (a nervous analyst named Oskar): That was me. I inherited it from Priya’s predecessor. Nobody told me to stop.

The room turns toward Oskar, who has been standing quietly near the door for most of the conversation, clearly hoping nobody would ask.

Marge: How does a manual accrual survive eighteen months and two go live attempts without anyone flagging it?

Oskar: (quietly) Because it closes clean every month. The number’s never wildly wrong. It’s just never provably right, either. Nobody’s ever had time to check.

Sable: (gently) You weren’t hiding this.

Oskar: I didn’t think I was allowed to raise it. I’m not senior enough to say the automated process was never built.

Ai. Cassiopeia: For the record, this is not a criticism of Oskar’s work. His estimation methodology is, if anything, unusually disciplined for a manual process running this long.

Thorne pulls up the original rebate contract terms. The numbers on the page do not match the numbers in the suspense account. Not by a little.

Thorne: Oskar. Whatever you were estimating against, it isn’t this contract.


The person doing the invisible work

Every implementation eventually finds an Oskar. Someone junior enough that they inherited a broken process rather than built it, and senior enough to keep it running competently for eighteen months without anyone above them noticing there was a problem at all. That combination, competent enough to hide the gap, junior enough to feel like they can’t raise it, is exactly what lets a process like this survive two failed go live attempts.

This isn’t a story about someone cutting corners. It’s the opposite. Oskar’s discipline is the reason nobody noticed sooner. A sloppier manual process would have thrown obviously wrong numbers and gotten caught in month two. A careful one, run by someone paying close attention every single month, can close clean for a year and a half while quietly drifting further and further from the actual contract terms underneath it.

The real failure here happened before Oskar ever touched the spreadsheet. Somewhere in the original ASC 606 transition, a line item for rebate automation didn’t get built, and instead of that gap showing up as a flagged risk, it got quietly absorbed by whoever happened to be sitting in the AP seat at the time. That’s the pattern worth naming: work that should have been a system’s job becomes a person’s invisible responsibility, and the person doing it often doesn’t feel like they have standing to say so out loud.


What happens next

The numbers on the contract and the numbers in the suspense account don’t match, and now the party has to figure out how far off they actually are, eighteen months deep, with go live four weeks away and a steering committee that’s about to start asking pointed questions about why the timeline hasn’t moved.

Next episode: Episode 7, The Steering Committee (coming soon)


If there’s a process on your team that’s being quietly kept alive by someone who doesn’t feel senior enough to flag it, this episode is for them too. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want to build the kind of system where gaps like this get caught in month two instead of month eighteen, start here: adnd365.com/start

Cause of death: nobody could agree on what the proof of concept was supposed to prove.


This one starts differently than the others. Every autopsy so far has been performed on a demo the vendor built to look better than reality. This one is performed on a demo the customer built, with the vendor’s help, to answer everything at once, and in doing so answered nothing.

The pattern is familiar to anyone who has scoped a proof of concept. It starts small and correct. One core question, one hypothesis, one thing that either works or doesn’t: can this system handle our multi-entity intercompany billing without a workaround, yes or no. Then someone from a different department hears there’s a POC happening and asks if it can also touch their process, since they’re curious too. Then someone senior asks for the trickiest edge case in the business to be included, because if it can’t handle that, what’s the point. By the time the scoping document is final, the POC that was supposed to answer one question is now attempting to demonstrate manufacturing, procurement, three approval hierarchies, a currency conversion edge case that occurs twice a year, and an integration to a system that isn’t even part of the actual project scope.

Nobody added any single piece of this in bad faith. That’s what makes it hard to stop once it’s moving.

What actually happened

A proof of concept exists to reduce uncertainty about one specific, high-risk question as cheaply and quickly as possible. The moment it starts trying to prove ten things instead of one, several things happen at once, all of them bad.

The build time stops being proportional to the risk being retired. A POC that answers one hard question can often be built in days, because the vendor and the team can focus every hour on the thing that actually matters. A POC trying to demonstrate ten things needs ten times the configuration, ten times the test data, ten times the edge cases handled, and none of that additional effort is reducing risk proportionally, because nine of the ten things were never actually in doubt.

The result also stops being interpretable. If the multi-entity billing scenario fails, but it failed inside a build that also included four other complex configurations layered on top of each other, you cannot cleanly attribute the failure. Was it the core capability that doesn’t exist, or a configuration conflict between two features that were never meant to be tested together, or a data setup error introduced trying to support scenario six while building scenario three. A focused POC gives you a clean signal. An overloaded one gives you noise that looks like a signal.

Why it happens to smart teams

The instinct behind scope creep in a POC is almost always defensible in isolation. Nobody wants to greenlight a six or seven figure implementation based on a narrow test, only to discover eight months in that some other critical process doesn’t fit either. The fear isn’t irrational. Systems do have gaps that only show up once you look in the right corner, and a POC feels like the cheapest moment to go looking.

The trouble is that “the cheapest moment to look” and “the cheapest way to look” are different questions. Looking broadly at low depth, a checklist of yes-or-no capability questions answered through documentation review, a reference call, or a scoped demo of specific features, retires broad risk cheaply. A single POC trying to go deep on ten things at once is neither cheap nor deep. It’s the expensive way to get a shallow answer to a question that didn’t need a POC to answer in the first place.

There’s also a political dimension that’s hard to name out loud in the room. Once word gets out that a POC is happening, being excluded from it can read as a signal that your department’s concerns don’t matter. Scope grows partly because saying no to an additional scenario feels like saying no to a person, not to a line item.

The actual damage

The POC that was supposed to take two weeks takes eight. The core question, the one thing that actually justified spending POC time and budget, gets buried under nine other questions that each needed their own edge case handling, and by the time results come back, the steering committee is looking at a partial success across ten dimensions instead of a clear answer on the one dimension that mattered. Decision paralysis follows almost automatically, because a partial, ambiguous result is much harder to act on than a clean pass or fail.

Worse, the actual high-risk question, the reason the POC existed in the first place, often gets the least rigorous testing of the ten, because it was scoped first and then diluted by everything added after it. The team spends real effort proving things that were never seriously in doubt, and comparatively little effort on the one thing that was.

The fix, if this is your situation

Separate what the POC needs to prove from what people merely want to see. Write down the single question, or at most two, whose answer would actually change the buying decision. Everything else, every “while we’re in there” request, goes on a second list explicitly labeled as out of scope for this exercise, with a stated plan for how it will get answered instead, whether that’s a reference call, a documentation review, or a second, later POC once the first question is settled.

When someone pushes back and asks why their scenario isn’t included, the honest answer is the useful one: this POC is designed to answer one hard question as cleanly as possible, and adding your scenario wouldn’t make the answer more trustworthy, it would make it harder to read. That’s not a dismissal of their concern. It’s a commitment to answering it properly, later, instead of poorly, now, buried inside somebody else’s test.

A POC that tries to prove everything proves nothing cleanly. The discipline isn’t in the build. It’s in what you refuse to put in it.


The same fragmentation shows up in individual work, not just project scoping. In The Forty One Percent Problem, I look at decades of research showing that a large, stable share of professional time gets eaten by low-judgment overhead scattered across too many things at once. A POC that tries to answer ten questions has the same disease as a workday that tries to touch ten priorities. Depth loses to breadth every time.

AI assistants are not a new idea. They are the first real attempt to fix a problem management research documented seventy years ago and never actually solved for most people.

Every pitch for an AI assistant makes some version of the same claim: it will take the administrative weight off your day so you can focus on the work that actually matters. That claim is usually presented as new. It is not. It is the exact same claim that got made about executive secretaries, and it is backed by the same research, run three separate times across seven decades, always finding the same number.

The short version: a large and remarkably stable share of professional work, something close to 40 percent, is low-judgment administrative overhead rather than the work someone was actually hired to do. For most of the twentieth century, the only fix on offer was a personal secretary, and that fix was rationed almost entirely by seniority. AI assistants are the first attempt to offer that same relief to everyone else. Here is where that number comes from, and why the history matters for judging whether the new fix is real.

The first study

In 1951, a Swedish economist named Sune Carlson did something nobody had done before. He asked a group of managing directors to keep detailed diaries of their working days, logged in real time rather than reconstructed from memory afterward. The result, published as Executive Behaviour, is generally regarded as the first systematic empirical study of what managers actually do with their time, as opposed to what people assumed they did.

The picture that emerged was not flattering. Executive days turned out to be fragmented, reactive, and dominated by short bursts of low-level activity. Phone calls. Correspondence. Brief conversations. Interruptions. The romantic image of the executive locked away making weighty strategic decisions bore little resemblance to the diaries. Most of the day was consumed by administrative traffic that did not require an executive’s judgment at all, it just required someone competent to handle it.

Carlson’s book was not widely read at the time. It was sparse on conclusions and thin on theory. But the method survived. Two decades later, Henry Mintzberg ran a similar observational study and arrived at the same finding, dressed in more modern language. Managers do not spend their days thinking. They spend their days responding.

The finding keeps getting rediscovered

What is striking is not that this was found once. It is that it has been found repeatedly, in different decades, using different methods, on different populations of workers, and the number keeps coming back roughly the same.

In 2013, Julian Birkinshaw and Jordan Cohen ran a three year study of knowledge workers and published the results in Harvard Business Review under the title “Make Time for the Work That Matters.” Their headline finding: knowledge workers spend an average of 41 percent of their time on discretionary activities that offer little personal satisfaction and could be handled competently by someone else. Participants who went through a structured process to identify and shed that work cut desk work by roughly six hours a week and meeting time by two more.

Five years later, Harvard Business School’s Michael Porter and Nitin Nohria published the results of tracking 27 Fortune 500 CEOs for a full year, more than 60,000 hours of coded time-use data. It remains the most detailed public look at how chief executives actually spend their days. The finding was not that these people were lazy or undisciplined. It was that even at the very top of an organization, with every resource available to protect their time, a large share of it still got eaten by things that did not need a CEO to do them.

Three studies, seven decades apart, using diaries, surveys, and direct observation, converging on the same basic fact: a large and remarkably stable percentage of professional work is low-judgment administrative overhead. Not incompetence. Not poor discipline. Just the physics of running an organization, where information has to move, calendars have to align, and someone has to read the email before anyone can decide whether it matters.

Why the fix was always rationed

Here is the part of the story that gets skipped over. The research kept finding the same problem, but the solution it kept prescribing, dedicated administrative support, was never distributed according to who had the problem. It was distributed according to seniority.

A junior employee drowning in the same 41 percent of low-value work as a CEO simply absorbed it. There was no equivalent relief further down the org chart. And as email, shared calendars, and self-service scheduling tools spread through the 1990s and 2000s, even the people who had once had dedicated secretarial support increasingly lost it, on the theory that the tools themselves had closed the gap. The secretary and executive assistant roles contracted sharply. The underlying problem the research had documented did not contract with them. It just became something everyone quietly managed alone.

This is worth sitting with, because it means the historical relationship between “having administrative support” and “being more productive” was never really tested at scale. It was tested at the top of organizations, on people who already had every other advantage, and the results were treated as proof of a general principle that was never actually given the chance to apply generally.

What changes when the support is not rationed by rank

The research gives a fairly precise description of what kind of work is worth taking off someone’s plate: correspondence, scheduling, screening, routine drafting, information retrieval, status tracking. Not judgment work. Not relationship work. The mechanical layer underneath both of those things.

That is also, not coincidentally, close to the exact list of tasks that current AI tools are best at and are being adopted for first. Which suggests the honest way to think about AI assistants is not as a replacement for a human secretary, doing the same job with different hardware. It is closer to the first real attempt to deliver the same category of relief that Carlson’s executives had, to people who were never senior enough to be given it.

That reframe matters because it changes what the interesting question is. The old question, does an executive with a secretary outperform one without, was really a proxy for a different question that took seventy years to ask properly: how much of anyone’s professional capacity is being spent on work that has nothing to do with why they were hired, and what happens when that overhead gets pushed down toward zero for everyone, not just the people at the top of the org chart.

There is a genuine limit to the analogy, and it is worth naming rather than skating past. A human assistant carries judgment and institutional memory that took Carlson’s own subjects years to build with the people supporting them, knowing which call to interrupt a meeting for and which one to let go to voicemail. Whether that layer of judgment gets replicated, approximated, or simply left undone is not a question the old research can answer, because the old research was never testing for it. What it can tell us, with unusual consistency across seventy years of data, is the size and shape of the problem being solved. That part, at least, is no longer a mystery.

An Advanced Dungeons & Dynamics 365 session. Episode 5 of 14.

Previously on: Episode 4, Page Seventeen, the party found the missing integration from the missing SOW page, an undocumented patch quietly feeding the shadow ledger, failing silently for three weeks with nobody watching.

Every implementation eventually produces a moment like this one: a real decision, a hard deadline, and not enough information to make the decision comfortably. This episode is that moment.


The Session

CFO: One hour. Fix it or freeze it. I have a board call.

Marge: Vex, can you patch the retry logic in an hour?

Vex: I can patch it. I can’t test it in an hour. Not against three weeks of backlog.

Thorne: Then we freeze it and post the backlog manually. Clean, auditable, slow.

Sable: Slow is fine. Wrong is not fine. I vote freeze.

Marge: I vote fix. If we freeze it now, it becomes tribal knowledge that “the integration is off” and nobody turns it back on. I’ve seen that before.

Thorne: Two and two.

Ai. Cassiopeia: I was not asked to vote. I will offer the information anyway. The backlog is not three weeks of data. It is three weeks of data plus a duplicate batch from a retry storm on day one. The real backlog is roughly half what Vex estimated.

Vex: …that changes the math. I can test that in an hour.

Marge: Then we fix it. Vex, go. Sable, you’re validating every posting before it touches the real ledger.

The room splits into motion. Vex pulls up the retry logic, hands moving fast. Sable stands over his shoulder, watching every line scroll past like she’s guarding a gate. Thorne quietly starts drafting a manual reconciliation plan anyway, just in case the fix doesn’t hold.

Thorne: (not looking up) Belt and suspenders.

Marge: Good instinct. Keep going.

Twenty minutes in, Vex hits something.

Vex: Found the actual bug. It’s not the retry logic itself. It’s a timeout value that’s too short for the freight audit feed on high volume days. The retries were never the problem. They were a symptom of a timeout nobody ever tuned.

Sable: So the real fix is smaller than we thought.

Vex: Much smaller. One config value.

Ai. Cassiopeia: I recommend documenting this timeout setting explicitly going forward. It has apparently been a default value since implementation, never revisited.

Forty one minutes later, Vex’s hands are still on the keyboard when the CFO walks back in early.

CFO: Well?


Why Cassiopeia’s vote mattered more than a vote

The most interesting moment in this scene isn’t the fix. It’s the tie. Two votes for freeze, two for fix, and the deciding information came from the one party member who doesn’t get a formal vote at all. That’s not a coincidence, and it’s not there to make a point about AI having opinions. It’s there because Ai. Cassiopeia had access to something the humans in the room didn’t have time to check under pressure: the actual shape of the backlog, not the estimated shape of it.

Under a real deadline, teams default to their gut. Freeze feels safer because it’s reversible. Fix feels riskier because it isn’t. Both instincts are reasonable, and both were wrong in the same way, because both were built on an estimate nobody had time to verify. The team wasn’t voting on values. They were voting on incomplete information, and neither side knew it.

This is the actual case for a grey collar workforce member in a room like this one. Not to replace judgment, Marge and Thorne still made the call, but to remove the guesswork underneath the judgment before the decision gets made instead of after. And notice what Vex found once the pressure came off enough to actually look: the retry logic everyone assumed was broken wasn’t the real problem at all. It was a five minute config fix hiding behind three weeks of assumed complexity.

That’s usually how it goes. The scary problem and the real problem are rarely the same problem.


What happens next

The fix holds, and landed cost starts posting clean for the first time in weeks. But validating every posting before it touches the real ledger means someone has to look closely at everything already in the system, and that closer look surfaces something nobody in the room was looking for: a second, older shadow process that’s been quietly running in the background of Contoso’s books for a lot longer than three weeks.

Next episode: Episode 6, The Ghost in AP (coming soon)


If your team has ever voted on a fix under a deadline without actually knowing the real size of the problem, you’re in good company. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a framework for surfacing the real problem before the deadline forces a guess, start here: adnd365.com/start

Cause of death: three years of production data quietly became four hundred clean sample records for the day of the pitch.


Somewhere in the middle of the demo, the presenter opens a grid. Customers, items, transactions, whatever the domain calls for. It scrolls smoothly. Every row has every field populated. Names are properly capitalized. Addresses have all their parts. There are no duplicate customer records for “Acme Corp,” “ACME Corp,” and “Acme Corp.” with a trailing space that your actual system has accumulated over a decade of different people typing the same name slightly differently.

The presenter doesn’t say “this is sample data.” They don’t need to. The grid looks so much like a real company’s data that the distinction quietly stops mattering to the room, and everyone leaves the meeting having watched a migration that never happened, of data that was never really yours.

What actually happened

Four hundred rows of clean, plausible-looking data is not a migration. It’s a mockup wearing a migration’s clothes. Somebody built that dataset specifically to demonstrate the target system’s data model, which means it was constructed backward from what the target system wants to receive, rather than forward from what your source system actually contains. It has never been through a real extract. It has never hit a field length limit, a character encoding mismatch, a required field that’s been null in your source system since 2019 because nobody enforced it, or a foreign key that points to a parent record that got deleted three reorganizations ago.

Real migrations die on exactly these details, and none of them are visible in a four hundred row demo grid, because the demo grid was never subjected to the process that would surface them. The sample data is a hypothesis about what your data looks like. It has not yet met your data.

There’s a second layer under this. Even when a vendor does an actual proof of concept against a real extract of your data, that extract is usually a snapshot, cleaned once, run through a mapping exercise once, and shown once. It demonstrates that a migration is possible for that slice, on that day, with that much attention paid to it. It does not demonstrate that the full historical dataset, run through the same process without the benefit of a team hand-tuning exceptions in real time, will produce the same result.

Why it works on smart people

Data problems are boring in a way that makes them easy to underestimate from the outside. Nobody gets excited describing thirty thousand customer records with inconsistent capitalization, or a decade of transactions where the currency field was optional for the first four years, and that lack of drama works against the diligence the problem deserves. A demo that skips the data reality skips the part of the story that was never going to be compelling to watch anyway, and audiences let it go for the same reason they let the seamlessness of the Golden Path Demo go: friction is a strange thing to ask someone to add back in.

There’s also a scale-blindness effect. A grid of four hundred rows and a database of four million rows look identical in a screen share, because you’re only ever looking at the same twenty rows on screen at once. The demo cannot visually communicate that the four hundred clean rows are a curated sliver, not a representative sample, so the brain does what it usually does with limited visual information: it extrapolates, and assumes the part it can see generalizes to the whole.

The actual damage

This is the one that blows up the project timeline more reliably than almost anything else, and it does it quietly, in the data cleansing and reconciliation phase that was budgeted as a two-week task because the demo made data migration look like a solved problem. Then someone runs the real extract, and it turns out eight percent of vendor records have no valid tax ID, eleven percent of item records reference a unit of measure that was deprecated four years ago, and there are nineteen thousand duplicate customer records that need to be identified and merged before go-live, none of which showed up in four hundred rows of hand-picked sample data.

The two-week task becomes a two-month task, the go-live date moves, and the business case that assumed a smooth data conversion now has to absorb a delay that nobody priced in, because the thing that actually determines a migration’s difficulty, the messiness of the real data, was the one thing the demo was specifically built not to show.

The fix, if you’re the one presenting

Run the demo against a real, ugly extract, even a small one, and don’t clean it first. Show the duplicate detection running against actual duplicates. Show what happens when a required field is null. Show the exception queue, and how many records land in it, and what the resolution workflow actually looks like for the person who has to work through that queue by hand. It’s a less polished five minutes. It’s also the only five minutes that tells the prospect anything real about what their conversion will cost.

A migration demo that never encounters bad data hasn’t demonstrated a migration. It’s demonstrated the destination.


The mapping exercise itself looks tidy in a lab too. In Familiar Ground: Mapping CRM to ERP, every concept has a clean twin on the other side, names changed, forms wider, same underlying logic. Real source data is rarely that cooperative, which is exactly the gap this autopsy is about.