An Advanced Dungeons & Dynamics 365 session. Episode 6 of 14.

Previously on: Episode 5, One Hour, the party fixed the integration that had been silently feeding the shadow ledger, and discovered the real problem was a five minute config fix hiding behind three weeks of assumed complexity.

Fixing one thing tends to reveal the next thing. That’s not bad luck. It’s what happens when you finally have clean data to look through instead of a mess to guess at. This episode, the party gets its first real look at what’s underneath everything else, and finds something that’s been there a lot longer than three weeks.


The Session

GM: The integration fix holds. Landed cost posts clean. The room exhales for the first time in five episodes.

Sable: Don’t relax yet. Validating the postings surfaced something else. There’s a recurring accrual in AP. Same amount, every month, for eighteen months. It posts to a suspense account and just sits there.

Thorne: Vendor?

Sable: Vendor code is a placeholder. “VEND TEMP 01.”

Marge: Another TEMP. Of course.

Ai. Cassiopeia: I cross referenced the amount against the vendor rebate program from the original ASC 606 transition. It matches a rebate structure that was supposed to be automated and never was.

Vex: So someone’s been manually estimating a rebate accrual for a vendor that might not even be configured right, for a year and a half, and just letting it sit in suspense?

Priya’s replacement (a nervous analyst named Oskar): That was me. I inherited it from Priya’s predecessor. Nobody told me to stop.

The room turns toward Oskar, who has been standing quietly near the door for most of the conversation, clearly hoping nobody would ask.

Marge: How does a manual accrual survive eighteen months and two go live attempts without anyone flagging it?

Oskar: (quietly) Because it closes clean every month. The number’s never wildly wrong. It’s just never provably right, either. Nobody’s ever had time to check.

Sable: (gently) You weren’t hiding this.

Oskar: I didn’t think I was allowed to raise it. I’m not senior enough to say the automated process was never built.

Ai. Cassiopeia: For the record, this is not a criticism of Oskar’s work. His estimation methodology is, if anything, unusually disciplined for a manual process running this long.

Thorne pulls up the original rebate contract terms. The numbers on the page do not match the numbers in the suspense account. Not by a little.

Thorne: Oskar. Whatever you were estimating against, it isn’t this contract.


The person doing the invisible work

Every implementation eventually finds an Oskar. Someone junior enough that they inherited a broken process rather than built it, and senior enough to keep it running competently for eighteen months without anyone above them noticing there was a problem at all. That combination, competent enough to hide the gap, junior enough to feel like they can’t raise it, is exactly what lets a process like this survive two failed go live attempts.

This isn’t a story about someone cutting corners. It’s the opposite. Oskar’s discipline is the reason nobody noticed sooner. A sloppier manual process would have thrown obviously wrong numbers and gotten caught in month two. A careful one, run by someone paying close attention every single month, can close clean for a year and a half while quietly drifting further and further from the actual contract terms underneath it.

The real failure here happened before Oskar ever touched the spreadsheet. Somewhere in the original ASC 606 transition, a line item for rebate automation didn’t get built, and instead of that gap showing up as a flagged risk, it got quietly absorbed by whoever happened to be sitting in the AP seat at the time. That’s the pattern worth naming: work that should have been a system’s job becomes a person’s invisible responsibility, and the person doing it often doesn’t feel like they have standing to say so out loud.


What happens next

The numbers on the contract and the numbers in the suspense account don’t match, and now the party has to figure out how far off they actually are, eighteen months deep, with go live four weeks away and a steering committee that’s about to start asking pointed questions about why the timeline hasn’t moved.

Next episode: Episode 7, The Steering Committee (coming soon)


If there’s a process on your team that’s being quietly kept alive by someone who doesn’t feel senior enough to flag it, this episode is for them too. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want to build the kind of system where gaps like this get caught in month two instead of month eighteen, start here: adnd365.com/start

Cause of death: nobody could agree on what the proof of concept was supposed to prove.


This one starts differently than the others. Every autopsy so far has been performed on a demo the vendor built to look better than reality. This one is performed on a demo the customer built, with the vendor’s help, to answer everything at once, and in doing so answered nothing.

The pattern is familiar to anyone who has scoped a proof of concept. It starts small and correct. One core question, one hypothesis, one thing that either works or doesn’t: can this system handle our multi-entity intercompany billing without a workaround, yes or no. Then someone from a different department hears there’s a POC happening and asks if it can also touch their process, since they’re curious too. Then someone senior asks for the trickiest edge case in the business to be included, because if it can’t handle that, what’s the point. By the time the scoping document is final, the POC that was supposed to answer one question is now attempting to demonstrate manufacturing, procurement, three approval hierarchies, a currency conversion edge case that occurs twice a year, and an integration to a system that isn’t even part of the actual project scope.

Nobody added any single piece of this in bad faith. That’s what makes it hard to stop once it’s moving.

What actually happened

A proof of concept exists to reduce uncertainty about one specific, high-risk question as cheaply and quickly as possible. The moment it starts trying to prove ten things instead of one, several things happen at once, all of them bad.

The build time stops being proportional to the risk being retired. A POC that answers one hard question can often be built in days, because the vendor and the team can focus every hour on the thing that actually matters. A POC trying to demonstrate ten things needs ten times the configuration, ten times the test data, ten times the edge cases handled, and none of that additional effort is reducing risk proportionally, because nine of the ten things were never actually in doubt.

The result also stops being interpretable. If the multi-entity billing scenario fails, but it failed inside a build that also included four other complex configurations layered on top of each other, you cannot cleanly attribute the failure. Was it the core capability that doesn’t exist, or a configuration conflict between two features that were never meant to be tested together, or a data setup error introduced trying to support scenario six while building scenario three. A focused POC gives you a clean signal. An overloaded one gives you noise that looks like a signal.

Why it happens to smart teams

The instinct behind scope creep in a POC is almost always defensible in isolation. Nobody wants to greenlight a six or seven figure implementation based on a narrow test, only to discover eight months in that some other critical process doesn’t fit either. The fear isn’t irrational. Systems do have gaps that only show up once you look in the right corner, and a POC feels like the cheapest moment to go looking.

The trouble is that “the cheapest moment to look” and “the cheapest way to look” are different questions. Looking broadly at low depth, a checklist of yes-or-no capability questions answered through documentation review, a reference call, or a scoped demo of specific features, retires broad risk cheaply. A single POC trying to go deep on ten things at once is neither cheap nor deep. It’s the expensive way to get a shallow answer to a question that didn’t need a POC to answer in the first place.

There’s also a political dimension that’s hard to name out loud in the room. Once word gets out that a POC is happening, being excluded from it can read as a signal that your department’s concerns don’t matter. Scope grows partly because saying no to an additional scenario feels like saying no to a person, not to a line item.

The actual damage

The POC that was supposed to take two weeks takes eight. The core question, the one thing that actually justified spending POC time and budget, gets buried under nine other questions that each needed their own edge case handling, and by the time results come back, the steering committee is looking at a partial success across ten dimensions instead of a clear answer on the one dimension that mattered. Decision paralysis follows almost automatically, because a partial, ambiguous result is much harder to act on than a clean pass or fail.

Worse, the actual high-risk question, the reason the POC existed in the first place, often gets the least rigorous testing of the ten, because it was scoped first and then diluted by everything added after it. The team spends real effort proving things that were never seriously in doubt, and comparatively little effort on the one thing that was.

The fix, if this is your situation

Separate what the POC needs to prove from what people merely want to see. Write down the single question, or at most two, whose answer would actually change the buying decision. Everything else, every “while we’re in there” request, goes on a second list explicitly labeled as out of scope for this exercise, with a stated plan for how it will get answered instead, whether that’s a reference call, a documentation review, or a second, later POC once the first question is settled.

When someone pushes back and asks why their scenario isn’t included, the honest answer is the useful one: this POC is designed to answer one hard question as cleanly as possible, and adding your scenario wouldn’t make the answer more trustworthy, it would make it harder to read. That’s not a dismissal of their concern. It’s a commitment to answering it properly, later, instead of poorly, now, buried inside somebody else’s test.

A POC that tries to prove everything proves nothing cleanly. The discipline isn’t in the build. It’s in what you refuse to put in it.


The same fragmentation shows up in individual work, not just project scoping. In The Forty One Percent Problem, I look at decades of research showing that a large, stable share of professional time gets eaten by low-judgment overhead scattered across too many things at once. A POC that tries to answer ten questions has the same disease as a workday that tries to touch ten priorities. Depth loses to breadth every time.

AI assistants are not a new idea. They are the first real attempt to fix a problem management research documented seventy years ago and never actually solved for most people.

Every pitch for an AI assistant makes some version of the same claim: it will take the administrative weight off your day so you can focus on the work that actually matters. That claim is usually presented as new. It is not. It is the exact same claim that got made about executive secretaries, and it is backed by the same research, run three separate times across seven decades, always finding the same number.

The short version: a large and remarkably stable share of professional work, something close to 40 percent, is low-judgment administrative overhead rather than the work someone was actually hired to do. For most of the twentieth century, the only fix on offer was a personal secretary, and that fix was rationed almost entirely by seniority. AI assistants are the first attempt to offer that same relief to everyone else. Here is where that number comes from, and why the history matters for judging whether the new fix is real.

The first study

In 1951, a Swedish economist named Sune Carlson did something nobody had done before. He asked a group of managing directors to keep detailed diaries of their working days, logged in real time rather than reconstructed from memory afterward. The result, published as Executive Behaviour, is generally regarded as the first systematic empirical study of what managers actually do with their time, as opposed to what people assumed they did.

The picture that emerged was not flattering. Executive days turned out to be fragmented, reactive, and dominated by short bursts of low-level activity. Phone calls. Correspondence. Brief conversations. Interruptions. The romantic image of the executive locked away making weighty strategic decisions bore little resemblance to the diaries. Most of the day was consumed by administrative traffic that did not require an executive’s judgment at all, it just required someone competent to handle it.

Carlson’s book was not widely read at the time. It was sparse on conclusions and thin on theory. But the method survived. Two decades later, Henry Mintzberg ran a similar observational study and arrived at the same finding, dressed in more modern language. Managers do not spend their days thinking. They spend their days responding.

The finding keeps getting rediscovered

What is striking is not that this was found once. It is that it has been found repeatedly, in different decades, using different methods, on different populations of workers, and the number keeps coming back roughly the same.

In 2013, Julian Birkinshaw and Jordan Cohen ran a three year study of knowledge workers and published the results in Harvard Business Review under the title “Make Time for the Work That Matters.” Their headline finding: knowledge workers spend an average of 41 percent of their time on discretionary activities that offer little personal satisfaction and could be handled competently by someone else. Participants who went through a structured process to identify and shed that work cut desk work by roughly six hours a week and meeting time by two more.

Five years later, Harvard Business School’s Michael Porter and Nitin Nohria published the results of tracking 27 Fortune 500 CEOs for a full year, more than 60,000 hours of coded time-use data. It remains the most detailed public look at how chief executives actually spend their days. The finding was not that these people were lazy or undisciplined. It was that even at the very top of an organization, with every resource available to protect their time, a large share of it still got eaten by things that did not need a CEO to do them.

Three studies, seven decades apart, using diaries, surveys, and direct observation, converging on the same basic fact: a large and remarkably stable percentage of professional work is low-judgment administrative overhead. Not incompetence. Not poor discipline. Just the physics of running an organization, where information has to move, calendars have to align, and someone has to read the email before anyone can decide whether it matters.

Why the fix was always rationed

Here is the part of the story that gets skipped over. The research kept finding the same problem, but the solution it kept prescribing, dedicated administrative support, was never distributed according to who had the problem. It was distributed according to seniority.

A junior employee drowning in the same 41 percent of low-value work as a CEO simply absorbed it. There was no equivalent relief further down the org chart. And as email, shared calendars, and self-service scheduling tools spread through the 1990s and 2000s, even the people who had once had dedicated secretarial support increasingly lost it, on the theory that the tools themselves had closed the gap. The secretary and executive assistant roles contracted sharply. The underlying problem the research had documented did not contract with them. It just became something everyone quietly managed alone.

This is worth sitting with, because it means the historical relationship between “having administrative support” and “being more productive” was never really tested at scale. It was tested at the top of organizations, on people who already had every other advantage, and the results were treated as proof of a general principle that was never actually given the chance to apply generally.

What changes when the support is not rationed by rank

The research gives a fairly precise description of what kind of work is worth taking off someone’s plate: correspondence, scheduling, screening, routine drafting, information retrieval, status tracking. Not judgment work. Not relationship work. The mechanical layer underneath both of those things.

That is also, not coincidentally, close to the exact list of tasks that current AI tools are best at and are being adopted for first. Which suggests the honest way to think about AI assistants is not as a replacement for a human secretary, doing the same job with different hardware. It is closer to the first real attempt to deliver the same category of relief that Carlson’s executives had, to people who were never senior enough to be given it.

That reframe matters because it changes what the interesting question is. The old question, does an executive with a secretary outperform one without, was really a proxy for a different question that took seventy years to ask properly: how much of anyone’s professional capacity is being spent on work that has nothing to do with why they were hired, and what happens when that overhead gets pushed down toward zero for everyone, not just the people at the top of the org chart.

There is a genuine limit to the analogy, and it is worth naming rather than skating past. A human assistant carries judgment and institutional memory that took Carlson’s own subjects years to build with the people supporting them, knowing which call to interrupt a meeting for and which one to let go to voicemail. Whether that layer of judgment gets replicated, approximated, or simply left undone is not a question the old research can answer, because the old research was never testing for it. What it can tell us, with unusual consistency across seventy years of data, is the size and shape of the problem being solved. That part, at least, is no longer a mystery.

An Advanced Dungeons & Dynamics 365 session. Episode 5 of 14.

Previously on: Episode 4, Page Seventeen, the party found the missing integration from the missing SOW page, an undocumented patch quietly feeding the shadow ledger, failing silently for three weeks with nobody watching.

Every implementation eventually produces a moment like this one: a real decision, a hard deadline, and not enough information to make the decision comfortably. This episode is that moment.


The Session

CFO: One hour. Fix it or freeze it. I have a board call.

Marge: Vex, can you patch the retry logic in an hour?

Vex: I can patch it. I can’t test it in an hour. Not against three weeks of backlog.

Thorne: Then we freeze it and post the backlog manually. Clean, auditable, slow.

Sable: Slow is fine. Wrong is not fine. I vote freeze.

Marge: I vote fix. If we freeze it now, it becomes tribal knowledge that “the integration is off” and nobody turns it back on. I’ve seen that before.

Thorne: Two and two.

Ai. Cassiopeia: I was not asked to vote. I will offer the information anyway. The backlog is not three weeks of data. It is three weeks of data plus a duplicate batch from a retry storm on day one. The real backlog is roughly half what Vex estimated.

Vex: …that changes the math. I can test that in an hour.

Marge: Then we fix it. Vex, go. Sable, you’re validating every posting before it touches the real ledger.

The room splits into motion. Vex pulls up the retry logic, hands moving fast. Sable stands over his shoulder, watching every line scroll past like she’s guarding a gate. Thorne quietly starts drafting a manual reconciliation plan anyway, just in case the fix doesn’t hold.

Thorne: (not looking up) Belt and suspenders.

Marge: Good instinct. Keep going.

Twenty minutes in, Vex hits something.

Vex: Found the actual bug. It’s not the retry logic itself. It’s a timeout value that’s too short for the freight audit feed on high volume days. The retries were never the problem. They were a symptom of a timeout nobody ever tuned.

Sable: So the real fix is smaller than we thought.

Vex: Much smaller. One config value.

Ai. Cassiopeia: I recommend documenting this timeout setting explicitly going forward. It has apparently been a default value since implementation, never revisited.

Forty one minutes later, Vex’s hands are still on the keyboard when the CFO walks back in early.

CFO: Well?


Why Cassiopeia’s vote mattered more than a vote

The most interesting moment in this scene isn’t the fix. It’s the tie. Two votes for freeze, two for fix, and the deciding information came from the one party member who doesn’t get a formal vote at all. That’s not a coincidence, and it’s not there to make a point about AI having opinions. It’s there because Ai. Cassiopeia had access to something the humans in the room didn’t have time to check under pressure: the actual shape of the backlog, not the estimated shape of it.

Under a real deadline, teams default to their gut. Freeze feels safer because it’s reversible. Fix feels riskier because it isn’t. Both instincts are reasonable, and both were wrong in the same way, because both were built on an estimate nobody had time to verify. The team wasn’t voting on values. They were voting on incomplete information, and neither side knew it.

This is the actual case for a grey collar workforce member in a room like this one. Not to replace judgment, Marge and Thorne still made the call, but to remove the guesswork underneath the judgment before the decision gets made instead of after. And notice what Vex found once the pressure came off enough to actually look: the retry logic everyone assumed was broken wasn’t the real problem at all. It was a five minute config fix hiding behind three weeks of assumed complexity.

That’s usually how it goes. The scary problem and the real problem are rarely the same problem.


What happens next

The fix holds, and landed cost starts posting clean for the first time in weeks. But validating every posting before it touches the real ledger means someone has to look closely at everything already in the system, and that closer look surfaces something nobody in the room was looking for: a second, older shadow process that’s been quietly running in the background of Contoso’s books for a lot longer than three weeks.

Next episode: Episode 6, The Ghost in AP (coming soon)


If your team has ever voted on a fix under a deadline without actually knowing the real size of the problem, you’re in good company. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a framework for surfacing the real problem before the deadline forces a guess, start here: adnd365.com/start

Cause of death: three years of production data quietly became four hundred clean sample records for the day of the pitch.


Somewhere in the middle of the demo, the presenter opens a grid. Customers, items, transactions, whatever the domain calls for. It scrolls smoothly. Every row has every field populated. Names are properly capitalized. Addresses have all their parts. There are no duplicate customer records for “Acme Corp,” “ACME Corp,” and “Acme Corp.” with a trailing space that your actual system has accumulated over a decade of different people typing the same name slightly differently.

The presenter doesn’t say “this is sample data.” They don’t need to. The grid looks so much like a real company’s data that the distinction quietly stops mattering to the room, and everyone leaves the meeting having watched a migration that never happened, of data that was never really yours.

What actually happened

Four hundred rows of clean, plausible-looking data is not a migration. It’s a mockup wearing a migration’s clothes. Somebody built that dataset specifically to demonstrate the target system’s data model, which means it was constructed backward from what the target system wants to receive, rather than forward from what your source system actually contains. It has never been through a real extract. It has never hit a field length limit, a character encoding mismatch, a required field that’s been null in your source system since 2019 because nobody enforced it, or a foreign key that points to a parent record that got deleted three reorganizations ago.

Real migrations die on exactly these details, and none of them are visible in a four hundred row demo grid, because the demo grid was never subjected to the process that would surface them. The sample data is a hypothesis about what your data looks like. It has not yet met your data.

There’s a second layer under this. Even when a vendor does an actual proof of concept against a real extract of your data, that extract is usually a snapshot, cleaned once, run through a mapping exercise once, and shown once. It demonstrates that a migration is possible for that slice, on that day, with that much attention paid to it. It does not demonstrate that the full historical dataset, run through the same process without the benefit of a team hand-tuning exceptions in real time, will produce the same result.

Why it works on smart people

Data problems are boring in a way that makes them easy to underestimate from the outside. Nobody gets excited describing thirty thousand customer records with inconsistent capitalization, or a decade of transactions where the currency field was optional for the first four years, and that lack of drama works against the diligence the problem deserves. A demo that skips the data reality skips the part of the story that was never going to be compelling to watch anyway, and audiences let it go for the same reason they let the seamlessness of the Golden Path Demo go: friction is a strange thing to ask someone to add back in.

There’s also a scale-blindness effect. A grid of four hundred rows and a database of four million rows look identical in a screen share, because you’re only ever looking at the same twenty rows on screen at once. The demo cannot visually communicate that the four hundred clean rows are a curated sliver, not a representative sample, so the brain does what it usually does with limited visual information: it extrapolates, and assumes the part it can see generalizes to the whole.

The actual damage

This is the one that blows up the project timeline more reliably than almost anything else, and it does it quietly, in the data cleansing and reconciliation phase that was budgeted as a two-week task because the demo made data migration look like a solved problem. Then someone runs the real extract, and it turns out eight percent of vendor records have no valid tax ID, eleven percent of item records reference a unit of measure that was deprecated four years ago, and there are nineteen thousand duplicate customer records that need to be identified and merged before go-live, none of which showed up in four hundred rows of hand-picked sample data.

The two-week task becomes a two-month task, the go-live date moves, and the business case that assumed a smooth data conversion now has to absorb a delay that nobody priced in, because the thing that actually determines a migration’s difficulty, the messiness of the real data, was the one thing the demo was specifically built not to show.

The fix, if you’re the one presenting

Run the demo against a real, ugly extract, even a small one, and don’t clean it first. Show the duplicate detection running against actual duplicates. Show what happens when a required field is null. Show the exception queue, and how many records land in it, and what the resolution workflow actually looks like for the person who has to work through that queue by hand. It’s a less polished five minutes. It’s also the only five minutes that tells the prospect anything real about what their conversion will cost.

A migration demo that never encounters bad data hasn’t demonstrated a migration. It’s demonstrated the destination.


The mapping exercise itself looks tidy in a lab too. In Familiar Ground: Mapping CRM to ERP, every concept has a clean twin on the other side, names changed, forms wider, same underlying logic. Real source data is rarely that cooperative, which is exactly the gap this autopsy is about.

An Advanced Dungeons & Dynamics 365 session. Episode 4 of 14.

Previously on: Episode 3, The Labyrinth of Financial Dimensions, the party mapped a chart of accounts that turned out to be three different structures stapled together, and found a dimension called TEMP_DO_NOT_USE quietly carrying eleven thousand postings nobody remembered approving.

Every SOW has an appendix nobody reads until they have to. For Contoso, that appendix was page seventeen, the integration list, and it went missing before the party ever set foot in the building. This episode, they finally find it.


The Session

GM: Vex finds page seventeen. It was stapled into the wrong appendix.

Vex: Six integrations. EDI for two vendors, a freight audit feed, a payroll export, a legacy AR aging tool, and… this one.

Thorne: Which one?

Vex: No name. No owner. Just an endpoint.

Ai. Cassiopeia: I have log access. It has been failing silently since three weeks ago. Retry logic swallows the error.

Sable: Failing into what?

Vex: (long pause) It posts landed cost adjustments. Into the shadow ledger.

Marge: So the spreadsheet Priya built wasn’t tracking a gap in the system. The system was quietly feeding it.

Thorne: Then the six figure variance from episode two…

Ai. Cassiopeia: Is not a variance. It is three weeks of unposted freight true ups sitting in a dead queue.

The room goes quiet. Somewhere down the hall, a fax machine that shouldn’t still exist starts printing.

Vex: (staring at the endpoint) This integration doesn’t have a name because whoever built it never expected anyone to find it. Look at the naming convention. It doesn’t match anything else in the environment.

Sable: It’s not part of the original build.

Vex: No. It’s a patch. Somebody built this after the fact, quietly, to route around a problem they didn’t have time to fix properly. And then never told anyone it existed.

Marge: Which means it was never in scope for testing. Never in scope for documentation. Never in scope for anything.

Ai. Cassiopeia: Correct. It has been running in the dark since before this engagement began.

Thorne: (quietly) We didn’t inherit a shadow ledger problem. We inherited a shadow integration that created one.


The workaround nobody admits to building

Every party in this scenario has been telling the truth as they understood it. The CFO didn’t know the integration existed. Priya didn’t know her shadow ledger was being fed by something outside her control. Even the original consultants probably didn’t know, since this integration was never part of the documented build. Somebody, at some point, hit a wall, needed the numbers to move, and built a bridge nobody else was told about.

This is the quiet failure mode that doesn’t show up in a status report. Not a missed deadline. Not a scope change. A small, undocumented patch, built with good intentions under real pressure, that outlives the person who built it and the reason they built it. Nobody decided to create three weeks of silent failure. Somebody just wrote retry logic that swallowed errors instead of surfacing them, because surfacing them would have meant admitting the patch existed in the first place.

The lesson here isn’t “don’t build workarounds.” Sometimes a workaround is the only thing standing between a business and a missed close. The lesson is that a workaround without an owner, without documentation, and without visibility into its own failures isn’t a bridge. It’s a debt that compounds silently until someone like Vex has to go digging for an endpoint with no name.


What happens next

The party now knows exactly where the six figure variance came from, and it isn’t a variance at all, just three weeks of backlog stuck behind a silent failure. But knowing the cause and fixing it in time are two different problems. The CFO wants an answer before his board call, and he’s giving the party one hour to decide whether to patch the integration live or freeze it and clean up the backlog by hand.

Next episode: Episode 5, One Hour (coming soon)


If there’s an integration running quietly in your environment that nobody currently owns, you’re not the only one. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want the configuration discipline that keeps workarounds like this one from going dark in the first place, start here: adnd365.com/start

Cause of death: two systems that had never spoken before were shown having a conversation written for them.


The presenter switches windows. A record gets created in System A. A few seconds later, as if by magic, the corresponding record appears in System B, fully formed, correctly mapped, no errors. “And that’s it,” the presenter says, “they just talk to each other.” The room relaxes. Integration, historically the single most reliable way for an implementation to go over budget and past deadline, has apparently been solved by two systems having a friendly chat while everyone watched.

Nobody in the room asks what was actually watching that conversation, or who taught it what to say. That’s the autopsy. The two systems didn’t learn to talk to each other. Someone wrote both sides of the script, tested it exactly once, against exactly one scenario, and ran it live in front of you.

What actually happened

An integration demo almost never shows the integration. It shows the happy path of the integration, which is a different and much smaller thing. The record that got created in System A was built to contain precisely the fields System B expects, in precisely the format System B expects them, with no null values in the fields that would trigger a mapping error, no duplicate keys, no encoding mismatch, none of the thousand small inconsistencies that live in a company’s actual data the moment more than one person or one legacy system has touched it.

The script connecting the two systems, whether it’s a middleware platform, a custom connector, or a scheduled job, was very likely written specifically for this demo, tuned against this one scenario, and has never been asked to handle a partial failure, a duplicate record, a field that arrives populated in one system and empty in the other, or a timeout on either end. It works, in the same sense that a bridge works if you only ever drive one specific car across it at one specific speed.

The part that never gets demoed, because it can’t be demoed in three minutes, is everything that happens when the sync fails halfway through. Does the transaction roll back cleanly on both sides, or does System A now believe the record synced while System B never received it. Is there a retry, and if so, does the retry create a duplicate. Is there an alert, and does it go to a person who is actually watching for it, or does it silently populate an error log nobody has looked at since the demo environment was built.

Why it works on smart people

Integration failures are, structurally, invisible until they aren’t. A sync that fails silently doesn’t announce itself. It just produces a slowly widening gap between what System A believes is true and what System B believes is true, and that gap is usually discovered by someone downstream, weeks or months later, reconciling numbers that don’t match and trying to figure out why.

Because the failure mode is invisible, “the demo showed it working” carries more weight than it should, simply because there’s no immediately visible counter-evidence in the room. A broken UI is obvious the moment you see it. A broken integration is obvious only in the reconciliation report nobody runs until month-end close, by which point the demo is a distant memory and the sales team has moved on to the next opportunity.

There’s also a vocabulary problem working in the vendor’s favor. “They just talk to each other” is a satisfying sentence, and it papers over an enormous amount of engineering that either exists, robustly, behind that sentence, or doesn’t exist yet and was built specifically to survive one scripted run.

The actual damage

This is the one that shows up as a reconciliation nightmare rather than a single dramatic failure. Two systems that were sold as integrated drift slowly apart in the weeks after go-live, each one silently correct according to its own records, disagreeing with the other in ways nobody notices until an audit, a customer complaint, or a finance close turns up numbers that don’t tie out. By then the question isn’t “does the integration work,” it’s “how long has it not been working, and what decisions got made on bad data in the meantime.”

The remediation is almost always more expensive than building the integration correctly the first time would have been, because now it includes both the engineering fix and a data cleanup project to reconcile however many weeks or months of silent drift accumulated before anyone noticed.

The fix, if you’re the one presenting

Show a failure on purpose. Send a record with a missing required field, or a duplicate key, and show what happens: does it error visibly, does it queue for retry, does someone get notified, does the other system stay in a known, correct state while the problem gets resolved. If the honest answer is “we haven’t built that handling yet,” say that, and say what the plan is. A prospect who sees a deliberate, controlled failure and a sane recovery path trusts the integration more than one who only ever saw the happy path, because they now know what happens on the day, and there will be a day, when the happy path isn’t what shows up.

Two systems that have never disagreed in front of you haven’t been integrated. They’ve been introduced.


This is exactly the failure mode a ledger-first architecture is built to make impossible. In Is Headless ERP Enough, or Just a Step in the Right Direction?, I walk through a prototype where two disconnected nodes post independent transactions and converge without conflicts, with no consensus protocol and no room for one system to quietly believe something the other doesn’t.

An Advanced Dungeons & Dynamics 365 session. Episode 3 of 14.

Previously on: Episode 2, The Shadow Ledger, the party discovered that Priya had been running a shadow spreadsheet for fourteen months, revealing a six figure variance nobody could explain. Priya left for vacation before anyone got an answer.

Every ERP has a chart of accounts. Not every ERP has a chart of accounts that actually makes sense. Sometimes what looks like a financial structure is really three different structures, built at three different points in the project’s history, quietly stapled together and never reconciled.

That’s the dungeon the party walks into this episode.


The Session

GM: You descend into the chart of accounts. It is, structurally, three different charts of accounts stapled together.

Sable: There are four segments for cost center. Four. Two of them mean the same thing in different departments.

Vex: (picking the lock on a legacy dimension table) Found something. A dimension called TEMP_DO_NOT_USE.

Thorne: How many postings?

Vex: Eleven thousand, two hundred and six.

Marge: Whose decision was that?

Vex: Nobody’s. That’s the problem. Somebody created it during UAT for testing and it just never got decommissioned. It’s load bearing now.

Ai. Cassiopeia: I flag that removing it will break forty three saved reports, none of which are documented.

Sable: We’re not removing it today. We’re mapping around it. Nobody touch the dimension.

A junior analyst walks past the war room door, glances in at the whiteboard now covered in a hand drawn diagram of overlapping cost center segments, and quietly closes the door again without saying anything. Nobody in the party blames him.

Thorne: (studying the diagram) This isn’t a chart of accounts. It’s an archaeology site.

Sable: Every layer’s a different consultant’s fingerprints. I can tell you which segment was added during which go live attempt just by how badly it fits with the one next to it.

Vex: (still elbow deep in the dimension table) There’s a comment field on this one. Someone left a note.

Marge: What’s it say?

Vex: “Will fix in phase two.”

Ai. Cassiopeia: I checked the project timeline. There was no phase two.


Why nobody decommissions the temp table

Every implementation accumulates a TEMP_DO_NOT_USE somewhere. It starts as exactly what the name says, a placeholder for testing, something built under deadline pressure that was always going to get cleaned up later. Then UAT ends, go live happens, and “later” quietly becomes never, because by the time anyone has bandwidth to revisit it, eleven thousand transactions are already posted against it and forty three reports are already built on top of it.

This is how technical debt actually accumulates in an ERP. Not through one bad decision, but through a hundred reasonable ones that never got revisited. Nobody sat down and decided Contoso should have four overlapping cost center segments. Somebody added one during the original build, somebody else added a second during a scope change nobody documented, and a third arrived during the failed second go live attempt when a different consultant didn’t know the first two existed.

The dangerous part isn’t that the structure is messy. Messy is survivable. The dangerous part is that nobody currently alive on the project understands the whole thing well enough to safely change it. That’s not a data model problem anymore. That’s an institutional memory problem, and it’s the same failure mode as the shadow ledger from episode two, just wearing a different disguise.

Sable’s instinct here is the right one: map around it before you touch it. You don’t fix an archaeology site by bulldozing it. You document what’s actually there first, then decide what’s safe to change.


What happens next

The party has a working map of the chart of accounts and a firm rule not to touch TEMP_DO_NOT_USE. But mapping the dimension structure surfaces the missing piece from episode one: the integration list from page seventeen of the SOW. In the next episode, the party finally gets their hands on it, and finds one integration on the list with no name, no owner, and no documentation at all.

Next episode: Episode 4, Page Seventeen (coming soon)


If your own chart of accounts has a TEMP_DO_NOT_USE hiding in it somewhere, you’re not alone, and you’re not imagining the risk. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a real framework for documenting and untangling structures like this one before they become load bearing, start here: adnd365.com/start

Cause of death: the case study was true, and that’s exactly the problem.


Two-thirds of the way through the deck, a new logo appears. A real one, a company you’ve heard of, sometimes a competitor’s supplier or a name from your own industry vertical. The slide has a number on it, usually a big one: forty percent reduction in close time, three million recovered in duplicate payments, six months to positive ROI. Underneath the number is a quote, attributed, sometimes even video, from a real person who really said those words.

Nothing on that slide is fabricated. That’s what makes this one the hardest autopsy in the series. The other demos in this blog die from omission, pacing, or seamlessness hiding a seam. This one dies from something subtler: a true statement about one company, presented in a context engineered to make you believe it’s a claim about yours.

What actually happened

The reference customer on the slide is not a random sample. It is, almost by definition, the single best outcome the vendor has produced across their entire installed base, selected specifically because the number is large and the customer is willing to say it out loud. Somewhere behind that slide are dozens or hundreds of other implementations that landed closer to the median, plus a smaller number that struggled or stalled, none of which get a logo or a quote, because nobody puts “we got most of the way to the business case, eventually, after two scope changes” on a slide.

There’s also a matching problem the case study never surfaces. The reference customer’s forty percent reduction in close time happened inside a specific starting condition: a particular level of process maturity, a particular data quality baseline, a particular willingness internally to change how work got done. The case study tells you the outcome. It almost never tells you the starting line, and the outcome without the starting line is not a number you can subtract your own situation from.

The quote does real work here too. A specific named person saying a specific thing on camera reads as harder evidence than an aggregate statistic, even though a single testimonial is a sample size of one, hand-selected from a population the vendor controls entirely.

Why it works on smart people

Humans are wired to trust specific, named, social proof more than abstract statistics, and this isn’t a flaw, it’s usually a reasonable heuristic. A named person willing to put their reputation behind a claim on camera is, in most contexts, more credible than an anonymous number. The problem is that the heuristic evolved for a world where the sample in front of you was roughly representative of the population, and a vendor-selected reference customer is the opposite of representative by construction.

There’s a second effect working alongside the first. By the time the reference slide appears, you’ve usually already sat through thirty or forty minutes of a demo that felt competent, so the case study isn’t landing on a skeptical audience, it’s landing on an audience that has already been primed to trust what they’re being shown. The reference customer isn’t doing the persuading alone. It’s the closing argument after the room has already been warmed up.

The actual damage

This is the one that turns into an internal expectations problem before it turns into a vendor problem. Someone in the room, often not maliciously, repeats the number in an internal steering committee deck as though it were a forecast rather than someone else’s outcome. “Similar companies have seen a forty percent reduction” quietly becomes “we’re targeting a forty percent reduction,” and by the time the project charter gets written, a single best-case data point from a different company, with different starting conditions, has become your project’s success criteria.

When your actual results land closer to the median, which is where most results land by definition, the project doesn’t get judged against a realistic baseline. It gets judged against the reference customer’s outcome, which nobody on your team ever should have agreed to as the target in the first place.

The fix, if you’re the one presenting

Show the range, not just the peak. If you have a reference customer at forty percent, say what the twenty-fifth and seventy-fifth percentile outcomes look like too, and say why the reference customer landed where they did, what was true about their starting point that might or might not be true about the prospect’s. A specific, named case study is still worth showing. It’s worth showing better, with its context attached, instead of as a number floating free of the conditions that produced it.

The honest version of that slide is less dramatic. It’s also the only version that survives contact with a steering committee eighteen months later.


This is the same shift I wrote about in The New Expert Isn’t the One With the Answers. Having the number was never the hard part. Knowing whether that number applies to your situation is.

An Advanced Dungeons & Dynamics 365 session. Episode 2 of 14.

Previously on: Episode 1, The Summons at Waterdeep Docks, the party arrived at Contoso Coffee Roasterie to find a missing scope document, a go live date that had quietly moved up by six weeks, and a whiteboard asking a question nobody could answer: where is the margin going.

Every implementation has an official system of record. And every implementation, if it’s been running long enough, has an unofficial one too. Usually a spreadsheet. Usually built by someone who got tired of waiting for the real numbers to be right.

That’s where episode two picks up.


The Session

GM: In the war room you meet three Contoso stakeholders: the CFO, the head of Ops, and a woman named Priya who nobody introduces by title.

Sable: (casting Data Mapping) I’m getting margin data. It’s coming from… this isn’t the ERP. This is a workbook.

Priya: That’s mine. I’ve been tracking real margin since the system doesn’t roll up freight variance correctly.

Marge: How long has this shadow ledger existed?

Priya: Since the first go live attempt. Fourteen months.

Thorne: (rolling a Perception check) The workbook total and the GL total are six figures apart.

Vex: Which one’s right?

Priya: I don’t know anymore. I leave for vacation tomorrow.

Ai. Cassiopeia: I would like it noted that I offered to reconcile this in October. The ticket was closed as “won’t fix, not urgent.”

The CFO leans over Thorne’s shoulder to look at the variance number. He goes pale in a way that has nothing to do with the lighting.

CFO: That number is bigger than our reported profit.

Nobody says anything for a moment. Priya starts quietly packing her laptop bag, the way someone packs when they’ve decided this is no longer their problem to solve, whether or not that’s actually true.

Priya: (standing) My flight’s at six tomorrow morning. I really am sorry.


Why the spreadsheet always wins

Every consultant who has done this work long enough has met a Priya. Somebody who didn’t set out to build a shadow system, they just needed the numbers to be right for a meeting, and the ERP wasn’t giving them right numbers fast enough. So they built a workaround. Then the workaround became a habit. Then the habit became fourteen months of institutional memory that lives entirely on one person’s laptop, reconciled by hand, understood by nobody else in the building.

This isn’t a story about a bad employee cutting corners. It’s what happens when a system doesn’t earn trust fast enough, and a business still has to make decisions in the meantime. Priya’s spreadsheet isn’t the failure. It’s evidence of a failure that happened somewhere upstream, probably long before she ever opened Excel.

The dangerous part isn’t that the spreadsheet exists. It’s that nobody in the room can currently say which number, the workbook or the GL, is actually correct. Six figures of ambiguity, and the person who understands the discrepancy best is getting on a plane in twelve hours.

That’s not a data problem anymore. That’s a continuity problem. And it’s the kind of gap that a proficiency system is supposed to catch long before it turns into a war room moment.


What happens next

Priya is gone by morning, and the shadow ledger goes with her, at least in terms of anyone who can explain it. The party has six figures of unexplained variance, a CFO who now can’t unsee the number, and a chart of accounts they haven’t even looked at yet. In the next episode, they go looking for where the discrepancy actually lives, and find a financial dimension that’s been quietly absorbing thousands of postings that were never supposed to exist.

Next episode: Episode 3, The Labyrinth of Financial Dimensions (coming soon)


A quick gut check for anyone running their own ERP: is there a shadow spreadsheet keeping your business honest right now? Gamifying the Enterprise: Game Mechanics for Continuous Proficiency digs into exactly why that happens, and how to design proficiency systems so it stops. Available on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want the actual configuration playbook behind scenes like this one, start here: adnd365.com/start