Archive

Tag Archives: erp

Cause of death: confidence borrowed from someone who was never actually going to be there.


Ten minutes into the meeting, or sometimes ten minutes before the end of it, a calendar notification pulls someone senior into the call. A VP, sometimes higher. They say a few warm, general things about strategic partnership and long-term vision, take one or two softball questions, and leave for their next meeting. The room’s posture changes almost immediately. If someone at that level is personally invested enough to show up, even briefly, this must really matter to the vendor. The deal, whatever doubts existed five minutes earlier, now feels more serious.

Nobody asks what “personally invested” is actually going to mean six months from now, when the invoices are contested, the timeline slips, or the implementation team needs an escalation path that goes higher than the account manager. That’s the autopsy. The executive’s presence proved that a calendar invite got accepted. It proved nothing about what happens after the meeting ends.

What actually happened

An executive sponsor cameo is a highly efficient transfer of credibility from a person to a deal, and the transfer costs the vendor almost nothing to make. Ten minutes of a VP’s time is genuinely cheap relative to the deal size being discussed, and the effect on the room is entirely disproportionate to that cost, because the audience reads presence as commitment. It rarely is. In most organizations selling anything of this size, senior leaders make these appearances routinely, across many simultaneous deals, as a normal part of their job, not as a signal that this particular account has been elevated to a special tier of personal attention.

The deeper issue is that the thing actually being evaluated, whether the vendor will show up when the implementation gets hard, is not a property of any one person’s goodwill. It’s a property of organizational structures: escalation paths, contractual service levels, account team continuity, whether the people doing the actual implementation work have the authority and resources to fix problems without waiting on approval from someone three levels up. None of that gets tested by a cameo. A cameo tests whether an executive’s assistant could find a ten-minute gap in a calendar.

There’s also a durability problem the cameo doesn’t address. The VP who dropped in with warm words about the partnership may be gone, reorganized, or reassigned to a different portfolio before the implementation is even a third of the way done, which happens routinely in any organization above a certain size. The relationship the room felt reassured by was never actually contracted. It was a mood, generated in a room, that has no mechanism for surviving contact with an org chart six months later.

Why it works on smart people

Status carries information in most human interactions, and that heuristic is usually reasonable: when someone senior spends scarce time on something, it often does signal real priority. The problem is that the heuristic breaks down specifically in situations, like enterprise sales, where the cost of the senior person’s time has been deliberately minimized to make the signal cheap to send. A genuine ten-minute cameo and a fully commissioned, resourced executive sponsorship look identical for the ten minutes you can observe them. They diverge entirely in the six months you can’t.

There’s also a reciprocity dynamic at play. A senior person taking time to personally reassure you creates a mild social obligation to receive that reassurance graciously, not to interrogate it. Asking a VP who just delivered warm remarks about partnership to specify exactly what escalation authority they’re personally committing to feels confrontational in a way that asking the account manager the same question doesn’t, so the question quietly goes unasked at exactly the moment it would have been most useful to ask it.

The actual damage

This is the one that surfaces the first time something actually goes wrong during implementation and the buyer tries to use the relationship they thought they’d built. The champion emails the VP directly, the way the cameo implicitly invited them to, and gets a response from an assistant, or a redirect back to the account team, or silence, because the VP’s actual involvement was never structured to include personal escalation on operational issues. The confidence the room felt in that meeting has no contractual or organizational form. It was real in the room and evaporated the moment it needed to be load-bearing.

The buyer is left in a worse position than if the executive had never appeared at all, because the cameo specifically substituted for asking the harder, more useful questions about actual escalation paths and account team continuity, the answers to which would have held up regardless of who was in what job six months later.

The fix, if you’re the one presenting, or the one buying

If you’re presenting, don’t let the cameo stand in for structure. If an executive is genuinely sponsoring the account, say specifically what that means: a defined escalation path with their name attached, a commitment to a quarterly check-in that’s on a calendar rather than implied, actual authority to authorize resources if the implementation hits trouble. If none of that exists, the honest version of the cameo is shorter and less dramatic, a courtesy visit rather than a commitment, and it should be presented as exactly that.

If you’re buying, ask the question the warm remarks were designed to make feel unnecessary: what happens, specifically, and who do we call, when this goes wrong. The answer that matters is a name and a process that survives a reorg. The ten minutes in the room, however sincere, was never that.


Genuine organizational attention rarely looks like a scheduled cameo. In The Penguins Knew Before Your Steering Committee Did, the character who actually notices the iceberg melting and does something about it is an unremarkable penguin, not the colony’s formal leadership showing up to reassure everyone for ten minutes. The steering committee structure is the cameo. The attention that actually matters usually comes from somewhere quieter.

I have heard rumours that some companies are starting to vibe code their own ERP systems. Is this the start of the decline of standardized ERP systems, and the start of company specific Business Operating Systems? If every business had its own OS, and an agentic framework, why is there any need to have standardized ERP systems?

It is a fair question, and one worth taking seriously instead of dismissing on reflex. But after digging into what is actually happening in the market, I do not think ERP is dying. I think something more interesting and more useful is happening at its edges.

What Companies Are Actually Vibe Coding

The companies experimenting with AI generated business software right now are almost entirely building at the periphery of ERP, not in its core. A Copilot Cowork style agent that reconciles a report. A custom approval workflow. A data cleanup script. A bolt on dashboard that nobody wanted to wait six weeks for IT to build. Very few organizations are vibe coding their general ledger, their costing engine, or their multi entity intercompany eliminations, and there is a good reason for that.

There is a real gap between “AI helped me build software” and “AI helped me build enterprise grade, auditable, multi user, transactionally consistent financial software.” That gap is where most of the current enthusiasm quietly runs out of road. One recent piece on the risks of vibe coding put it plainly: the danger is not that these tools fail to work. The software appears to work right up until the business reaches a level of complexity that exposes what was never built in the first place. Security, monitoring, compliance, and disaster recovery do not show up in a demo. They show up eighteen months later when something breaks and there is no vendor to call.

We Have Been Here Before

There is also a useful historical echo worth remembering. Standardized platforms like SAP, Oracle, PeopleSoft, and later NetSuite and Salesforce did not emerge in a vacuum. They emerged because “build your own” had already been tried extensively in the eighties and nineties, and the industry collectively decided that a professionally maintained, continuously improved, shared system was the better bargain. One system, supported by a vendor with thousands of customers absorbing the cost of getting tax law and audit controls right, beat a thousand bespoke systems each carrying that burden alone.

Vibe coding does not remove the reasons that decision was made three decades ago. It just lowers the upfront cost of relearning them the hard way.

Two Layers That Want Different Economics

The question of “why standardize if every business can have its own OS” collapses two layers of the business that actually want very different things.

The first layer is the system of record. This is the ledger, inventory, costing, tax, and consolidation logic. This layer wants low variance, not customization. Its value comes from being provably correct under audit, from producing the same math every quarter regardless of who is running the company, and from a vendor carrying the liability of getting VAT rules or SOX controls right across thousands of customers instead of one company carrying that risk alone. Agentic capability does not change the incentive to pool that risk and cost across an ecosystem. If anything it raises the stakes, because an agent posting a bad journal entry at two in the morning with no human review is a considerably scarier proposition than a person doing it.

The second layer is orchestration. This is how a specific company actually gets things done around that system of record. This is exactly where variance is valuable, and exactly where agentic frameworks and vibe coding are already winning, because it is low blast radius, fast changing, and does not need to survive a financial audit on its own merits.

The Composable Core, Not the Bespoke Empire

So the more likely outcome is not that every business ends up building its own ERP from scratch. It is something closer to a composable core: a thin, standardized, vendor maintained transactional spine, surrounded by a thick, company specific, largely agent authored layer of workflow, automation, and decision logic that talks to that spine through APIs. This is not a new architecture invented for the agentic era. It is basically the direction that platforms like Dynamics 365 Finance and Operations were already heading, with data entities and business events designed precisely so that logic could live outside the core without touching it.

The Business Operating System that gets discussed in these conversations is real. But it gets built on top of standardized ERP, not instead of it, because nobody actually wants to own tax engine liability, and no CFO wants to explain to auditors that the general ledger was vibe coded last quarter.

Where the Real Disruption Sits

If I had to point at where this trend genuinely threatens the standard ERP model, it would be three places.

The mid market and small business segment, where compliance and audit stakes are lower and the cost of a full ERP license and implementation is disproportionate to the size of the operation. This is where skipping a Tier 2 ERP entirely in favour of an agentic build of the operations layer is most economically rational.

Verticals with no good off the shelf fit, where companies already customize ERP into something barely recognizable compared to the base product. Vibe coding just makes that customization faster and cheaper. It is not a fundamentally new behaviour, it is an acceleration of an old one.

And the vendors themselves, who will not sit still and watch this happen. Expect Microsoft, SAP, and Oracle to lean hard into a “we are the trusted core, bring your own agents to the edges” positioning rather than fight the trend directly. Given where Copilot and Cowork are already headed inside the Microsoft ecosystem, this is not speculation. It is already the direction of travel.

The Verdict

Vibe coding is not the death of standardized ERP. It is the arrival of a cheap, fast way to build the layer of the business that ERP vendors were never well suited to build in the first place. The ledger stays boring, standardized, and vendor owned, because boring is exactly what you want from something an auditor has to sign off on. Everything wrapped around it just got a lot more interesting.

John Kotter wrote Our Iceberg Is Melting as a fable for people who already knew his eight step change model and needed a way to explain it to everyone else without putting them to sleep. A colony of Antarctic penguins lives on an iceberg. One of them, an unremarkable penguin named Fred, notices the iceberg is riddled with cracks and will not survive the winter. Nobody wants to hear it. The leadership council is comfortable. The colony has always lived on this iceberg. Fred is not important enough to be believed.

The book is aimed at children and executives in roughly equal measure, which tells you something about how leaders actually respond to bad news about the platform they built their careers on.

I have spent enough years around D365 F&O implementations to recognize every penguin in that story. Most ERP failures are not technical failures. The system usually works. What fails is the eight step sequence Kotter is illustrating, and it fails in the same order every time.

Step One: Someone Has to Notice the Cracks

Fred does not manufacture urgency. He finds it, through unglamorous observation, and then has to fight to get anyone to look at what he found. In an ERP program this is usually the business analyst three levels down who has actually mapped the current state process and knows the general ledger reconciliation takes eleven days by hand. Nobody asked her opinion because the steering committee already decided the go-live date.

Urgency that is manufactured by a consultant in a kickoff deck does not survive contact with month four. Urgency that comes from someone inside the organization who found the actual crack tends to.

Step Two through Four: The Guiding Team Is Not the Org Chart

Louis, the head penguin, does not solve the problem alone, and he does not delegate it to a committee chosen by rank. He assembles a mixed group that includes NoNo the professional skeptic, because a guiding coalition that only contains believers cannot survive contact with the rest of the colony. NoNo’s objections get answered in advance instead of ambushing the project in month six.

I have sat in enough steering committees to know most of them are built from an org chart, not from who actually has credibility on the floor. A finance director with veto power and zero trust from the warehouse team is not a guiding coalition. He is a bottleneck with a title.

Steps Five through Seven: Short Wins or No Buy In

The colony does not wait for the whole migration plan before it sees a win. The penguins send scouts, they find a new iceberg, and they celebrate finding it long before anyone has actually moved. This is the part of the model most ERP programs skip entirely. Everything gets held for the big bang go-live, and by the time it arrives the organization has spent a year hearing about a future state it has never once been allowed to touch.

Every fit-gap workshop I have run that included a working demo of one solved process, even a small one, bought more goodwill than any status report ever has. People do not commit to a vision. They commit to evidence the vision is real.

Step Eight: The New Iceberg Has to Become Normal

The fable ends with the colony living on the new iceberg as though they had always lived there, and this is the step almost nobody in ERP world budgets for. Go-live is treated as the finish line. It is closer to the halfway mark. The old habits, the shadow spreadsheets, the workaround someone built in Excel because the new system was inconvenient in week two, all of that quietly recolonizes the organization unless someone is watching for it.

I have seen systems go live successfully and then watch the business slide back into its old processes within a fiscal quarter because nobody treated adoption as an ongoing job. The iceberg was new. The habits were not.

Why the Fable Works Better Than the Framework

Kotter could have written the eight steps as a slide deck. He wrote a story instead, because a story lets you see NoNo’s objection land and get resolved instead of reading a bullet that says manage resistance. It lets you watch Louis choose the guiding team instead of reading a bullet that says build a coalition. The mechanism becomes visible instead of asserted.

That is the actual argument for using fables and analogies in change management work generally, ERP or otherwise. An eight step framework tells people what to do. A story shows them what it looks like when someone does not do it, which is usually the more persuasive lesson.

If your organization is mid migration and something feels off, it is worth asking which penguin is missing. Often it is not a step in the model that got skipped. It is a specific person, the one who noticed the crack early and was not believed.

Our Iceberg Is Melting by John Kotter and Holger Rathgeber: https://www.amazon.com/Our-Iceberg-Melting-Succeeding-Conditions/dp/0399563911

Cause of death: the entire session proved the system worked, and none of it proved anyone would use it.


Ninety minutes of demo. Every module clicked through, every workflow shown, every objection about functionality answered on the spot. Then, in the last five minutes, someone asks about training and adoption, and the presenter says something reassuring and short: “the interface is intuitive, users pick it up quickly,” maybe gestures at a slide with a generic icon of people around a laptop, and the meeting ends on schedule.

Nothing in those ninety minutes tested the only variable that determines whether an ERP implementation actually pays off: whether the several hundred people who currently do their jobs a certain way will actually do them a different way starting on a specific Monday morning. That variable got five minutes and a slide. Everything else got ninety.

What actually happened

A demo is, by construction, a test of the system in isolation. It shows what the software can do when operated by someone who already knows exactly which buttons to press, in exactly the right order, with no muscle memory pulling them back toward the old way of doing things. That’s a test of the product. It says almost nothing about the much harder problem sitting underneath every ERP rollout, which is that the system doesn’t fail because it can’t do the work. It fails, when it fails, because the people who were supposed to start using it on day one didn’t, or did so inconsistently, or found a workaround that quietly recreated the old process inside the new tool.

The demo cannot show this risk because the risk doesn’t live in the software. It lives in a warehouse supervisor who has run the same process for eleven years and has a system that works well enough for them, a controller who doesn’t trust the new close process until they’ve watched it succeed three months running, a data entry team that will, absent enforcement, keep using the spreadsheet they built in 2019 because it’s faster for them personally even if it creates downstream problems for everyone else. None of that shows up in a screen share. All of it shows up in the first ninety days of production.

Why it works on smart people

Functionality is legible in a way that adoption risk isn’t. You can watch a screen and evaluate, with reasonable confidence, whether a feature does what it claims to do. You cannot watch a screen and evaluate whether the accounts payable team, specifically, at your company, specifically, will actually change how they process an invoice exception, because that isn’t a property of the software being demonstrated. It’s a property of an organization that isn’t in the room.

Because functionality is the thing that’s easy to evaluate in real time, it naturally consumes the meeting. Evaluators default to spending their scrutiny where scrutiny is possible to apply, and adoption risk, being diffuse, organizational, and only observable months later, gets the leftover attention at the end of the agenda, not because anyone decided it mattered less, but because there was no way to spend ninety minutes evaluating it the way there was for the features.

There’s also a comforting assumption doing a lot of quiet work: that a system good enough to buy will be a system people are willing to use. That assumption is often wrong in a specific, predictable direction. The features that make a system good for the business, tighter controls, more required fields, more visible audit trails, are frequently the exact features that make individual users’ jobs feel harder in the short term, which is precisely the friction that produces workarounds and shadow processes.

The actual damage

This is the failure mode that shows up as an implementation that technically went live and never actually delivered its business case. The system works. The training happened, in the generic sense that sessions were held and attendance was tracked. And six months later, half the intended users are still keeping a parallel spreadsheet, a workaround team has formed around one specific process nobody bothered to redesign for the new tool, and the data quality problems the new system was supposed to fix are still there, because garbage still goes in whenever someone routes around the system instead of through it.

Nobody can point to a single moment this failed. It failed gradually, in a thousand small individual decisions to keep doing things the old way, none of which showed up as a defect, an error, or a support ticket, because from the system’s point of view nothing went wrong. The people just didn’t come.

The fix, if you’re the one presenting, or the one buying

Treat adoption as a first-class item in the evaluation, not a closing slide. Ask the vendor, specifically, what happens to the specific roles in your organization who will feel the most friction from the new process, not the roles who benefit most. Ask what the actual training plan looks like beyond a generic session count, and who owns reinforcement after go-live, when the temptation to slide back into the old process is highest. If you’re the one presenting, bring a real adoption story, with a real friction point that was anticipated and addressed, instead of a slide with a stock photo and the word “intuitive.”

A system that works and a system that gets used are not the same claim, and only one of them was tested in that room.


This is exactly the gap I dug into in Gamification of ERP: Turning Drudgery into Dopamine. Points and badges won’t fix a bad rollout, but the underlying problem, that a system’s success depends on individual motivation, not just individual capability, is precisely the variable this autopsy argues nobody tests before go-live.

Cause of death: a button that already worked got wrapped in a chat box and called intelligent.


The feature existed before the demo did. Somewhere in the product, there was a dropdown, a rule, a scheduled job, something deterministic that took an input and produced a correct, predictable output every time. Then, at some point in the last two product cycles, that same feature got a new front door: a text box, a little sparkle icon, a placeholder that says “ask me anything.” Now, instead of picking a value from a dropdown, you type a sentence, wait a beat, and the same output appears. The presenter calls this AI. The room, primed by two years of hearing the word everywhere, nods.

Nobody asks the obvious question: was this better before, and did anyone check.

What actually happened

Somewhere inside a lot of “AI-powered” features sits a deterministic operation that a rule, a formula, or a simple lookup already handled correctly and quickly. Wrapping that operation in a natural-language interface doesn’t make the underlying logic smarter. It adds a translation layer, one that has to interpret an unstructured sentence and map it back onto the same structured operation the dropdown was already doing directly, with total accuracy, in a fraction of the time.

That translation layer isn’t free. It introduces a new failure mode that didn’t exist before: the AI misreading the sentence and selecting the wrong option, a phrasing the model hasn’t seen before returning an unhelpful answer, latency where there used to be an instant response. None of this is inherent to AI as a category. Plenty of genuinely AI-native capabilities do things no dropdown ever could: summarizing an unstructured document, drafting a first pass at something that has no single correct answer, flagging an anomaly buried in a pattern too complex for a fixed rule to catch. The problem in this specific demo isn’t that AI was used. It’s that AI was used to solve a problem that was already solved, by something more reliable, and the swap happened anyway because AI is what gets funded, marketed, and put on a keynote slide this year.

The tell is almost always the same: watch what happens when you ask the natural-language version to do something slightly outside the phrasing it was tuned on. The dropdown never had this problem, because a dropdown has no phrasing to be tuned on. It just has options.

Why it works on smart people

Nobody wants to be the person in the room who seems skeptical of AI in a year when every vendor, every board deck, and every competitor’s marketing has decided AI adoption is the metric that matters. Asking “why does this need to be a chat interface instead of the three-click menu it replaced” risks sounding like you’re behind, not like you’re asking a reasonable engineering question, and that social pressure runs in exactly the wrong direction. It rewards accepting the AI wrapper uncritically and punishes the person who’d actually stop to check whether it improved anything.

There’s a second, quieter dynamic on the vendor side that compounds this. Once a company’s roadmap and marketing commit to being an “AI-first” platform, there’s organizational pressure to retrofit AI onto existing features whether or not doing so improves them, because a product review, an investor update, or a competitive comparison chart wants to see the AI checkbox filled in across the board, not filled in only where it was actually the right tool.

The actual damage

This is the one that costs you reliability you already had. A feature that used to return the same correct answer every time now returns a slightly different answer depending on phrasing, and the variance itself becomes a support burden, because now every unexpected result needs to be triaged as either a real bug or a model doing something technically defensible but unhelpful. Power users, the ones who had the old dropdown workflow memorized and could execute it in seconds, are now slower, because the natural-language version, for all its friendliness, requires more typing and more waiting than the three clicks it replaced.

There’s a subtler cost too. Every AI feature added for marketing reasons rather than capability reasons dilutes the credibility of the AI features that are actually doing something new and valuable. Once a buyer has been burned by one chat-wrapped dropdown, they bring that skepticism to the next AI claim in the deck, including the one that might have genuinely deserved their trust.

The fix, if you’re the one presenting

Ask, honestly, before the feature ships: does the natural-language interface let the user do something the deterministic version couldn’t, or is it doing the same operation with more ambiguity and more latency. If the honest answer is that the underlying logic hasn’t changed, keep the dropdown, or offer both, and don’t spend the marketing budget claiming AI where the actual improvement is zero or negative. If the AI genuinely does something new, lead with that specifically, in concrete terms, instead of leaning on the word “AI” to do the persuading by itself.

The word was never the feature. The capability was, and a capability that already existed doesn’t become new because it learned to accept a sentence instead of a click.


The distinction here matters enough that I built around it. In Is Headless ERP Enough, or Just a Step in the Right Direction?, the “AI as the sole interface” section argues for AI genuinely load-bearing in the architecture, not a chat box glued in front of the same dropdown. That’s the version of AI worth demoing. Everything else in this autopsy is what it looks like when a team ships the coat of paint instead.

Cause of death: the box got checked without anyone asking what checking it actually meant.


Security comes up in most enterprise demos as a single slide, usually near the end, usually delivered fast. Role-based access control. Single sign-on. Field-level security. Audit logging. SOC 2. Encryption at rest and in transit. Each term gets a checkmark, a confident nod from the presenter, and about four seconds of screen time before the deck moves on to something more visually interesting.

Nobody in the room stops the slide. That’s the autopsy. Every term on that slide is real, in the sense that the feature genuinely exists somewhere in the product. What’s missing is any demonstration that it does what the room assumes it does, configured the way your company would actually need it configured, at the depth your actual risk profile requires.

What actually happened

“Role-based access control” is not one feature. It’s a spectrum that runs from a handful of fixed roles with no customization, through role-based permissions you can tailor at the menu-item level, through field-level and record-level security that can restrict a single sensitive column or a single customer’s data from a specific user, through fully dynamic, condition-based access rules that change what someone can see based on context. A vendor can put “role-based access control” on a slide truthfully at any point on that spectrum, and the room has no way of knowing, from the slide alone, which end they’re getting.

The same collapse happens to every other term on the list. “Audit logging” might mean every field change on every table is captured with before and after values and an immutable timestamp, or it might mean a handful of high-level events get logged with no field-level detail. “SOC 2” tells you an audit happened and a report exists. It doesn’t tell you which trust service criteria were in scope, what the exceptions were, or whether the report is even current. “Encryption at rest” is true of nearly every modern cloud platform and tells you almost nothing about key management, who holds the keys, or what happens in a subpoena scenario.

None of this is the vendor lying. Every term is technically accurate. The compression happens in the translation from a nuanced, configurable capability into a single reassuring word on a slide, and the room does the rest of the work by filling in the most generous plausible interpretation of what that word means.

Why it works on smart people

Security and compliance are exactly the kind of topic where nobody in a sales meeting wants to be the person who asks a question that reveals they don’t fully understand the acronym. SOC 2 Type I versus Type II, the difference between authentication and authorization, what “field-level security” actually restricts versus what it merely hides in the UI while leaving accessible through an API, these are legitimate technical distinctions that most people in the room, including some people whose job title suggests they should know them, have only a fuzzy grasp of.

The slide exploits that fuzziness efficiently, not through malice but through pace. There’s no room in a four-second checkmark to unpack what a term actually covers, and stopping to ask “which SOC 2 trust service criteria were in scope” in the middle of a fast-moving demo carries a social cost that quietly discourages the question from being asked, the same social cost that lets the AI Magic Demo’s chat window go unexamined.

The actual damage

This is the one that surfaces during an actual security review, an actual audit, or worse, an actual incident, months or years after the contract was signed. Someone in security or compliance, doing real diligence for the first time, discovers that “field-level security” in this product means the field is hidden in the standard UI but fully readable through the API with the right permission, which is a meaningfully different security posture than what was assumed at signing. Or the SOC 2 report, once actually read line by line, turns out to have scoped out exactly the subsystem your data lives in.

At that point the conversation isn’t a demo follow-up. It’s a risk finding, sometimes one that has to be reported up to a board or disclosed to a regulator, and remediating it after the fact, whether that means additional configuration, a compensating control, or a vendor conversation about contractual commitments, is far more expensive and far more visible than it would have been to ask the specific question up front.

The fix, if you’re the one presenting

Don’t let the slide stand alone. For each term, say what it actually covers and, just as importantly, what it doesn’t. “Field-level security restricts visibility in the standard interface. It does not currently restrict API access to the same field, here’s how customers typically compensate for that.” “Our SOC 2 report covers these three trust service criteria, and here’s the exception list from the most recent period.” That level of specificity costs a few extra minutes and makes the product sound less flawless. It also means the prospect’s actual security review, whenever it happens, confirms what they were told instead of contradicting it.

A checkmark is not a control. It’s a promise that a control exists, and promises deserve exactly as much scrutiny as anything else on the slide.


The same collapse happens to job titles, not just checklist items. In AI Makers Don’t Build Models. They Build Value., a single contested word, “maker,” was doing more work than it could support, and the gap got filled by whoever was listening. A security slide runs on the exact same mechanism, one word standing in for a spectrum, and the room filling in whichever end feels most reassuring.

Cause of death: ten records don’t behave like ten million, and nobody in the room was thinking in millions.


The presenter clicks a button. A report renders instantly. A search returns results before the loading spinner has time to spin. A batch job that would, in your world, run overnight, completes in the time it takes to say “and here’s the output.” The room absorbs all of this the way it absorbs everything else in a demo: as evidence of how the product performs.

It is not evidence of how the product performs. It is evidence of how the product performs against a database with four hundred customers, running on hardware provisioned for a demo, with exactly one user logged in, who is the only person generating any load at all. Your environment, on day one of production, will have none of those three things in common with it.

What actually happened

Performance is not a property of software in isolation. It’s a property of software under a specific load, against a specific dataset size, with a specific number of concurrent users doing specific things at the same time. A demo environment is engineered, whether deliberately or just by the natural economics of running a sales demo, to minimize all three of these variables simultaneously. Small dataset, dedicated hardware, single user. That is close to the best-case condition the software will ever run under, and it’s the condition you’re being shown as though it were representative.

The gap between that best case and your actual production reality is usually invisible in the room because nothing about the demo interface tells you the dataset is small. A grid with four hundred rows and a grid with four million rows look identical on screen, for the same reason they looked identical in the migration autopsy: you’re only ever looking at the same twenty rows at a time. Query performance, index behavior, and lock contention all degrade in ways that simply don’t exist yet at four hundred rows, and none of that shows up until the dataset, and the concurrent user count, both climb to something resembling your real operation.

Batch and integration jobs hide the same gap differently. A nightly job that processes four hundred records in three seconds tells you almost nothing about how it will behave processing four hundred thousand records at 2 a.m. while five other scheduled jobs are also competing for the same database connections. The three-second version and the six-hour version can be, technically, the exact same code.

Why it works on smart people

Performance is one of the few things a demo can show without narrating, which makes it feel more objective than the parts that require a presenter’s framing. Nobody has to make a claim about speed. You just watch it happen, in real time, and speed you watch with your own eyes reads as harder evidence than speed someone tells you about. That instinct is usually correct. It’s just being applied to a measurement taken under conditions that will never recur once the system goes live.

There’s also a scale-intuition gap that’s genuinely hard to close without direct experience. Most people don’t have a strong internal sense of how nonlinearly performance can degrade as data volume and concurrency grow. A system that feels instant at four hundred records doesn’t necessarily feel merely a little slower at four million. Depending on how indexes, queries, and locking are built, it can fall off a cliff at some threshold nobody in the demo room has any way of anticipating.

The actual damage

This is the one that shows up as a go-live incident rather than a slow discovery, because performance problems under real load tend to announce themselves immediately and all at once, usually during month-end close or the first day the full user base logs in simultaneously. Reports that took two seconds in the demo take four minutes. A batch process that ran in three seconds against sample data doesn’t finish before the next scheduled job needs the same resources, and the two start colliding every night.

The remediation at that point is expensive and disruptive in a way that early testing would not have been: emergency performance tuning, index rebuilding, sometimes an infrastructure upgrade that wasn’t budgeted, all happening under the worst possible conditions, with live users blocked and a go-live date already spent.

The fix, if you’re the one presenting

Test and show performance against something that resembles your prospect’s actual scale, not the vendor’s default demo dataset. If a true load test isn’t feasible in the sales cycle, at minimum say so explicitly: “this demo is running against four hundred sample records on dedicated hardware, here’s what we know about performance at your expected volume, and here’s how we’d validate it before go-live.” That sentence costs you the illusion of effortlessness. It buys you a prospect who understands what they’re actually being shown, and a performance conversation that happens in scoping instead of in a production incident.

Speed you watched with your own eyes is still only evidence of the conditions you watched it under. Ten records were never going to tell you what ten million would do.


I’ve argued in The Same Four Systems that the same organizational patterns show up whether you’re a corner store or a Fortune 500 company, just with higher stakes. That holds for structure. It doesn’t hold for performance. A query pattern that’s invisible at four hundred rows can become the whole story at four million, and no amount of pattern-matching from a small scale prepares you for exactly where that threshold sits.

Cause of death: nobody could agree on what the proof of concept was supposed to prove.


This one starts differently than the others. Every autopsy so far has been performed on a demo the vendor built to look better than reality. This one is performed on a demo the customer built, with the vendor’s help, to answer everything at once, and in doing so answered nothing.

The pattern is familiar to anyone who has scoped a proof of concept. It starts small and correct. One core question, one hypothesis, one thing that either works or doesn’t: can this system handle our multi-entity intercompany billing without a workaround, yes or no. Then someone from a different department hears there’s a POC happening and asks if it can also touch their process, since they’re curious too. Then someone senior asks for the trickiest edge case in the business to be included, because if it can’t handle that, what’s the point. By the time the scoping document is final, the POC that was supposed to answer one question is now attempting to demonstrate manufacturing, procurement, three approval hierarchies, a currency conversion edge case that occurs twice a year, and an integration to a system that isn’t even part of the actual project scope.

Nobody added any single piece of this in bad faith. That’s what makes it hard to stop once it’s moving.

What actually happened

A proof of concept exists to reduce uncertainty about one specific, high-risk question as cheaply and quickly as possible. The moment it starts trying to prove ten things instead of one, several things happen at once, all of them bad.

The build time stops being proportional to the risk being retired. A POC that answers one hard question can often be built in days, because the vendor and the team can focus every hour on the thing that actually matters. A POC trying to demonstrate ten things needs ten times the configuration, ten times the test data, ten times the edge cases handled, and none of that additional effort is reducing risk proportionally, because nine of the ten things were never actually in doubt.

The result also stops being interpretable. If the multi-entity billing scenario fails, but it failed inside a build that also included four other complex configurations layered on top of each other, you cannot cleanly attribute the failure. Was it the core capability that doesn’t exist, or a configuration conflict between two features that were never meant to be tested together, or a data setup error introduced trying to support scenario six while building scenario three. A focused POC gives you a clean signal. An overloaded one gives you noise that looks like a signal.

Why it happens to smart teams

The instinct behind scope creep in a POC is almost always defensible in isolation. Nobody wants to greenlight a six or seven figure implementation based on a narrow test, only to discover eight months in that some other critical process doesn’t fit either. The fear isn’t irrational. Systems do have gaps that only show up once you look in the right corner, and a POC feels like the cheapest moment to go looking.

The trouble is that “the cheapest moment to look” and “the cheapest way to look” are different questions. Looking broadly at low depth, a checklist of yes-or-no capability questions answered through documentation review, a reference call, or a scoped demo of specific features, retires broad risk cheaply. A single POC trying to go deep on ten things at once is neither cheap nor deep. It’s the expensive way to get a shallow answer to a question that didn’t need a POC to answer in the first place.

There’s also a political dimension that’s hard to name out loud in the room. Once word gets out that a POC is happening, being excluded from it can read as a signal that your department’s concerns don’t matter. Scope grows partly because saying no to an additional scenario feels like saying no to a person, not to a line item.

The actual damage

The POC that was supposed to take two weeks takes eight. The core question, the one thing that actually justified spending POC time and budget, gets buried under nine other questions that each needed their own edge case handling, and by the time results come back, the steering committee is looking at a partial success across ten dimensions instead of a clear answer on the one dimension that mattered. Decision paralysis follows almost automatically, because a partial, ambiguous result is much harder to act on than a clean pass or fail.

Worse, the actual high-risk question, the reason the POC existed in the first place, often gets the least rigorous testing of the ten, because it was scoped first and then diluted by everything added after it. The team spends real effort proving things that were never seriously in doubt, and comparatively little effort on the one thing that was.

The fix, if this is your situation

Separate what the POC needs to prove from what people merely want to see. Write down the single question, or at most two, whose answer would actually change the buying decision. Everything else, every “while we’re in there” request, goes on a second list explicitly labeled as out of scope for this exercise, with a stated plan for how it will get answered instead, whether that’s a reference call, a documentation review, or a second, later POC once the first question is settled.

When someone pushes back and asks why their scenario isn’t included, the honest answer is the useful one: this POC is designed to answer one hard question as cleanly as possible, and adding your scenario wouldn’t make the answer more trustworthy, it would make it harder to read. That’s not a dismissal of their concern. It’s a commitment to answering it properly, later, instead of poorly, now, buried inside somebody else’s test.

A POC that tries to prove everything proves nothing cleanly. The discipline isn’t in the build. It’s in what you refuse to put in it.


The same fragmentation shows up in individual work, not just project scoping. In The Forty One Percent Problem, I look at decades of research showing that a large, stable share of professional time gets eaten by low-judgment overhead scattered across too many things at once. A POC that tries to answer ten questions has the same disease as a workday that tries to touch ten priorities. Depth loses to breadth every time.

Cause of death: three years of production data quietly became four hundred clean sample records for the day of the pitch.


Somewhere in the middle of the demo, the presenter opens a grid. Customers, items, transactions, whatever the domain calls for. It scrolls smoothly. Every row has every field populated. Names are properly capitalized. Addresses have all their parts. There are no duplicate customer records for “Acme Corp,” “ACME Corp,” and “Acme Corp.” with a trailing space that your actual system has accumulated over a decade of different people typing the same name slightly differently.

The presenter doesn’t say “this is sample data.” They don’t need to. The grid looks so much like a real company’s data that the distinction quietly stops mattering to the room, and everyone leaves the meeting having watched a migration that never happened, of data that was never really yours.

What actually happened

Four hundred rows of clean, plausible-looking data is not a migration. It’s a mockup wearing a migration’s clothes. Somebody built that dataset specifically to demonstrate the target system’s data model, which means it was constructed backward from what the target system wants to receive, rather than forward from what your source system actually contains. It has never been through a real extract. It has never hit a field length limit, a character encoding mismatch, a required field that’s been null in your source system since 2019 because nobody enforced it, or a foreign key that points to a parent record that got deleted three reorganizations ago.

Real migrations die on exactly these details, and none of them are visible in a four hundred row demo grid, because the demo grid was never subjected to the process that would surface them. The sample data is a hypothesis about what your data looks like. It has not yet met your data.

There’s a second layer under this. Even when a vendor does an actual proof of concept against a real extract of your data, that extract is usually a snapshot, cleaned once, run through a mapping exercise once, and shown once. It demonstrates that a migration is possible for that slice, on that day, with that much attention paid to it. It does not demonstrate that the full historical dataset, run through the same process without the benefit of a team hand-tuning exceptions in real time, will produce the same result.

Why it works on smart people

Data problems are boring in a way that makes them easy to underestimate from the outside. Nobody gets excited describing thirty thousand customer records with inconsistent capitalization, or a decade of transactions where the currency field was optional for the first four years, and that lack of drama works against the diligence the problem deserves. A demo that skips the data reality skips the part of the story that was never going to be compelling to watch anyway, and audiences let it go for the same reason they let the seamlessness of the Golden Path Demo go: friction is a strange thing to ask someone to add back in.

There’s also a scale-blindness effect. A grid of four hundred rows and a database of four million rows look identical in a screen share, because you’re only ever looking at the same twenty rows on screen at once. The demo cannot visually communicate that the four hundred clean rows are a curated sliver, not a representative sample, so the brain does what it usually does with limited visual information: it extrapolates, and assumes the part it can see generalizes to the whole.

The actual damage

This is the one that blows up the project timeline more reliably than almost anything else, and it does it quietly, in the data cleansing and reconciliation phase that was budgeted as a two-week task because the demo made data migration look like a solved problem. Then someone runs the real extract, and it turns out eight percent of vendor records have no valid tax ID, eleven percent of item records reference a unit of measure that was deprecated four years ago, and there are nineteen thousand duplicate customer records that need to be identified and merged before go-live, none of which showed up in four hundred rows of hand-picked sample data.

The two-week task becomes a two-month task, the go-live date moves, and the business case that assumed a smooth data conversion now has to absorb a delay that nobody priced in, because the thing that actually determines a migration’s difficulty, the messiness of the real data, was the one thing the demo was specifically built not to show.

The fix, if you’re the one presenting

Run the demo against a real, ugly extract, even a small one, and don’t clean it first. Show the duplicate detection running against actual duplicates. Show what happens when a required field is null. Show the exception queue, and how many records land in it, and what the resolution workflow actually looks like for the person who has to work through that queue by hand. It’s a less polished five minutes. It’s also the only five minutes that tells the prospect anything real about what their conversion will cost.

A migration demo that never encounters bad data hasn’t demonstrated a migration. It’s demonstrated the destination.


The mapping exercise itself looks tidy in a lab too. In Familiar Ground: Mapping CRM to ERP, every concept has a clean twin on the other side, names changed, forms wider, same underlying logic. Real source data is rarely that cooperative, which is exactly the gap this autopsy is about.

Cause of death: two systems that had never spoken before were shown having a conversation written for them.


The presenter switches windows. A record gets created in System A. A few seconds later, as if by magic, the corresponding record appears in System B, fully formed, correctly mapped, no errors. “And that’s it,” the presenter says, “they just talk to each other.” The room relaxes. Integration, historically the single most reliable way for an implementation to go over budget and past deadline, has apparently been solved by two systems having a friendly chat while everyone watched.

Nobody in the room asks what was actually watching that conversation, or who taught it what to say. That’s the autopsy. The two systems didn’t learn to talk to each other. Someone wrote both sides of the script, tested it exactly once, against exactly one scenario, and ran it live in front of you.

What actually happened

An integration demo almost never shows the integration. It shows the happy path of the integration, which is a different and much smaller thing. The record that got created in System A was built to contain precisely the fields System B expects, in precisely the format System B expects them, with no null values in the fields that would trigger a mapping error, no duplicate keys, no encoding mismatch, none of the thousand small inconsistencies that live in a company’s actual data the moment more than one person or one legacy system has touched it.

The script connecting the two systems, whether it’s a middleware platform, a custom connector, or a scheduled job, was very likely written specifically for this demo, tuned against this one scenario, and has never been asked to handle a partial failure, a duplicate record, a field that arrives populated in one system and empty in the other, or a timeout on either end. It works, in the same sense that a bridge works if you only ever drive one specific car across it at one specific speed.

The part that never gets demoed, because it can’t be demoed in three minutes, is everything that happens when the sync fails halfway through. Does the transaction roll back cleanly on both sides, or does System A now believe the record synced while System B never received it. Is there a retry, and if so, does the retry create a duplicate. Is there an alert, and does it go to a person who is actually watching for it, or does it silently populate an error log nobody has looked at since the demo environment was built.

Why it works on smart people

Integration failures are, structurally, invisible until they aren’t. A sync that fails silently doesn’t announce itself. It just produces a slowly widening gap between what System A believes is true and what System B believes is true, and that gap is usually discovered by someone downstream, weeks or months later, reconciling numbers that don’t match and trying to figure out why.

Because the failure mode is invisible, “the demo showed it working” carries more weight than it should, simply because there’s no immediately visible counter-evidence in the room. A broken UI is obvious the moment you see it. A broken integration is obvious only in the reconciliation report nobody runs until month-end close, by which point the demo is a distant memory and the sales team has moved on to the next opportunity.

There’s also a vocabulary problem working in the vendor’s favor. “They just talk to each other” is a satisfying sentence, and it papers over an enormous amount of engineering that either exists, robustly, behind that sentence, or doesn’t exist yet and was built specifically to survive one scripted run.

The actual damage

This is the one that shows up as a reconciliation nightmare rather than a single dramatic failure. Two systems that were sold as integrated drift slowly apart in the weeks after go-live, each one silently correct according to its own records, disagreeing with the other in ways nobody notices until an audit, a customer complaint, or a finance close turns up numbers that don’t tie out. By then the question isn’t “does the integration work,” it’s “how long has it not been working, and what decisions got made on bad data in the meantime.”

The remediation is almost always more expensive than building the integration correctly the first time would have been, because now it includes both the engineering fix and a data cleanup project to reconcile however many weeks or months of silent drift accumulated before anyone noticed.

The fix, if you’re the one presenting

Show a failure on purpose. Send a record with a missing required field, or a duplicate key, and show what happens: does it error visibly, does it queue for retry, does someone get notified, does the other system stay in a known, correct state while the problem gets resolved. If the honest answer is “we haven’t built that handling yet,” say that, and say what the plan is. A prospect who sees a deliberate, controlled failure and a sane recovery path trusts the integration more than one who only ever saw the happy path, because they now know what happens on the day, and there will be a day, when the happy path isn’t what shows up.

Two systems that have never disagreed in front of you haven’t been integrated. They’ve been introduced.


This is exactly the failure mode a ledger-first architecture is built to make impossible. In Is Headless ERP Enough, or Just a Step in the Right Direction?, I walk through a prototype where two disconnected nodes post independent transactions and converge without conflicts, with no consensus protocol and no room for one system to quietly believe something the other doesn’t.