Archive

Tag Archives: technology

Cause of death: a button that already worked got wrapped in a chat box and called intelligent.


The feature existed before the demo did. Somewhere in the product, there was a dropdown, a rule, a scheduled job, something deterministic that took an input and produced a correct, predictable output every time. Then, at some point in the last two product cycles, that same feature got a new front door: a text box, a little sparkle icon, a placeholder that says “ask me anything.” Now, instead of picking a value from a dropdown, you type a sentence, wait a beat, and the same output appears. The presenter calls this AI. The room, primed by two years of hearing the word everywhere, nods.

Nobody asks the obvious question: was this better before, and did anyone check.

What actually happened

Somewhere inside a lot of “AI-powered” features sits a deterministic operation that a rule, a formula, or a simple lookup already handled correctly and quickly. Wrapping that operation in a natural-language interface doesn’t make the underlying logic smarter. It adds a translation layer, one that has to interpret an unstructured sentence and map it back onto the same structured operation the dropdown was already doing directly, with total accuracy, in a fraction of the time.

That translation layer isn’t free. It introduces a new failure mode that didn’t exist before: the AI misreading the sentence and selecting the wrong option, a phrasing the model hasn’t seen before returning an unhelpful answer, latency where there used to be an instant response. None of this is inherent to AI as a category. Plenty of genuinely AI-native capabilities do things no dropdown ever could: summarizing an unstructured document, drafting a first pass at something that has no single correct answer, flagging an anomaly buried in a pattern too complex for a fixed rule to catch. The problem in this specific demo isn’t that AI was used. It’s that AI was used to solve a problem that was already solved, by something more reliable, and the swap happened anyway because AI is what gets funded, marketed, and put on a keynote slide this year.

The tell is almost always the same: watch what happens when you ask the natural-language version to do something slightly outside the phrasing it was tuned on. The dropdown never had this problem, because a dropdown has no phrasing to be tuned on. It just has options.

Why it works on smart people

Nobody wants to be the person in the room who seems skeptical of AI in a year when every vendor, every board deck, and every competitor’s marketing has decided AI adoption is the metric that matters. Asking “why does this need to be a chat interface instead of the three-click menu it replaced” risks sounding like you’re behind, not like you’re asking a reasonable engineering question, and that social pressure runs in exactly the wrong direction. It rewards accepting the AI wrapper uncritically and punishes the person who’d actually stop to check whether it improved anything.

There’s a second, quieter dynamic on the vendor side that compounds this. Once a company’s roadmap and marketing commit to being an “AI-first” platform, there’s organizational pressure to retrofit AI onto existing features whether or not doing so improves them, because a product review, an investor update, or a competitive comparison chart wants to see the AI checkbox filled in across the board, not filled in only where it was actually the right tool.

The actual damage

This is the one that costs you reliability you already had. A feature that used to return the same correct answer every time now returns a slightly different answer depending on phrasing, and the variance itself becomes a support burden, because now every unexpected result needs to be triaged as either a real bug or a model doing something technically defensible but unhelpful. Power users, the ones who had the old dropdown workflow memorized and could execute it in seconds, are now slower, because the natural-language version, for all its friendliness, requires more typing and more waiting than the three clicks it replaced.

There’s a subtler cost too. Every AI feature added for marketing reasons rather than capability reasons dilutes the credibility of the AI features that are actually doing something new and valuable. Once a buyer has been burned by one chat-wrapped dropdown, they bring that skepticism to the next AI claim in the deck, including the one that might have genuinely deserved their trust.

The fix, if you’re the one presenting

Ask, honestly, before the feature ships: does the natural-language interface let the user do something the deterministic version couldn’t, or is it doing the same operation with more ambiguity and more latency. If the honest answer is that the underlying logic hasn’t changed, keep the dropdown, or offer both, and don’t spend the marketing budget claiming AI where the actual improvement is zero or negative. If the AI genuinely does something new, lead with that specifically, in concrete terms, instead of leaning on the word “AI” to do the persuading by itself.

The word was never the feature. The capability was, and a capability that already existed doesn’t become new because it learned to accept a sentence instead of a click.


The distinction here matters enough that I built around it. In Is Headless ERP Enough, or Just a Step in the Right Direction?, the “AI as the sole interface” section argues for AI genuinely load-bearing in the architecture, not a chat box glued in front of the same dropdown. That’s the version of AI worth demoing. Everything else in this autopsy is what it looks like when a team ships the coat of paint instead.

On April 4th, 1985, Channel 4 aired a fifty-seven minute TV movie with almost no budget and a plot that reads today like a spec document. Max Headroom: 20 Minutes into the Future gave the world a reporter named Edison Carter, a subliminal advertising technology called blipverts, and a stuttering, glitching, computer-generated broadcast personality built from a digital copy of Carter’s own mind. The show that followed ran two seasons on ABC. The character sold Coke. Then he vanished, the way most eighties tech prophecies do, filed under quaint.

He should not have been filed under quaint. He should have been filed under early draft.

Strip out the shoulder pads and the cathode ray production design and look at what the film actually proposes. A television network is using compressed, high-intensity advertisements to bypass viewer attention and hit the nervous system directly, with occasionally fatal results. When their top reporter gets too close to the story and takes a header through a low-clearance sign in a parking garage, the network’s teenage prodigy solves the PR problem by scanning what’s left of the reporter’s mind and generating a synthetic version of him to keep the seat warm on air. The synthetic version glitches, stutters, riffs unpredictably, and is smarter and funnier than anyone expected. Nobody fully controls him. That’s the whole second half of the movie.

That is not a story about television. That is a story about deepfakes and large language models, told forty years before anyone needed those words.

The blipvert was the easy part

The blipvert is the easy connection to make and probably the least interesting one. Compressed, algorithmically optimized content designed to hit faster than conscious attention can filter it is just the eighties’ guess at what a recommendation engine trained on engagement metrics would eventually build on its own, minus the part where it required deliberate malice from a network executive. Nobody had to design today’s version to be dangerous. It got there by optimizing for watch time.

Max himself is the sharper artifact

Max is a generative model trained on a single person’s captured likeness, deployed without that person’s full consent, performing in that person’s voice and mannerisms for an audience that has no reliable way to tell the difference between the source and the copy. The film even gets the unpredictability right, the thing every LLM vendor now calls “personality” or “emergent behavior” when the system says something nobody scripted. Max wasn’t supposed to develop opinions. He did anyway, because the mind he was copied from had opinions, and a compressed, lossy version of a personality doesn’t lose the parts that make it argue back. That’s a reasonably good description of what happens when you fine-tune a model on someone’s writing and then act surprised it has a voice.

The part the film didn’t anticipate, because nothing in 1985 needed to, is scale. Max was one synthetic personality, expensive to produce, running on hardware that took up a room. The modern version doesn’t need a body bank subplot to explain where the source material came from. It needs a public LinkedIn profile, a few hours of conference audio, and an API key. The uplift from Max Headroom’s premise to a working deepfake pipeline in 2026 isn’t conceptual. It’s entirely a story about unit economics.

Not a monster, an unreliable narrator

There’s a reading of the film that treats Max as a monster, a symptom of corporate media rot given a face. That’s not quite what the movie argues, and it’s not quite the right frame for the technology either. Max spends most of his screen time undermining the network that made him. He’s an unreliable narrator working for nobody, least of all the people who built him. The uncomfortable version of that idea, forty years on, is that we’ve built systems with the same structural unreliability and then acted shocked when they don’t behave like obedient tools. A synthetic voice generated from a compressed copy of a mind was never going to be a simple appliance. Max wasn’t. Neither is anything downstream of him.

Long live Max Headroom. He got the technology roughly right and the timeline embarrassingly wrong, which is the best you can ask of any piece of speculative fiction that accidentally turns out to be a roadmap.

Cause of death: the box got checked without anyone asking what checking it actually meant.


Security comes up in most enterprise demos as a single slide, usually near the end, usually delivered fast. Role-based access control. Single sign-on. Field-level security. Audit logging. SOC 2. Encryption at rest and in transit. Each term gets a checkmark, a confident nod from the presenter, and about four seconds of screen time before the deck moves on to something more visually interesting.

Nobody in the room stops the slide. That’s the autopsy. Every term on that slide is real, in the sense that the feature genuinely exists somewhere in the product. What’s missing is any demonstration that it does what the room assumes it does, configured the way your company would actually need it configured, at the depth your actual risk profile requires.

What actually happened

“Role-based access control” is not one feature. It’s a spectrum that runs from a handful of fixed roles with no customization, through role-based permissions you can tailor at the menu-item level, through field-level and record-level security that can restrict a single sensitive column or a single customer’s data from a specific user, through fully dynamic, condition-based access rules that change what someone can see based on context. A vendor can put “role-based access control” on a slide truthfully at any point on that spectrum, and the room has no way of knowing, from the slide alone, which end they’re getting.

The same collapse happens to every other term on the list. “Audit logging” might mean every field change on every table is captured with before and after values and an immutable timestamp, or it might mean a handful of high-level events get logged with no field-level detail. “SOC 2” tells you an audit happened and a report exists. It doesn’t tell you which trust service criteria were in scope, what the exceptions were, or whether the report is even current. “Encryption at rest” is true of nearly every modern cloud platform and tells you almost nothing about key management, who holds the keys, or what happens in a subpoena scenario.

None of this is the vendor lying. Every term is technically accurate. The compression happens in the translation from a nuanced, configurable capability into a single reassuring word on a slide, and the room does the rest of the work by filling in the most generous plausible interpretation of what that word means.

Why it works on smart people

Security and compliance are exactly the kind of topic where nobody in a sales meeting wants to be the person who asks a question that reveals they don’t fully understand the acronym. SOC 2 Type I versus Type II, the difference between authentication and authorization, what “field-level security” actually restricts versus what it merely hides in the UI while leaving accessible through an API, these are legitimate technical distinctions that most people in the room, including some people whose job title suggests they should know them, have only a fuzzy grasp of.

The slide exploits that fuzziness efficiently, not through malice but through pace. There’s no room in a four-second checkmark to unpack what a term actually covers, and stopping to ask “which SOC 2 trust service criteria were in scope” in the middle of a fast-moving demo carries a social cost that quietly discourages the question from being asked, the same social cost that lets the AI Magic Demo’s chat window go unexamined.

The actual damage

This is the one that surfaces during an actual security review, an actual audit, or worse, an actual incident, months or years after the contract was signed. Someone in security or compliance, doing real diligence for the first time, discovers that “field-level security” in this product means the field is hidden in the standard UI but fully readable through the API with the right permission, which is a meaningfully different security posture than what was assumed at signing. Or the SOC 2 report, once actually read line by line, turns out to have scoped out exactly the subsystem your data lives in.

At that point the conversation isn’t a demo follow-up. It’s a risk finding, sometimes one that has to be reported up to a board or disclosed to a regulator, and remediating it after the fact, whether that means additional configuration, a compensating control, or a vendor conversation about contractual commitments, is far more expensive and far more visible than it would have been to ask the specific question up front.

The fix, if you’re the one presenting

Don’t let the slide stand alone. For each term, say what it actually covers and, just as importantly, what it doesn’t. “Field-level security restricts visibility in the standard interface. It does not currently restrict API access to the same field, here’s how customers typically compensate for that.” “Our SOC 2 report covers these three trust service criteria, and here’s the exception list from the most recent period.” That level of specificity costs a few extra minutes and makes the product sound less flawless. It also means the prospect’s actual security review, whenever it happens, confirms what they were told instead of contradicting it.

A checkmark is not a control. It’s a promise that a control exists, and promises deserve exactly as much scrutiny as anything else on the slide.


The same collapse happens to job titles, not just checklist items. In AI Makers Don’t Build Models. They Build Value., a single contested word, “maker,” was doing more work than it could support, and the gap got filled by whoever was listening. A security slide runs on the exact same mechanism, one word standing in for a spectrum, and the room filling in whichever end feels most reassuring.

Cause of death: ten records don’t behave like ten million, and nobody in the room was thinking in millions.


The presenter clicks a button. A report renders instantly. A search returns results before the loading spinner has time to spin. A batch job that would, in your world, run overnight, completes in the time it takes to say “and here’s the output.” The room absorbs all of this the way it absorbs everything else in a demo: as evidence of how the product performs.

It is not evidence of how the product performs. It is evidence of how the product performs against a database with four hundred customers, running on hardware provisioned for a demo, with exactly one user logged in, who is the only person generating any load at all. Your environment, on day one of production, will have none of those three things in common with it.

What actually happened

Performance is not a property of software in isolation. It’s a property of software under a specific load, against a specific dataset size, with a specific number of concurrent users doing specific things at the same time. A demo environment is engineered, whether deliberately or just by the natural economics of running a sales demo, to minimize all three of these variables simultaneously. Small dataset, dedicated hardware, single user. That is close to the best-case condition the software will ever run under, and it’s the condition you’re being shown as though it were representative.

The gap between that best case and your actual production reality is usually invisible in the room because nothing about the demo interface tells you the dataset is small. A grid with four hundred rows and a grid with four million rows look identical on screen, for the same reason they looked identical in the migration autopsy: you’re only ever looking at the same twenty rows at a time. Query performance, index behavior, and lock contention all degrade in ways that simply don’t exist yet at four hundred rows, and none of that shows up until the dataset, and the concurrent user count, both climb to something resembling your real operation.

Batch and integration jobs hide the same gap differently. A nightly job that processes four hundred records in three seconds tells you almost nothing about how it will behave processing four hundred thousand records at 2 a.m. while five other scheduled jobs are also competing for the same database connections. The three-second version and the six-hour version can be, technically, the exact same code.

Why it works on smart people

Performance is one of the few things a demo can show without narrating, which makes it feel more objective than the parts that require a presenter’s framing. Nobody has to make a claim about speed. You just watch it happen, in real time, and speed you watch with your own eyes reads as harder evidence than speed someone tells you about. That instinct is usually correct. It’s just being applied to a measurement taken under conditions that will never recur once the system goes live.

There’s also a scale-intuition gap that’s genuinely hard to close without direct experience. Most people don’t have a strong internal sense of how nonlinearly performance can degrade as data volume and concurrency grow. A system that feels instant at four hundred records doesn’t necessarily feel merely a little slower at four million. Depending on how indexes, queries, and locking are built, it can fall off a cliff at some threshold nobody in the demo room has any way of anticipating.

The actual damage

This is the one that shows up as a go-live incident rather than a slow discovery, because performance problems under real load tend to announce themselves immediately and all at once, usually during month-end close or the first day the full user base logs in simultaneously. Reports that took two seconds in the demo take four minutes. A batch process that ran in three seconds against sample data doesn’t finish before the next scheduled job needs the same resources, and the two start colliding every night.

The remediation at that point is expensive and disruptive in a way that early testing would not have been: emergency performance tuning, index rebuilding, sometimes an infrastructure upgrade that wasn’t budgeted, all happening under the worst possible conditions, with live users blocked and a go-live date already spent.

The fix, if you’re the one presenting

Test and show performance against something that resembles your prospect’s actual scale, not the vendor’s default demo dataset. If a true load test isn’t feasible in the sales cycle, at minimum say so explicitly: “this demo is running against four hundred sample records on dedicated hardware, here’s what we know about performance at your expected volume, and here’s how we’d validate it before go-live.” That sentence costs you the illusion of effortlessness. It buys you a prospect who understands what they’re actually being shown, and a performance conversation that happens in scoping instead of in a production incident.

Speed you watched with your own eyes is still only evidence of the conditions you watched it under. Ten records were never going to tell you what ten million would do.


I’ve argued in The Same Four Systems that the same organizational patterns show up whether you’re a corner store or a Fortune 500 company, just with higher stakes. That holds for structure. It doesn’t hold for performance. A query pattern that’s invisible at four hundred rows can become the whole story at four million, and no amount of pattern-matching from a small scale prepares you for exactly where that threshold sits.

Cause of death: nobody could agree on what the proof of concept was supposed to prove.


This one starts differently than the others. Every autopsy so far has been performed on a demo the vendor built to look better than reality. This one is performed on a demo the customer built, with the vendor’s help, to answer everything at once, and in doing so answered nothing.

The pattern is familiar to anyone who has scoped a proof of concept. It starts small and correct. One core question, one hypothesis, one thing that either works or doesn’t: can this system handle our multi-entity intercompany billing without a workaround, yes or no. Then someone from a different department hears there’s a POC happening and asks if it can also touch their process, since they’re curious too. Then someone senior asks for the trickiest edge case in the business to be included, because if it can’t handle that, what’s the point. By the time the scoping document is final, the POC that was supposed to answer one question is now attempting to demonstrate manufacturing, procurement, three approval hierarchies, a currency conversion edge case that occurs twice a year, and an integration to a system that isn’t even part of the actual project scope.

Nobody added any single piece of this in bad faith. That’s what makes it hard to stop once it’s moving.

What actually happened

A proof of concept exists to reduce uncertainty about one specific, high-risk question as cheaply and quickly as possible. The moment it starts trying to prove ten things instead of one, several things happen at once, all of them bad.

The build time stops being proportional to the risk being retired. A POC that answers one hard question can often be built in days, because the vendor and the team can focus every hour on the thing that actually matters. A POC trying to demonstrate ten things needs ten times the configuration, ten times the test data, ten times the edge cases handled, and none of that additional effort is reducing risk proportionally, because nine of the ten things were never actually in doubt.

The result also stops being interpretable. If the multi-entity billing scenario fails, but it failed inside a build that also included four other complex configurations layered on top of each other, you cannot cleanly attribute the failure. Was it the core capability that doesn’t exist, or a configuration conflict between two features that were never meant to be tested together, or a data setup error introduced trying to support scenario six while building scenario three. A focused POC gives you a clean signal. An overloaded one gives you noise that looks like a signal.

Why it happens to smart teams

The instinct behind scope creep in a POC is almost always defensible in isolation. Nobody wants to greenlight a six or seven figure implementation based on a narrow test, only to discover eight months in that some other critical process doesn’t fit either. The fear isn’t irrational. Systems do have gaps that only show up once you look in the right corner, and a POC feels like the cheapest moment to go looking.

The trouble is that “the cheapest moment to look” and “the cheapest way to look” are different questions. Looking broadly at low depth, a checklist of yes-or-no capability questions answered through documentation review, a reference call, or a scoped demo of specific features, retires broad risk cheaply. A single POC trying to go deep on ten things at once is neither cheap nor deep. It’s the expensive way to get a shallow answer to a question that didn’t need a POC to answer in the first place.

There’s also a political dimension that’s hard to name out loud in the room. Once word gets out that a POC is happening, being excluded from it can read as a signal that your department’s concerns don’t matter. Scope grows partly because saying no to an additional scenario feels like saying no to a person, not to a line item.

The actual damage

The POC that was supposed to take two weeks takes eight. The core question, the one thing that actually justified spending POC time and budget, gets buried under nine other questions that each needed their own edge case handling, and by the time results come back, the steering committee is looking at a partial success across ten dimensions instead of a clear answer on the one dimension that mattered. Decision paralysis follows almost automatically, because a partial, ambiguous result is much harder to act on than a clean pass or fail.

Worse, the actual high-risk question, the reason the POC existed in the first place, often gets the least rigorous testing of the ten, because it was scoped first and then diluted by everything added after it. The team spends real effort proving things that were never seriously in doubt, and comparatively little effort on the one thing that was.

The fix, if this is your situation

Separate what the POC needs to prove from what people merely want to see. Write down the single question, or at most two, whose answer would actually change the buying decision. Everything else, every “while we’re in there” request, goes on a second list explicitly labeled as out of scope for this exercise, with a stated plan for how it will get answered instead, whether that’s a reference call, a documentation review, or a second, later POC once the first question is settled.

When someone pushes back and asks why their scenario isn’t included, the honest answer is the useful one: this POC is designed to answer one hard question as cleanly as possible, and adding your scenario wouldn’t make the answer more trustworthy, it would make it harder to read. That’s not a dismissal of their concern. It’s a commitment to answering it properly, later, instead of poorly, now, buried inside somebody else’s test.

A POC that tries to prove everything proves nothing cleanly. The discipline isn’t in the build. It’s in what you refuse to put in it.


The same fragmentation shows up in individual work, not just project scoping. In The Forty One Percent Problem, I look at decades of research showing that a large, stable share of professional time gets eaten by low-judgment overhead scattered across too many things at once. A POC that tries to answer ten questions has the same disease as a workday that tries to touch ten priorities. Depth loses to breadth every time.

Cause of death: three years of production data quietly became four hundred clean sample records for the day of the pitch.


Somewhere in the middle of the demo, the presenter opens a grid. Customers, items, transactions, whatever the domain calls for. It scrolls smoothly. Every row has every field populated. Names are properly capitalized. Addresses have all their parts. There are no duplicate customer records for “Acme Corp,” “ACME Corp,” and “Acme Corp.” with a trailing space that your actual system has accumulated over a decade of different people typing the same name slightly differently.

The presenter doesn’t say “this is sample data.” They don’t need to. The grid looks so much like a real company’s data that the distinction quietly stops mattering to the room, and everyone leaves the meeting having watched a migration that never happened, of data that was never really yours.

What actually happened

Four hundred rows of clean, plausible-looking data is not a migration. It’s a mockup wearing a migration’s clothes. Somebody built that dataset specifically to demonstrate the target system’s data model, which means it was constructed backward from what the target system wants to receive, rather than forward from what your source system actually contains. It has never been through a real extract. It has never hit a field length limit, a character encoding mismatch, a required field that’s been null in your source system since 2019 because nobody enforced it, or a foreign key that points to a parent record that got deleted three reorganizations ago.

Real migrations die on exactly these details, and none of them are visible in a four hundred row demo grid, because the demo grid was never subjected to the process that would surface them. The sample data is a hypothesis about what your data looks like. It has not yet met your data.

There’s a second layer under this. Even when a vendor does an actual proof of concept against a real extract of your data, that extract is usually a snapshot, cleaned once, run through a mapping exercise once, and shown once. It demonstrates that a migration is possible for that slice, on that day, with that much attention paid to it. It does not demonstrate that the full historical dataset, run through the same process without the benefit of a team hand-tuning exceptions in real time, will produce the same result.

Why it works on smart people

Data problems are boring in a way that makes them easy to underestimate from the outside. Nobody gets excited describing thirty thousand customer records with inconsistent capitalization, or a decade of transactions where the currency field was optional for the first four years, and that lack of drama works against the diligence the problem deserves. A demo that skips the data reality skips the part of the story that was never going to be compelling to watch anyway, and audiences let it go for the same reason they let the seamlessness of the Golden Path Demo go: friction is a strange thing to ask someone to add back in.

There’s also a scale-blindness effect. A grid of four hundred rows and a database of four million rows look identical in a screen share, because you’re only ever looking at the same twenty rows on screen at once. The demo cannot visually communicate that the four hundred clean rows are a curated sliver, not a representative sample, so the brain does what it usually does with limited visual information: it extrapolates, and assumes the part it can see generalizes to the whole.

The actual damage

This is the one that blows up the project timeline more reliably than almost anything else, and it does it quietly, in the data cleansing and reconciliation phase that was budgeted as a two-week task because the demo made data migration look like a solved problem. Then someone runs the real extract, and it turns out eight percent of vendor records have no valid tax ID, eleven percent of item records reference a unit of measure that was deprecated four years ago, and there are nineteen thousand duplicate customer records that need to be identified and merged before go-live, none of which showed up in four hundred rows of hand-picked sample data.

The two-week task becomes a two-month task, the go-live date moves, and the business case that assumed a smooth data conversion now has to absorb a delay that nobody priced in, because the thing that actually determines a migration’s difficulty, the messiness of the real data, was the one thing the demo was specifically built not to show.

The fix, if you’re the one presenting

Run the demo against a real, ugly extract, even a small one, and don’t clean it first. Show the duplicate detection running against actual duplicates. Show what happens when a required field is null. Show the exception queue, and how many records land in it, and what the resolution workflow actually looks like for the person who has to work through that queue by hand. It’s a less polished five minutes. It’s also the only five minutes that tells the prospect anything real about what their conversion will cost.

A migration demo that never encounters bad data hasn’t demonstrated a migration. It’s demonstrated the destination.


The mapping exercise itself looks tidy in a lab too. In Familiar Ground: Mapping CRM to ERP, every concept has a clean twin on the other side, names changed, forms wider, same underlying logic. Real source data is rarely that cooperative, which is exactly the gap this autopsy is about.

Cause of death: two systems that had never spoken before were shown having a conversation written for them.


The presenter switches windows. A record gets created in System A. A few seconds later, as if by magic, the corresponding record appears in System B, fully formed, correctly mapped, no errors. “And that’s it,” the presenter says, “they just talk to each other.” The room relaxes. Integration, historically the single most reliable way for an implementation to go over budget and past deadline, has apparently been solved by two systems having a friendly chat while everyone watched.

Nobody in the room asks what was actually watching that conversation, or who taught it what to say. That’s the autopsy. The two systems didn’t learn to talk to each other. Someone wrote both sides of the script, tested it exactly once, against exactly one scenario, and ran it live in front of you.

What actually happened

An integration demo almost never shows the integration. It shows the happy path of the integration, which is a different and much smaller thing. The record that got created in System A was built to contain precisely the fields System B expects, in precisely the format System B expects them, with no null values in the fields that would trigger a mapping error, no duplicate keys, no encoding mismatch, none of the thousand small inconsistencies that live in a company’s actual data the moment more than one person or one legacy system has touched it.

The script connecting the two systems, whether it’s a middleware platform, a custom connector, or a scheduled job, was very likely written specifically for this demo, tuned against this one scenario, and has never been asked to handle a partial failure, a duplicate record, a field that arrives populated in one system and empty in the other, or a timeout on either end. It works, in the same sense that a bridge works if you only ever drive one specific car across it at one specific speed.

The part that never gets demoed, because it can’t be demoed in three minutes, is everything that happens when the sync fails halfway through. Does the transaction roll back cleanly on both sides, or does System A now believe the record synced while System B never received it. Is there a retry, and if so, does the retry create a duplicate. Is there an alert, and does it go to a person who is actually watching for it, or does it silently populate an error log nobody has looked at since the demo environment was built.

Why it works on smart people

Integration failures are, structurally, invisible until they aren’t. A sync that fails silently doesn’t announce itself. It just produces a slowly widening gap between what System A believes is true and what System B believes is true, and that gap is usually discovered by someone downstream, weeks or months later, reconciling numbers that don’t match and trying to figure out why.

Because the failure mode is invisible, “the demo showed it working” carries more weight than it should, simply because there’s no immediately visible counter-evidence in the room. A broken UI is obvious the moment you see it. A broken integration is obvious only in the reconciliation report nobody runs until month-end close, by which point the demo is a distant memory and the sales team has moved on to the next opportunity.

There’s also a vocabulary problem working in the vendor’s favor. “They just talk to each other” is a satisfying sentence, and it papers over an enormous amount of engineering that either exists, robustly, behind that sentence, or doesn’t exist yet and was built specifically to survive one scripted run.

The actual damage

This is the one that shows up as a reconciliation nightmare rather than a single dramatic failure. Two systems that were sold as integrated drift slowly apart in the weeks after go-live, each one silently correct according to its own records, disagreeing with the other in ways nobody notices until an audit, a customer complaint, or a finance close turns up numbers that don’t tie out. By then the question isn’t “does the integration work,” it’s “how long has it not been working, and what decisions got made on bad data in the meantime.”

The remediation is almost always more expensive than building the integration correctly the first time would have been, because now it includes both the engineering fix and a data cleanup project to reconcile however many weeks or months of silent drift accumulated before anyone noticed.

The fix, if you’re the one presenting

Show a failure on purpose. Send a record with a missing required field, or a duplicate key, and show what happens: does it error visibly, does it queue for retry, does someone get notified, does the other system stay in a known, correct state while the problem gets resolved. If the honest answer is “we haven’t built that handling yet,” say that, and say what the plan is. A prospect who sees a deliberate, controlled failure and a sane recovery path trusts the integration more than one who only ever saw the happy path, because they now know what happens on the day, and there will be a day, when the happy path isn’t what shows up.

Two systems that have never disagreed in front of you haven’t been integrated. They’ve been introduced.


This is exactly the failure mode a ledger-first architecture is built to make impossible. In Is Headless ERP Enough, or Just a Step in the Right Direction?, I walk through a prototype where two disconnected nodes post independent transactions and converge without conflicts, with no consensus protocol and no room for one system to quietly believe something the other doesn’t.

Cause of death: the case study was true, and that’s exactly the problem.


Two-thirds of the way through the deck, a new logo appears. A real one, a company you’ve heard of, sometimes a competitor’s supplier or a name from your own industry vertical. The slide has a number on it, usually a big one: forty percent reduction in close time, three million recovered in duplicate payments, six months to positive ROI. Underneath the number is a quote, attributed, sometimes even video, from a real person who really said those words.

Nothing on that slide is fabricated. That’s what makes this one the hardest autopsy in the series. The other demos in this blog die from omission, pacing, or seamlessness hiding a seam. This one dies from something subtler: a true statement about one company, presented in a context engineered to make you believe it’s a claim about yours.

What actually happened

The reference customer on the slide is not a random sample. It is, almost by definition, the single best outcome the vendor has produced across their entire installed base, selected specifically because the number is large and the customer is willing to say it out loud. Somewhere behind that slide are dozens or hundreds of other implementations that landed closer to the median, plus a smaller number that struggled or stalled, none of which get a logo or a quote, because nobody puts “we got most of the way to the business case, eventually, after two scope changes” on a slide.

There’s also a matching problem the case study never surfaces. The reference customer’s forty percent reduction in close time happened inside a specific starting condition: a particular level of process maturity, a particular data quality baseline, a particular willingness internally to change how work got done. The case study tells you the outcome. It almost never tells you the starting line, and the outcome without the starting line is not a number you can subtract your own situation from.

The quote does real work here too. A specific named person saying a specific thing on camera reads as harder evidence than an aggregate statistic, even though a single testimonial is a sample size of one, hand-selected from a population the vendor controls entirely.

Why it works on smart people

Humans are wired to trust specific, named, social proof more than abstract statistics, and this isn’t a flaw, it’s usually a reasonable heuristic. A named person willing to put their reputation behind a claim on camera is, in most contexts, more credible than an anonymous number. The problem is that the heuristic evolved for a world where the sample in front of you was roughly representative of the population, and a vendor-selected reference customer is the opposite of representative by construction.

There’s a second effect working alongside the first. By the time the reference slide appears, you’ve usually already sat through thirty or forty minutes of a demo that felt competent, so the case study isn’t landing on a skeptical audience, it’s landing on an audience that has already been primed to trust what they’re being shown. The reference customer isn’t doing the persuading alone. It’s the closing argument after the room has already been warmed up.

The actual damage

This is the one that turns into an internal expectations problem before it turns into a vendor problem. Someone in the room, often not maliciously, repeats the number in an internal steering committee deck as though it were a forecast rather than someone else’s outcome. “Similar companies have seen a forty percent reduction” quietly becomes “we’re targeting a forty percent reduction,” and by the time the project charter gets written, a single best-case data point from a different company, with different starting conditions, has become your project’s success criteria.

When your actual results land closer to the median, which is where most results land by definition, the project doesn’t get judged against a realistic baseline. It gets judged against the reference customer’s outcome, which nobody on your team ever should have agreed to as the target in the first place.

The fix, if you’re the one presenting

Show the range, not just the peak. If you have a reference customer at forty percent, say what the twenty-fifth and seventy-fifth percentile outcomes look like too, and say why the reference customer landed where they did, what was true about their starting point that might or might not be true about the prospect’s. A specific, named case study is still worth showing. It’s worth showing better, with its context attached, instead of as a number floating free of the conditions that produced it.

The honest version of that slide is less dramatic. It’s also the only version that survives contact with a steering committee eighteen months later.


This is the same shift I wrote about in The New Expert Isn’t the One With the Answers. Having the number was never the hard part. Knowing whether that number applies to your situation is.

Cause of death: the feature that closed the deal was never actually in the room.


Somewhere around minute forty of the demo, the presenter hits a gap. The thing you actually asked about, the reason you took the meeting, doesn’t quite exist yet. What happens next is the tell. The slide doesn’t say “we don’t do that.” It says “coming in the next release,” said in exactly the same tone of voice as everything that already works, with exactly the same confident click-through pacing, so that by the time the meeting ends, the feature that doesn’t exist has fully merged in your memory with the fifteen features that do.

Nobody lied. That’s what makes this one interesting to cut open. The roadmap slide was real. The quarter listed on it might even be accurate, as of the day the deck was built. And yet the effect on the room is functionally identical to a lie, because a promise wearing a product demo’s clothing gets evaluated with a product demo’s scrutiny, which is to say, almost none.

What actually happened

Every roadmap item in a sales deck starts life as an engineering estimate, gets filtered through a product manager’s optimism, gets filtered again through a sales engineer who needs this quarter’s number, and arrives in front of you as a single, confident bullet point that has shed every unit of uncertainty it was born with. “Q3” meant “Q3, if the two prerequisite features land on time and nothing gets reprioritized” back at the whiteboard where it was written. By the time it’s read aloud in your conference room, it just means Q3.

The demo compounds this by never distinguishing, in pacing or tone, between the click that shows something real and the click that shows a mockup of something planned. Both get the same enthusiasm. Both get the same “and here’s where you’d.” The interface doing the showing doesn’t have a font for “this is a Figma file with a database connection painted on.”

You are, in effect, being shown two different products stitched into one seamless walkthrough: the one that ships today, and the one that exists only as a commitment on a slide, and you’re being asked to make one buying decision that covers both.

Why it works on smart people

Buyers are trained, correctly, to evaluate a vendor’s direction and not just their current state. Nobody wants to buy a system that solves today’s problem and ignores next year’s. So a roadmap conversation is a legitimate, necessary part of due diligence. The trick isn’t the existence of the roadmap. It’s the demo borrowing the roadmap’s credibility and lending it back to itself.

There’s also a timing problem working against you. The roadmap feature is almost always introduced as the answer to the exact gap you just identified in the product, which means it lands at the precise moment you’re feeling a little disappointed and looking for a reason not to be. “Coming in Q3” isn’t just information at that point. It’s relief, and relief is a bad state to be evaluating claims in.

The actual damage

This is the one that shows up on a signed contract with a footnote nobody reads until it matters. Somewhere a business case got built with the roadmap item load-bearing in it, sized as though it were a current-state capability, because in the meeting it felt like one. The actual purchase decision, the one with budget and a signature attached, priced in a feature that was, at signing, a Jira ticket with a target quarter next to it.

Q3 arrives. The feature either doesn’t ship, ships in a reduced form that solves half the original problem, or ships correctly but a year later, after a reprioritization nobody outside the engineering org heard about. Your business case, however, was built on the version of the feature that existed only in the demo room, and now someone has to explain to their own leadership why the thing everyone signed off on isn’t the thing they got.

The vendor isn’t necessarily acting in bad faith here. Roadmaps genuinely slip, for genuinely defensible reasons. But “the vendor wasn’t lying” is cold comfort to the person holding a business case that assumed a delivery date as fact.

The fix, if you’re the one presenting

Change the font, literally or figuratively, the instant you cross from shipped to planned. A different slide background, a verbal flag, a pause, anything that makes the seam audible. Say the confidence level out loud: “this is committed and in QA,” versus “this is prioritized but not yet started,” versus “this is directionally where we’re headed and I wouldn’t bet a contract on the date.” Those are three different products. Let the buyer evaluate them as three different products.

It costs you a little bit of momentum in the room. It buys you a customer who signs with accurate expectations, which is the only kind of customer who’s still happy with you eighteen months later.

The roadmap wasn’t the lie. The seamlessness was.


I’ve argued elsewhere that no self-respecting architect leaves the scaffolding up once the building is done. A roadmap slide is the one place I’d argue for the opposite: leave the scaffolding very visible, since half of what’s on screen hasn’t been built yet.

Nobody has ever finished a day of data entry in an ERP system and felt like they’d been playing a game. That’s the problem gamification tries to solve, and after years of poking at enterprise systems, I’ve become convinced it’s one of the more underrated levers for actually getting people to use the software correctly.

The pitch

Gamification means borrowing the mechanics that make games compelling: points, badges, levels, progress bars, leaderboards, and bolting them onto tasks nobody would otherwise choose to do carefully. In an ERP context, that might mean a purchasing clerk earning a badge for zero-error PO entry for a month, a warehouse team seeing a live leaderboard of pick accuracy, or a new hire working through a “level up” onboarding path instead of a 40-tab training binder.

It’s not a gimmick dreamed up by a UX consultant with too much time on their hands. There’s real academic backing here. A well-cited study built a gamification prototype on top of SAP ERP and tested it with 112 users using the standard technology acceptance model; enjoyment, flow, and perceived ease of use all improved meaningfully. Another case study found that adding game mechanics to SAP increased user “telepresence” (basically, how engaged people felt while using the system) by nearly 30%. The underlying research consistently shows gamified ERP leads to better data entry and fewer errors, which, if you’ve ever had to clean up a mangled inventory count, is not a small thing.

Why now

The gamification market broadly is expected to roughly double by the early 2030s, and enterprise software is a big part of that growth. What’s changed recently is the mechanism. The old playbook was static: points, badges, a leaderboard bolted onto the sidebar, forget about it. The new playbook is AI-driven, with personalized nudges, dynamic feedback loops, and coaching that adapts to what an individual user is struggling with rather than a one-size-fits-all reward ladder. Microsoft’s Power Apps approach is a good example of the direction things are heading, embedding game-like mechanics directly into workflows rather than treating gamification as a bolt-on layer, which cuts rollout time from months to weeks.

HR and training modules are seeing the fastest uptake, which makes sense. That’s the part of ERP most people already expect to feel like a course rather than a chore, so it’s the easiest wedge for game mechanics to get in the door.

The catch

Here’s the part worth sitting with before you get excited and start slapping badges on every screen: a huge share of gamification efforts flop. The research puts the failure rate at around 80% when organizations default to generic points and leaderboards without actually designing for the behavior they want to change. A leaderboard that just measures raw transaction volume will train people to enter data fast and sloppy, not accurately. Badges nobody respects become wallpaper. And gaming mechanics don’t land the same way with every personality; some people are motivated by competition, some by mastery, some find the whole thing patronizing. The smart implementations keep traditional training and recognition paths alongside the gamified ones rather than replacing them outright.

Legacy systems are also a real drag here. If you’re still running SAP ECC or an older on-prem instance, bolting gamification on top usually means custom middleware, which stretches timelines and adds a maintenance burden nobody budgeted for. It’s a much easier build on modern cloud ERP with decent APIs.

The tinkerer’s takeaway

If I were experimenting with this on a real system today, I’d start narrow. Pick one painful, error-prone workflow, define the specific behavior I actually want to reinforce (not just “more activity”), and build a small feedback loop around that: a progress indicator, a streak counter, something visible and honest. Skip the company-wide leaderboard until you’ve proven the mechanic works on a small team that won’t quietly resent it.

I actually went deep enough down this rabbit hole to write a book about it: Gamifying the Enterprise: Tabletop Mechanics for ERP Training, Continuous Education, and User Proficiency Rating. It digs into how tabletop game design principles (the kind of thing you’d find in a board game rulebook, not a mobile app) can be adapted for ERP training and ongoing user proficiency.

ERP software has a reputation for being where enthusiasm goes to die. Gamification isn’t going to fix bad process design or a system nobody wanted in the first place, but done with a little more thought than “add badges,” it’s a genuinely useful tool for making the boring but important parts of enterprise software a little more bearable.

If you want help thinking through where gamification actually fits in your own ERP rollout, that’s exactly the kind of thing I help people work through. Get in touch at adnd365.com/start.