An Advanced Dungeons & Dynamics 365 session. Episode 3 of 14.

Previously on: Episode 2, The Shadow Ledger, the party discovered that Priya had been running a shadow spreadsheet for fourteen months, revealing a six figure variance nobody could explain. Priya left for vacation before anyone got an answer.

Every ERP has a chart of accounts. Not every ERP has a chart of accounts that actually makes sense. Sometimes what looks like a financial structure is really three different structures, built at three different points in the project’s history, quietly stapled together and never reconciled.

That’s the dungeon the party walks into this episode.


The Session

GM: You descend into the chart of accounts. It is, structurally, three different charts of accounts stapled together.

Sable: There are four segments for cost center. Four. Two of them mean the same thing in different departments.

Vex: (picking the lock on a legacy dimension table) Found something. A dimension called TEMP_DO_NOT_USE.

Thorne: How many postings?

Vex: Eleven thousand, two hundred and six.

Marge: Whose decision was that?

Vex: Nobody’s. That’s the problem. Somebody created it during UAT for testing and it just never got decommissioned. It’s load bearing now.

Ai. Cassiopeia: I flag that removing it will break forty three saved reports, none of which are documented.

Sable: We’re not removing it today. We’re mapping around it. Nobody touch the dimension.

A junior analyst walks past the war room door, glances in at the whiteboard now covered in a hand drawn diagram of overlapping cost center segments, and quietly closes the door again without saying anything. Nobody in the party blames him.

Thorne: (studying the diagram) This isn’t a chart of accounts. It’s an archaeology site.

Sable: Every layer’s a different consultant’s fingerprints. I can tell you which segment was added during which go live attempt just by how badly it fits with the one next to it.

Vex: (still elbow deep in the dimension table) There’s a comment field on this one. Someone left a note.

Marge: What’s it say?

Vex: “Will fix in phase two.”

Ai. Cassiopeia: I checked the project timeline. There was no phase two.


Why nobody decommissions the temp table

Every implementation accumulates a TEMP_DO_NOT_USE somewhere. It starts as exactly what the name says, a placeholder for testing, something built under deadline pressure that was always going to get cleaned up later. Then UAT ends, go live happens, and “later” quietly becomes never, because by the time anyone has bandwidth to revisit it, eleven thousand transactions are already posted against it and forty three reports are already built on top of it.

This is how technical debt actually accumulates in an ERP. Not through one bad decision, but through a hundred reasonable ones that never got revisited. Nobody sat down and decided Contoso should have four overlapping cost center segments. Somebody added one during the original build, somebody else added a second during a scope change nobody documented, and a third arrived during the failed second go live attempt when a different consultant didn’t know the first two existed.

The dangerous part isn’t that the structure is messy. Messy is survivable. The dangerous part is that nobody currently alive on the project understands the whole thing well enough to safely change it. That’s not a data model problem anymore. That’s an institutional memory problem, and it’s the same failure mode as the shadow ledger from episode two, just wearing a different disguise.

Sable’s instinct here is the right one: map around it before you touch it. You don’t fix an archaeology site by bulldozing it. You document what’s actually there first, then decide what’s safe to change.


What happens next

The party has a working map of the chart of accounts and a firm rule not to touch TEMP_DO_NOT_USE. But mapping the dimension structure surfaces the missing piece from episode one: the integration list from page seventeen of the SOW. In the next episode, the party finally gets their hands on it, and finds one integration on the list with no name, no owner, and no documentation at all.

Next episode: Episode 4, Page Seventeen (coming soon)


If your own chart of accounts has a TEMP_DO_NOT_USE hiding in it somewhere, you’re not alone, and you’re not imagining the risk. Gamifying the Enterprise: Game Mechanics for Continuous Proficiency is available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want a real framework for documenting and untangling structures like this one before they become load bearing, start here: adnd365.com/start

Cause of death: the case study was true, and that’s exactly the problem.


Two-thirds of the way through the deck, a new logo appears. A real one, a company you’ve heard of, sometimes a competitor’s supplier or a name from your own industry vertical. The slide has a number on it, usually a big one: forty percent reduction in close time, three million recovered in duplicate payments, six months to positive ROI. Underneath the number is a quote, attributed, sometimes even video, from a real person who really said those words.

Nothing on that slide is fabricated. That’s what makes this one the hardest autopsy in the series. The other demos in this blog die from omission, pacing, or seamlessness hiding a seam. This one dies from something subtler: a true statement about one company, presented in a context engineered to make you believe it’s a claim about yours.

What actually happened

The reference customer on the slide is not a random sample. It is, almost by definition, the single best outcome the vendor has produced across their entire installed base, selected specifically because the number is large and the customer is willing to say it out loud. Somewhere behind that slide are dozens or hundreds of other implementations that landed closer to the median, plus a smaller number that struggled or stalled, none of which get a logo or a quote, because nobody puts “we got most of the way to the business case, eventually, after two scope changes” on a slide.

There’s also a matching problem the case study never surfaces. The reference customer’s forty percent reduction in close time happened inside a specific starting condition: a particular level of process maturity, a particular data quality baseline, a particular willingness internally to change how work got done. The case study tells you the outcome. It almost never tells you the starting line, and the outcome without the starting line is not a number you can subtract your own situation from.

The quote does real work here too. A specific named person saying a specific thing on camera reads as harder evidence than an aggregate statistic, even though a single testimonial is a sample size of one, hand-selected from a population the vendor controls entirely.

Why it works on smart people

Humans are wired to trust specific, named, social proof more than abstract statistics, and this isn’t a flaw, it’s usually a reasonable heuristic. A named person willing to put their reputation behind a claim on camera is, in most contexts, more credible than an anonymous number. The problem is that the heuristic evolved for a world where the sample in front of you was roughly representative of the population, and a vendor-selected reference customer is the opposite of representative by construction.

There’s a second effect working alongside the first. By the time the reference slide appears, you’ve usually already sat through thirty or forty minutes of a demo that felt competent, so the case study isn’t landing on a skeptical audience, it’s landing on an audience that has already been primed to trust what they’re being shown. The reference customer isn’t doing the persuading alone. It’s the closing argument after the room has already been warmed up.

The actual damage

This is the one that turns into an internal expectations problem before it turns into a vendor problem. Someone in the room, often not maliciously, repeats the number in an internal steering committee deck as though it were a forecast rather than someone else’s outcome. “Similar companies have seen a forty percent reduction” quietly becomes “we’re targeting a forty percent reduction,” and by the time the project charter gets written, a single best-case data point from a different company, with different starting conditions, has become your project’s success criteria.

When your actual results land closer to the median, which is where most results land by definition, the project doesn’t get judged against a realistic baseline. It gets judged against the reference customer’s outcome, which nobody on your team ever should have agreed to as the target in the first place.

The fix, if you’re the one presenting

Show the range, not just the peak. If you have a reference customer at forty percent, say what the twenty-fifth and seventy-fifth percentile outcomes look like too, and say why the reference customer landed where they did, what was true about their starting point that might or might not be true about the prospect’s. A specific, named case study is still worth showing. It’s worth showing better, with its context attached, instead of as a number floating free of the conditions that produced it.

The honest version of that slide is less dramatic. It’s also the only version that survives contact with a steering committee eighteen months later.


This is the same shift I wrote about in The New Expert Isn’t the One With the Answers. Having the number was never the hard part. Knowing whether that number applies to your situation is.

An Advanced Dungeons & Dynamics 365 session. Episode 2 of 14.

Previously on: Episode 1, The Summons at Waterdeep Docks, the party arrived at Contoso Coffee Roasterie to find a missing scope document, a go live date that had quietly moved up by six weeks, and a whiteboard asking a question nobody could answer: where is the margin going.

Every implementation has an official system of record. And every implementation, if it’s been running long enough, has an unofficial one too. Usually a spreadsheet. Usually built by someone who got tired of waiting for the real numbers to be right.

That’s where episode two picks up.


The Session

GM: In the war room you meet three Contoso stakeholders: the CFO, the head of Ops, and a woman named Priya who nobody introduces by title.

Sable: (casting Data Mapping) I’m getting margin data. It’s coming from… this isn’t the ERP. This is a workbook.

Priya: That’s mine. I’ve been tracking real margin since the system doesn’t roll up freight variance correctly.

Marge: How long has this shadow ledger existed?

Priya: Since the first go live attempt. Fourteen months.

Thorne: (rolling a Perception check) The workbook total and the GL total are six figures apart.

Vex: Which one’s right?

Priya: I don’t know anymore. I leave for vacation tomorrow.

Ai. Cassiopeia: I would like it noted that I offered to reconcile this in October. The ticket was closed as “won’t fix, not urgent.”

The CFO leans over Thorne’s shoulder to look at the variance number. He goes pale in a way that has nothing to do with the lighting.

CFO: That number is bigger than our reported profit.

Nobody says anything for a moment. Priya starts quietly packing her laptop bag, the way someone packs when they’ve decided this is no longer their problem to solve, whether or not that’s actually true.

Priya: (standing) My flight’s at six tomorrow morning. I really am sorry.


Why the spreadsheet always wins

Every consultant who has done this work long enough has met a Priya. Somebody who didn’t set out to build a shadow system, they just needed the numbers to be right for a meeting, and the ERP wasn’t giving them right numbers fast enough. So they built a workaround. Then the workaround became a habit. Then the habit became fourteen months of institutional memory that lives entirely on one person’s laptop, reconciled by hand, understood by nobody else in the building.

This isn’t a story about a bad employee cutting corners. It’s what happens when a system doesn’t earn trust fast enough, and a business still has to make decisions in the meantime. Priya’s spreadsheet isn’t the failure. It’s evidence of a failure that happened somewhere upstream, probably long before she ever opened Excel.

The dangerous part isn’t that the spreadsheet exists. It’s that nobody in the room can currently say which number, the workbook or the GL, is actually correct. Six figures of ambiguity, and the person who understands the discrepancy best is getting on a plane in twelve hours.

That’s not a data problem anymore. That’s a continuity problem. And it’s the kind of gap that a proficiency system is supposed to catch long before it turns into a war room moment.


What happens next

Priya is gone by morning, and the shadow ledger goes with her, at least in terms of anyone who can explain it. The party has six figures of unexplained variance, a CFO who now can’t unsee the number, and a chart of accounts they haven’t even looked at yet. In the next episode, they go looking for where the discrepancy actually lives, and find a financial dimension that’s been quietly absorbing thousands of postings that were never supposed to exist.

Next episode: Episode 3, The Labyrinth of Financial Dimensions (coming soon)


A quick gut check for anyone running their own ERP: is there a shadow spreadsheet keeping your business honest right now? Gamifying the Enterprise: Game Mechanics for Continuous Proficiency digs into exactly why that happens, and how to design proficiency systems so it stops. Available on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want the actual configuration playbook behind scenes like this one, start here: adnd365.com/start

Cause of death: the feature that closed the deal was never actually in the room.


Somewhere around minute forty of the demo, the presenter hits a gap. The thing you actually asked about, the reason you took the meeting, doesn’t quite exist yet. What happens next is the tell. The slide doesn’t say “we don’t do that.” It says “coming in the next release,” said in exactly the same tone of voice as everything that already works, with exactly the same confident click-through pacing, so that by the time the meeting ends, the feature that doesn’t exist has fully merged in your memory with the fifteen features that do.

Nobody lied. That’s what makes this one interesting to cut open. The roadmap slide was real. The quarter listed on it might even be accurate, as of the day the deck was built. And yet the effect on the room is functionally identical to a lie, because a promise wearing a product demo’s clothing gets evaluated with a product demo’s scrutiny, which is to say, almost none.

What actually happened

Every roadmap item in a sales deck starts life as an engineering estimate, gets filtered through a product manager’s optimism, gets filtered again through a sales engineer who needs this quarter’s number, and arrives in front of you as a single, confident bullet point that has shed every unit of uncertainty it was born with. “Q3” meant “Q3, if the two prerequisite features land on time and nothing gets reprioritized” back at the whiteboard where it was written. By the time it’s read aloud in your conference room, it just means Q3.

The demo compounds this by never distinguishing, in pacing or tone, between the click that shows something real and the click that shows a mockup of something planned. Both get the same enthusiasm. Both get the same “and here’s where you’d.” The interface doing the showing doesn’t have a font for “this is a Figma file with a database connection painted on.”

You are, in effect, being shown two different products stitched into one seamless walkthrough: the one that ships today, and the one that exists only as a commitment on a slide, and you’re being asked to make one buying decision that covers both.

Why it works on smart people

Buyers are trained, correctly, to evaluate a vendor’s direction and not just their current state. Nobody wants to buy a system that solves today’s problem and ignores next year’s. So a roadmap conversation is a legitimate, necessary part of due diligence. The trick isn’t the existence of the roadmap. It’s the demo borrowing the roadmap’s credibility and lending it back to itself.

There’s also a timing problem working against you. The roadmap feature is almost always introduced as the answer to the exact gap you just identified in the product, which means it lands at the precise moment you’re feeling a little disappointed and looking for a reason not to be. “Coming in Q3” isn’t just information at that point. It’s relief, and relief is a bad state to be evaluating claims in.

The actual damage

This is the one that shows up on a signed contract with a footnote nobody reads until it matters. Somewhere a business case got built with the roadmap item load-bearing in it, sized as though it were a current-state capability, because in the meeting it felt like one. The actual purchase decision, the one with budget and a signature attached, priced in a feature that was, at signing, a Jira ticket with a target quarter next to it.

Q3 arrives. The feature either doesn’t ship, ships in a reduced form that solves half the original problem, or ships correctly but a year later, after a reprioritization nobody outside the engineering org heard about. Your business case, however, was built on the version of the feature that existed only in the demo room, and now someone has to explain to their own leadership why the thing everyone signed off on isn’t the thing they got.

The vendor isn’t necessarily acting in bad faith here. Roadmaps genuinely slip, for genuinely defensible reasons. But “the vendor wasn’t lying” is cold comfort to the person holding a business case that assumed a delivery date as fact.

The fix, if you’re the one presenting

Change the font, literally or figuratively, the instant you cross from shipped to planned. A different slide background, a verbal flag, a pause, anything that makes the seam audible. Say the confidence level out loud: “this is committed and in QA,” versus “this is prioritized but not yet started,” versus “this is directionally where we’re headed and I wouldn’t bet a contract on the date.” Those are three different products. Let the buyer evaluate them as three different products.

It costs you a little bit of momentum in the room. It buys you a customer who signs with accurate expectations, which is the only kind of customer who’s still happy with you eighteen months later.

The roadmap wasn’t the lie. The seamlessness was.


I’ve argued elsewhere that no self-respecting architect leaves the scaffolding up once the building is done. A roadmap slide is the one place I’d argue for the opposite: leave the scaffolding very visible, since half of what’s on screen hasn’t been built yet.

Cause of death: nobody asked what the prompt actually did.


The setup is always the same. The presenter opens a chat panel next to the application. Types a sentence, something like “create a purchase order for our top vendor and route it for approval.” Hits enter. Three seconds later, a purchase order exists, fully populated, correctly routed, and the room exhales the specific kind of gasp that sales decks are built around.

Nobody in the room asks the only question that matters: what happened between the sentence and the purchase order.

That’s the autopsy. Not whether the AI worked. It worked. The patient died of the room’s collective decision not to look inside the box.

What actually happened

There are, broadly, three things that could have occurred in those three seconds, and they have wildly different implications for what you’re buying.

The prompt could have triggered a genuinely flexible model reasoning over your live data and your actual configuration, in which case, remarkable, and worth every follow-up question you can throw at it. Or it could have matched against a narrow, pre-built intent, one of a finite list the vendor trained and tested specifically for this demo, in which case the “AI” is doing roughly what a well-labeled button would do, dressed in a chat window. Or, and this is the quietly common one, it could have worked because the demo environment was built so there was exactly one top vendor, exactly one approval path, and exactly one plausible interpretation of the sentence, which means the model didn’t have to be smart. It had nothing to be confused by.

You cannot tell these three apart from the audience seat. That’s not an accident. It’s the format working as intended.

Why it works on smart people

Chat interfaces borrow credibility from every other chat interface you’ve ever used. You already know how to type a sentence and get a reasonable response, because you do it daily with a general-purpose assistant that genuinely is flexible. The demo imports that trust wholesale and applies it to a narrow, scripted intent recognizer that shares none of the underlying capability.

There’s a second layer, too. Asking “what exactly did the model do” mid-demo feels like a strange thing to interrupt with, technical in a way that makes you look like you’re missing the point of the show. So the room lets it go, same as it lets the Golden Path Demo’s pristine data go. The social cost of asking is higher than the informational value anyone expects to get from asking, so nobody does, and the vendor never has to answer a question nobody asked.

The actual damage

This is the one that costs real money, because “AI does it for you” gets sized into the business case. Someone in procurement multiplies the demo by the number of purchase orders your company cuts in a year and arrives at a headcount reduction that made it into a slide before a single production prompt had been run against a real vendor master with your actual data quality problems in it.

Then go-live happens, and the model that flawlessly handled “our top vendor” in the demo turns out to need a disambiguation step for the fourteen vendors named some variant of “Smith Industries” in your real vendor table, and the “flexible AI agent” turns out to be a narrow intent match against six pre-built scenarios, and scenario seven is where your actual business lives.

The gap between what was promised and what shipped doesn’t show up as a bug. It shows up as a quiet, expensive redefinition of what “handled by AI” meant, discovered by whoever now has to explain the missed headcount number.

The fix, if you’re the one presenting

Show your work, on purpose, before anyone has to ask. After the prompt returns its clean result, pull back the curtain for ten seconds: here’s the intent it matched, here’s the data it pulled from, here’s what happens if the vendor name is ambiguous. If the honest answer is “the model reasoned over live data,” say that, and let it be more impressive for being specific. If the honest answer is “this is a pre-built scenario for the categories we support today,” say that too. A buyer who knows exactly what they bought doesn’t come back angry in month four. A buyer who assumed general intelligence and got a narrow, well-executed lookup does, and they remember whose demo sold it to them.

The trick was never the model. It was the silence right after it worked.


If the “narrow intent versus genuine reasoning” distinction sounds familiar, it’s the same fault line I dug into in Is Headless ERP Enough, or Just a Step in the Right Direction?, on what it would actually take for AI to be the architecture instead of a coat of paint on top of it.

An Advanced Dungeons & Dynamics 365 session. Episode 1 of 14.

Every consulting engagement has an origin story. Most of them start with a kickoff call, a signed SOW, and someone on the client side saying “we’re really excited to get started” in a tone that suggests they are not.

This one starts at the docks, at dawn, with a folder that’s still warm from the printer.

Welcome to the Contoso Convergence. This is the campaign we’re running here on the blog: a full Advanced Dungeons & Dynamics 365 session, told in fourteen episodes, following a five person party sent to rescue a D365 Finance & Operations go live that has already failed twice. If you’ve ever lived through a stalled implementation, you’ll recognize the dungeon. We’ve just given it a party and some dice.

Here’s the party, for reference:

  • Thorne Ledgerkeep, Paladin (Functional Consultant, Finance). Cannot tell a lie about a trial balance. Weak against scope creep.
  • Vex Nullpointer, Rogue/Artificer (Technical Consultant). Picks locks on legacy integrations. Allergic to undocumented customizations.
  • Sable Query, Wizard (Solution Architect). Casts Data Mapping at range. Vulnerable to “just one more report” requests.
  • Captain Marge Dunwell, Fighter (Project Manager). Frontline tank. Immune to fear, mostly immune to steering committees.
  • Ai. Cassiopeia, grey collar party member, an AI agent bound to the Wayfinder’s Network. Casts Copilot Suggestion. Cannot be blamed in the retro, which everyone resents.

They work for the Waterdeep Trading Company. Their assignment: Contoso Coffee Roasterie, a mid market roaster whose D365 go live has now stalled twice, and whose current state can be summarized by one line on a whiteboard that nobody has erased.

Let’s begin.


The Session

GM (Dessa, dispatching the party): Contoso Coffee Roasterie. Third go live attempt. The first two consultants who touched this engagement now work in agriculture. You leave at dawn.

Thorne: What’s the scope?

GM: That’s the thing. Here’s the SOW. It’s forty one pages. Page seventeen is missing.

Vex: (flipping pages) Page seventeen is always the page with the integration list on it. Always.

Sable: I’ll cast Data Mapping when we arrive. Can’t do it blind.

Marge: Team, standard rules. We document the golden path, we do not perform the golden path. If Contoso’s demo deck has a slide that says “and then the order just flows through,” that slide is lying to us.

Ai. Cassiopeia: I have ingested the prior two consultants’ status reports. Both end mid sentence.

The party arrives at Contoso’s gates. The receptionist hands Thorne a folder. It is warm, as though recently printed in a panic.

Receptionist: They’re expecting you in the war room. Also, go live is in six weeks.

Thorne: The SOW says twelve.

Receptionist: That was the old go live date.

The party is led down a hallway that smells faintly of burnt coffee and cold pizza. Someone, at some point, taped a printed banner to the wall that reads “WE ARE ALMOST THERE.” Someone else has crossed out “ALMOST” and written “NOT” above it in marker.

The war room door creaks open. Inside, a whiteboard covered in red ink reads only:

“WHERE IS THE MARGIN GOING.”

No question mark. Nobody in the room has energy left for punctuation.


Why this scene, why first

Every stalled implementation has a version of this moment: the point where a new team walks in and the first thing they learn is that the paperwork doesn’t match the reality on the ground. A missing page seventeen. A go live date that’s moved without anyone updating the SOW. A room that’s been living inside the same unanswered question for months.

None of that is dysfunction, exactly. It’s what happens when a project runs long enough that the documentation stops keeping pace with the truth. The SOW says twelve weeks because that’s when someone last had time to update it. The scope says one thing because rewriting it would mean admitting how much has drifted. Nobody’s lying. Everybody’s just too busy surviving the current sprint to fix the paper trail.

Marge’s line matters here, and it’s the whole thesis of this campaign: document the golden path, don’t perform it. A demo that only shows the happy path teaches the team nothing about the exceptions they’ll actually live in. That’s not a throwaway rule for a fictional party. That’s the difference between a go live that holds and a third failed attempt.

We’ll come back to that idea more than once over the next thirteen episodes.


What happens next

The party has a whiteboard, an unanswered question, and a go live date that just got six weeks shorter. In the next episode, they meet the Contoso stakeholders, including a woman named Priya who nobody introduces by title, and the party learns that somebody has been quietly running the real numbers in a spreadsheet the entire time.

Next episode: Episode 2, The Shadow Ledger (coming soon)


If any of this feels familiar, that’s on purpose. This campaign is built on the same framework behind Gamifying the Enterprise: Game Mechanics for Continuous Proficiency, available now on Amazon: https://www.amazon.com/dp/B0GY3VWLVX

And if you want to run something like this for your own team, not the dragons, the actual configuration, the AD&D365 guides are the place to start: adnd365.com/start

Nobody has ever finished a day of data entry in an ERP system and felt like they’d been playing a game. That’s the problem gamification tries to solve, and after years of poking at enterprise systems, I’ve become convinced it’s one of the more underrated levers for actually getting people to use the software correctly.

The pitch

Gamification means borrowing the mechanics that make games compelling: points, badges, levels, progress bars, leaderboards, and bolting them onto tasks nobody would otherwise choose to do carefully. In an ERP context, that might mean a purchasing clerk earning a badge for zero-error PO entry for a month, a warehouse team seeing a live leaderboard of pick accuracy, or a new hire working through a “level up” onboarding path instead of a 40-tab training binder.

It’s not a gimmick dreamed up by a UX consultant with too much time on their hands. There’s real academic backing here. A well-cited study built a gamification prototype on top of SAP ERP and tested it with 112 users using the standard technology acceptance model; enjoyment, flow, and perceived ease of use all improved meaningfully. Another case study found that adding game mechanics to SAP increased user “telepresence” (basically, how engaged people felt while using the system) by nearly 30%. The underlying research consistently shows gamified ERP leads to better data entry and fewer errors, which, if you’ve ever had to clean up a mangled inventory count, is not a small thing.

Why now

The gamification market broadly is expected to roughly double by the early 2030s, and enterprise software is a big part of that growth. What’s changed recently is the mechanism. The old playbook was static: points, badges, a leaderboard bolted onto the sidebar, forget about it. The new playbook is AI-driven, with personalized nudges, dynamic feedback loops, and coaching that adapts to what an individual user is struggling with rather than a one-size-fits-all reward ladder. Microsoft’s Power Apps approach is a good example of the direction things are heading, embedding game-like mechanics directly into workflows rather than treating gamification as a bolt-on layer, which cuts rollout time from months to weeks.

HR and training modules are seeing the fastest uptake, which makes sense. That’s the part of ERP most people already expect to feel like a course rather than a chore, so it’s the easiest wedge for game mechanics to get in the door.

The catch

Here’s the part worth sitting with before you get excited and start slapping badges on every screen: a huge share of gamification efforts flop. The research puts the failure rate at around 80% when organizations default to generic points and leaderboards without actually designing for the behavior they want to change. A leaderboard that just measures raw transaction volume will train people to enter data fast and sloppy, not accurately. Badges nobody respects become wallpaper. And gaming mechanics don’t land the same way with every personality; some people are motivated by competition, some by mastery, some find the whole thing patronizing. The smart implementations keep traditional training and recognition paths alongside the gamified ones rather than replacing them outright.

Legacy systems are also a real drag here. If you’re still running SAP ECC or an older on-prem instance, bolting gamification on top usually means custom middleware, which stretches timelines and adds a maintenance burden nobody budgeted for. It’s a much easier build on modern cloud ERP with decent APIs.

The tinkerer’s takeaway

If I were experimenting with this on a real system today, I’d start narrow. Pick one painful, error-prone workflow, define the specific behavior I actually want to reinforce (not just “more activity”), and build a small feedback loop around that: a progress indicator, a streak counter, something visible and honest. Skip the company-wide leaderboard until you’ve proven the mechanic works on a small team that won’t quietly resent it.

I actually went deep enough down this rabbit hole to write a book about it: Gamifying the Enterprise: Tabletop Mechanics for ERP Training, Continuous Education, and User Proficiency Rating. It digs into how tabletop game design principles (the kind of thing you’d find in a board game rulebook, not a mobile app) can be adapted for ERP training and ongoing user proficiency.

ERP software has a reputation for being where enthusiasm goes to die. Gamification isn’t going to fix bad process design or a system nobody wanted in the first place, but done with a little more thought than “add badges,” it’s a genuinely useful tool for making the boring but important parts of enterprise software a little more bearable.

If you want help thinking through where gamification actually fits in your own ERP rollout, that’s exactly the kind of thing I help people work through. Get in touch at adnd365.com/start.

I’ve been building and implementing ERP systems for a long time. Most of them share the same skeleton underneath: a monolithic database that owns the authoritative state, application code layered on top to enforce business rules, and a UI wrapped around the whole thing so humans can interact with it.

That pattern has served enterprise operations well for thirty years. It’s proven, stable, and deeply integrated into how organizations run.

Headless ERP Is Real Progress – and a Useful Stepping Stone

The industry’s shift toward headless ERP is a genuine improvement. The pitch is compelling: decompose the monolith, expose business capabilities as APIs, and let any interface – web, mobile, AI copilot – consume them. This breaks the tight coupling between the UI and the data, which creates real flexibility.

But it’s worth being clear about what headless ERP changes and what it doesn’t.

The database is still the authoritative system of record. The business logic is still baked into the same application layer. What’s changed is that the user interface abstraction has moved one level up the stack. An AI agent now calls the same procedure that a form used to call. The underlying state management, the write model, and the integrity model are largely the same.

That’s a meaningful evolution. It’s just not the same thing as rethinking the architecture from the ground up. The major ERP vendors are building AI agents, tool APIs, and conversational interfaces as fast as they can – and those capabilities are genuinely valuable. But in most cases, the mutable relational database remains the authoritative core. AI is a new interaction layer, not a new architecture.

What Would a Truly AI-Native ERP Look Like?

That raises a useful design question: if we weren’t constrained by the existing architecture, what would we actually need an ERP to do?

An ERP needs to:

  1. Record that business events happened and cannot be undone
  2. Derive current state from those records
  3. Enforce rules about which events are permitted
  4. Report on any slice of history or current state
  5. Coordinate with other parties and nodes

Notice that “store mutable rows in a relational schema” is not in that list. That’s an implementation choice that has become a deeply embedded assumption – but it isn’t the only one available.

What if the ledger was the database?

Not a blockchain. Not distributed consensus. Not tokens. Just immutable, append-only, typed records – structured text that any human or machine can read, that hashes itself into a verifiable chain, and that never requires a rollback because facts don’t change, only new facts get added.

I’ve been building exactly this. The prototype uses plain Markdown files as the authoritative store. Every business posting – a journal entry, an inventory movement, a document – is a Markdown file containing machine-canonical JSON and human-readable context. There is no database. Current state is a projection, rebuilt from the ledger on startup and kept live in memory. The only write is an atomic append.

Fifty-seven passing tests cover balanced accounting, inventory authority, immutable document chains, verified replication, disconnected-node convergence, and same-origin fork rejection. The balance sheet balances. The inventory reconciles. Two disconnected nodes post independent transactions and converge without conflicts – without any consensus protocol.

AI Becomes the Only Interface

Here’s where the architecture shift becomes genuinely interesting.

In a ledger-first, text-native ERP, there is no form to fill out. There is no screen to navigate. There’s a typed command interface and a set of hard invariants. An AI agent – or a human typing natural language – submits an intent. The system validates it against policy, checks the invariants, and either appends an immutable record or returns a precise rejection.

The UI doesn’t exist until it needs to exist. When a warehouse manager asks “what’s on hand at site B?”, the answer is a live projection query. When a CFO asks “show me the aging receivables by customer segment”, that’s a reporting query – constructed dynamically against the ledger, not a pre-built report someone maintained for years. When a purchasing agent says “draft a PO for the coffee beans I bought last quarter at the same price,” the AI reads the ledger history, constructs the command, and submits it for approval.

There’s no screen configuration. There’s no report builder. There’s no workflow designer. The AI is the interface, the report, and the workflow – and the ledger makes sure it can’t lie about the numbers.

Policies Are the Contracts, Not Code

The hardest part of any ERP implementation isn’t the software. It’s the rules: approval thresholds, posting profiles, accounting mappings, regulatory requirements, credit limits. In traditional ERP these are baked into configuration tables, workflow definitions, and custom code – all of which require IT to change and none of which an AI can inspect or reason about directly.

In a ledger-first design, policy is a first-class ledger record.

An approved policy revision is itself immutable, effective-dated, and hash-linked. It says: “For purchase orders above $10,000, two approvals are required, and the expense must post to account 6100.” That policy record is readable by both the enforcement engine and an AI agent explaining a rejection. When the CFO updates the threshold, a new policy revision is posted to the ledger – the old one doesn’t disappear, it simply becomes historical. You can audit every rule change the same way you audit every financial transaction.

The objects – products, customers, accounts, locations – are extensible definition records, not schema columns. Adding a new attribute to a product doesn’t require a database migration. It’s a new field in the next definition revision. The schema is the business model; the ledger stores the evolution of the business model alongside the business events.

A Different Starting Point

This model raises questions worth sitting with, especially for architects and business leaders thinking about long-term platform choices.

If the authoritative store is a text ledger that any AI can read directly, the relationship between data and the systems that manage it changes. If policy is data rather than embedded configuration, it becomes inspectable and auditable in ways that configuration screens aren’t. If the UI is assembled dynamically rather than maintained as a static artifact, the cost of adapting to change drops.

Existing ERP platforms are adding AI agents, tool APIs, and conversational interfaces at a rapid pace – and those are genuine improvements on top of proven foundations. The question this architecture raises isn’t whether those platforms have value. They clearly do. It’s whether the database-first foundation is the only viable path, or whether there are scenarios where a ledger-first approach offers distinct advantages.

What I’m describing is a different starting point – not a replacement for everything an ERP does today, but an exploration of whether the database-first model is the permanent default or one point on a longer architectural trajectory.

What This Isn’t

To be direct about what I’m not claiming:

This is a prototype. It has 57 passing tests, not 57,000 customers. It doesn’t have digital signatures, authenticated transport, encryption, crash recovery, or performance benchmarks at scale. The hard work of production ERP – currencies, taxes, period close, consolidation, regulatory certification – hasn’t been done.

I’m also not claiming that organizations should abandon well-functioning ERP implementations. For continuously connected operations under one administrative authority, a traditional ERP database remains a strong and well-understood answer.

But for organizations that need:

  • Independent nodes that operate offline and converge later
  • AI governance with a hard execution boundary and auditable decisions
  • Extensible business objects without vendor schema lock-in
  • Deterministic history reconstruction – what was true on day one through today
  • Policies as auditable data – not configuration only IT can change

…the ledger-first, AI-native model deserves serious evaluation. It’s not a theoretical construct. The balance sheet balances. The inventory reconciles. The tests pass.

The Question Worth Asking

The ERP industry’s investment in AI capabilities is accelerating. In most implementations today, those capabilities are built as interfaces to the existing architecture. The database is still authoritative, and AI is the new interaction layer.

What if the AI wasn’t the client – what if the ledger was, and the AI was one more authorized agent submitting typed, policy-governed commands alongside the humans?

That’s not headless ERP. That’s a different architecture.


I’m writing this as someone who has implemented ERP systems commercially and is now prototyping an alternative architecture – not to sell a product, but to test whether the assumptions we’ve been building on are as permanent as we’ve treated them.

The source code, architecture specification, and worked scenarios are in an active research project. Happy to dig into specifics in the comments below.

For years, expertise was measured by how much you knew.

The best ERP consultants could recite configuration settings from memory. They knew which parameters to enable, where to find the obscure options, and how one setting quietly broke another three modules over. Experience meant accumulating answers.

AI changes that.

Today, the answer is often seconds away. Ask an AI how to configure inventory dimensions, build a procurement workflow, design a chart of accounts, or set up a security role, and you’ll get a reasonable starting point. Often it will just generate the configuration itself.

So does that make years of ERP experience less valuable?

Quite the opposite.

The Skill Has Shifted

The value was never really in typing values into a form. It’s in knowing what the business actually needs, translating that into the right request, catching it when AI makes the wrong assumption, and steering it back to the correct solution.

Think of AI as the world’s fastest configuration specialist. Tell it exactly what you want, and it will produce an impressive amount of work in very little time. Give it a vague or incomplete request, or one built on a misread of the business, and it will build the wrong thing just as fast. Faster, in fact, than any consultant ever could.

The consultant’s role hasn’t disappeared. It has moved up a level.

Instead of asking “How do I configure this?” we’re asking “What should this business process look like?”

Instead of worrying which checkbox to select, we’re deciding why that checkbox should exist in the first place.

That’s a much harder problem.

A Procurement Example

Take Dynamics 365 Finance & Supply Chain. Suppose a company wants to automate purchasing.

A few years ago, a consultant would spend hours clicking through forms: building workflows, defining approval hierarchies, configuring procurement policies, testing every scenario by hand. Today, AI can generate much of that scaffolding.

But someone still has to answer:

  • Should approvals run by department, cost center, project, or spending threshold?
  • Which purchases should bypass approval entirely?
  • How are emergency purchases handled?
  • What happens when an approver is out on vacation?
  • Which controls satisfy audit requirements without slowing the business down?

Those aren’t configuration questions. They’re business questions. And business questions have always been where experienced consultants create the most value.

The Pattern Is Everywhere

Developers spend less time writing boilerplate and more time describing the architecture they want. Data analysts spend less time writing SQL and more time deciding which questions are worth asking. Architects spend less time drawing every line and more time defining the building.

Doctors increasingly have AI helping interpret medical images, but the physician still decides which tests to order, how to read the results in context, and what treatment actually fits the patient.

The profession doesn’t disappear. The center of gravity moves.

Prompting Is Requirements Gathering, Rebranded

This is why prompt writing gets misunderstood. People treat it as clever wording. It isn’t.

Good prompts are evidence of good thinking. A well-written prompt reflects a clear grasp of the business objective, the constraints, the desired outcome, and the tradeoffs involved. The better you understand the problem, the better your prompt becomes.

In many ways, prompting is just the modern version of requirements gathering. The consultant who asks twenty thoughtful questions before asking AI a single one will consistently outperform the consultant who jumps straight to generating configurations.

Where This Leaves the Next ERP Expert

The future ERP expert will spend less time configuring systems and more time shaping solutions. That means curiosity beats memorization. Business knowledge beats navigation skills. Judgment beats mechanics.

AI is making execution cheaper. It’s making thinking more valuable.

The companies that win won’t just have access to better AI. They’ll have people who know what to ask, why they’re asking it, how to challenge the results, and when to change direction.

The next generation of ERP consultants won’t be measured by how fast they can configure a system. They’ll be measured by how well they can define the problem AI is being asked to solve.

Because in the age of AI, the competitive advantage isn’t having all the answers. It’s asking the questions that lead to the right ones.

In 1899, the Norwegian mathematician Niels Henrik Abel complained about Carl Friedrich Gauss’s papers. Gauss’s proofs were airtight but gave no hint of how he’d gotten there. Abel said Gauss was “like the fox, who effaces his tracks in the sand with his tail.” Gauss’s reply has outlived the complaint: “No self-respecting architect leaves the scaffolding in place after completing the building.”

He wasn’t being cagey. Gauss ran through dozens of false starts for every theorem he published, and he thought showing that mess would do the reader a disservice. A finished proof is not a transcript of how the proof was found. It’s a clean path from A to B, built after the fact, once you already know where B is.

This is exactly what happens in a good product demo, and it’s why demos so often feel like a kind of dishonesty even when every word in them is true.

The gap between finding and showing

When you’re building something hard, most of the time is spent wrong. You try an approach, it breaks in a way you didn’t expect, you try another. Six weeks in, you’ve got a whiteboard full of dead ends and one narrow path that actually works. Then you build the demo, and the demo only shows the narrow path. Someone watches it and sees three clicks. They have no way of seeing the six weeks.

That’s not a flaw in the demo. A demo that included the six weeks would be unwatchable, and it would also be the wrong artifact. Nobody wants the scaffolding. They want the building.

But it does create a predictable failure mode: the audience underestimates what it took, and then either assumes it must have been easy (so why did it take so long) or assumes it must be fragile (since nothing that clean survives contact with reality). Both reactions come from the same place: they’re reacting to the absence of visible effort, not to the actual state of the work.

Related shapes of the same idea

This tension shows up everywhere people do hard work and then have to present it simply:

Einstein’s Zurich Notebook, from 1912–13, is page after page of failed derivations. He got within a hair of general relativity, talked himself out of it for the wrong reasons, and lost three years before arriving at the fifteen or so symbols that appear in the final field equations. Nobody reads the field equations and sees the three lost years.

There’s an old story, probably apocryphal, about the engineer Charles Steinmetz, called in to fix a broken generator no one else could diagnose. He listened to it, made one chalk mark on the housing, and told them to replace that part. His bill was ten thousand dollars, itemized: “making the mark, one dollar; knowing where to put it, nine thousand nine hundred ninety-nine dollars.” The mark is the demo. The knowing is the scaffolding.

And there’s the duck on the pond: gliding, apparently motionless, paddling hard just under the surface where no one’s looking. It’s a cliché precisely because it describes something true about most visible competence.

What to do with this, if you’re the one giving the demo

You’re not obligated to show your work, and you shouldn’t try to cram it in. But it’s worth occasionally naming that it exists, even in one sentence: “this took longer than it looks like it should have” or “we tried three other approaches before this one.” Not as a boast, and not as an apology. Just as a way of telling the audience that the gap they’re sensing is real, so they don’t have to fill it in themselves with a worse story, like it must have been trivial, or it must be held together with tape.

Gauss removed his scaffolding because he trusted the building to stand on its own. That’s usually the right call. The failure isn’t leaving the scaffolding down. It’s letting people assume there was never any scaffolding at all.