Autopsy #15: The Broken Windows Demo

Cause of death: the small stuff got waved off as cosmetic, and cosmetic was never what it was actually signaling.


Partway through the demo, something small is visibly off. A field labeled “Custmer Grp” with the typo still in it. A report total that’s correct but formatted with the wrong currency symbol. An old menu item labeled “Legacy Approval, do not use” still sitting in the navigation where everyone can see it. Someone in the room notices, points at the screen, and the presenter waves it off without breaking stride. “That’s cosmetic, we’ll clean that up before go live, doesn’t affect the numbers.” The room nods and moves on, because the thing being demonstrated, the actual functional capability, did in fact work correctly.

Nobody asks the question that actually matters here, which isn’t “does this affect the numbers.” It’s “what does leaving this visible tell the two hundred people who are about to start using this system every day.” That’s the autopsy. The defect itself was never the problem. What it signals is.

What actually happened

Broken windows theory, originally a criminology idea from the early nineteen eighties, argues that visible, unaddressed signs of disorder, a broken window left unrepaired, graffiti left up, don’t just reflect neglect, they actively invite more of it, because their presence signals that nobody is watching and nothing gets enforced. The mechanism isn’t about the broken window causing damage on its own. It’s about what an unrepaired window communicates to everyone who sees it about whether this place has an owner who cares.

The same mechanism runs through a go live environment with unusual precision. A typo in a field label, a leftover legacy menu item, a report that’s functionally correct but visually sloppy, none of these break anything on their own. What they do is tell the first cohort of real users, on day one, before anyone has formed a habit yet, that quality here is negotiable. A user who sees an obviously wrong field label on their first day draws a reasonable, entirely rational inference: if that got through, my own sloppy data entry probably will too. Nobody enforces the small stuff here. The inference isn’t about that one field. It’s about the whole system’s apparent tolerance for error, and it gets formed in the first few days, before the system has any track record to override it.

Small defects waved off in a demo compound specifically because they arrive stacked. One typo alone probably wouldn’t shift anyone’s behavior. A demo that’s waved off three or four small things in a row, cosmetic each time, individually defensible each time, has quietly established a pattern before go live even happens.

Why it works on smart people

Triage is a genuinely good practice, and drawing a line between functional bugs and cosmetic ones is a reasonable way to prioritize limited time before a deadline. The instinct to wave off a typo isn’t wrong on its own terms. Where it goes wrong is in applying an engineering severity framework, does this affect the calculation, to a question that was never actually about calculation. Users don’t experience a system the way an engineer categorizes a bug ticket. They experience it as a single continuous impression of whether the place is cared for, and that impression doesn’t sort itself into functional and cosmetic buckets the way a backlog does.

There’s also a time pressure effect specific to go live weekends. By the time anyone’s looking at a stray legacy menu item or a mislabeled field, the team is usually exhausted, focused on the handful of things that could actually cause a financial misstatement or a blocked transaction, and a cosmetic issue genuinely does rank lower by every reasonable prioritization framework available in that moment. The framework is correct. It’s just answering a different question than the one that actually determines early adoption behavior.

The actual damage

This is the one that shows up as a slow erosion in data quality that nobody can point to a single cause for. Three months after go live, free text fields that were supposed to use standardized values are full of inconsistent entries. A workaround process has formed around the one screen that still has a visibly awkward layout, because early users concluded, correctly, that nobody was watching that screen closely. None of this traces back cleanly to the typo in the field label from the go live demo, but the typo, and the two or three cosmetic issues that sat alongside it, unaddressed, in the very first days anyone used the system, set the tone that made the rest of it feel acceptable.

The remediation cost here is unusually high relative to how small the original defects were, because by the time the erosion is visible enough to act on, it’s a data quality and behavior problem spread across the user base, not a five minute fix to a field label. Fixing the typo now costs the same five minutes it always would have. Fixing six months of inconsistent data entry that grew out of the signal the typo sent does not.

The fix, if you’re the one presenting, or the one buying

Fix the visible small stuff before go live specifically because of what it signals, not because of whether it moves a number. If a genuine deadline forces a trade off, say so honestly and specifically to the people who’ll be using the system, “we know this field label is wrong, it’ll be fixed by Friday, here’s how to report anything else you notice,” rather than letting it sit silently and be discovered on its own. The difference between an acknowledged, tracked flaw and a silent, ignored one is exactly the difference between an owner who’s watching and one who isn’t, and users read that difference correctly, every time.

The typo was never going to break a journal entry. What it broke, quietly, was the first impression of whether anyone was going to notice if they didn’t get their own entries right either.


Standardization is exactly the governance broken windows theory argues for, and it’s worth asking what happens when it erodes at scale, not just in one field label. In Is Vibe Coding the End of Standardized ERP?, I look at what happens when every business starts generating its own bespoke operating logic with no shared standard to enforce consistency. A thousand small ungoverned deviations, each individually reasonable, is the broken windows problem running at the scale of an entire industry instead of one field label.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.