· Systems · 9 min read
When Blameless Became Meaningless
The blameless post-mortem was one of the best ideas in operational culture. Then we kept the word and threw away the meaning — and turned incident reports into horoscopes with a severity rating.

Somewhere along the way, we took one of the best ideas in operational culture and sanded it down until it couldn’t cut anything anymore.
The blameless post-mortem was supposed to fix a real problem. People hide failures when failures get them punished. Hide the failure and you lose the lesson, so you end up paying for the same outage twice — once when it happens, and again when it happens to someone who never heard about the first one. The fix was elegant: stop punishing people for honest mistakes, and they’ll tell you the truth. Truth in, learning out. Beautiful.
Then the industry did what the industry always does with a good idea. It kept the word and threw away the meaning.
Blameless doesn’t mean nameless
Here’s the confusion sitting at the bottom of most modern post-mortems: people think blameless means you’re not allowed to say who broke it.
It doesn’t. Blameless means you won’t punish the person who broke it. Those are completely different promises, and we quietly swapped the hard one for the easy one.
Not punishing someone requires a culture. It requires managers who can look at a person who just took down production and say “okay — what did the system let you do, and how do we make sure the next person can’t” — and mean it, and have it hold at performance review six months later. That’s expensive. That takes years of trust to build and one bad calibration meeting to burn.
Not naming someone, on the other hand, requires a find-and-replace. So that’s the one we shipped.
But you cannot get to a good fix without the sentence you’re not allowed to say. “The deploy took down checkout” is not the useful sentence. The useful sentence is “Kathy pushed a config change that took down checkout, because nothing in the pipeline told her checkout depended on it.” The name isn’t there to shame Kathy. The name is there because the moment you erase it, you also erase how it happened — and how it happened is the entire point. Scrub the actor and you’re left with “an incident occurred,” which is the operational equivalent of “mistakes were made.” Nobody learns anything from the passive voice.
And it’s worse than dull, because the passive voice doesn’t just hide the actor — it changes what kind of event you’re looking at. “Kathy pushed a bad config” is a human action, and human actions have causes you can fix. “An incident occurred” is weather. It’s a coincidence, an act of God, a law of nature — and you don’t file action items against the laws of nature. The moment the sentence loses its subject, the outage stops being something we did and becomes something that happened to us, and you can already feel the room relax. Nobody’s fault. Nothing to fix. Just one of those things.
And there’s the first logic trap, sitting right there in the grammar. Turning “Kathy made a mistake” into “a mistake was made” doesn’t just delete a name. It deletes agency. And once you’ve told agency to get out of the room, you don’t get to be surprised when it declines to come back five minutes later, exactly when you start talking about fixes. You can’t build a remediation on top of a cause you’ve grammatically erased. There’s nothing there to remediate. You removed the hand that touched the thing, so all that’s left is the thing, sitting there, having apparently broken itself.
So the first way blameless becomes meaningless is simple: we protect the person by deleting the facts, and then wonder why the report reads like a horoscope.
Fixing the symptom and calling it a cure
The second way is subtler, and it’s the one I actually want to talk about, because it’s the one nobody flags in the retro.
We had an incident. The experiment service went down during a database upgrade — routine, planned, a few minutes of downtime nobody thought twice about. Except the frontend depended on the experiment service. Nobody knew that. So a “few minutes of expected downtime” became a customer-facing outage, and everyone spent the afternoon rediscovering an architecture diagram that had been lying to us for a year.
Post-mortem. Action item, unanimous, sensible-sounding: add a check that verifies the frontend doesn’t call the experiment service during maintenance.
Great. So I asked the question that made the room quiet.
What about next week, when it turns out checkout is quietly reading from remote config? Or that the mailing system leans on auth in some way nobody documented? We didn’t find the hidden dependency. We found a hidden dependency, and then wrote a test for that exact one, like a man who trips over a paving stone and responds by memorizing the location of that specific stone.
And notice what made the room go quiet — because it wasn’t a hard question. It was an unfocused one. Remember agency, the thing we spent the last section watching get thrown out of the room? Here it is again, except this time it gets invited back, and it looks like wisdom. “Focus on what you can impact.” “Don’t point fingers — just fix what’s yours.” On the surface, lovely. Grown-up. Very stoic. But watch where it seats you: at your own little square of the map. Your service, your test, your runbook. The check you can add without asking anyone else to change anything. Agency, but housebroken — allowed back in, as long as it keeps its eyes on its own feet and doesn’t wander two teams over or, God forbid, upstairs. And that domesticated ownership is exactly what makes the small fix feel not just acceptable but virtuous. You’re not being narrow. You’re being focused.
Which makes “what else is silently coupled like this?” a heresy. Not a wrong question — a rude one. It unfocuses. It wanders off your square and starts pointing at the whole map.
This is the part that gets skipped: the incident is a sample, not the disease. The specific failure — experiment service, database upgrade, frontend — is one draw from a distribution you don’t understand. The engineering move isn’t to patch the draw. It’s to fix the whole distribution: to ask what class of thing you just discovered you’re blind to, and go looking for the rest of it. Not as an intellectual exercise. As a hunt for the other silent dependencies that are, right now, waiting for their own bad afternoon.
Generalizing the failure is harder than patching it. That’s exactly why it doesn’t happen. A specific test is a closed ticket. “We don’t actually know what depends on what, and we should fix that” is a quarter of work with no obvious owner. One of those fits in the sprint. Guess which one gets written down.
The ceiling nobody mentions
And here’s where the two failures turn out to be the same failure wearing different clothes.
The reason we patch the specific stone instead of surveying the whole sidewalk isn’t only laziness. It’s that surveying the whole sidewalk means touching things outside the room. It means telling team B, C, D and E to change how they work. And to justify that, you have to stand up and say: this happened because team A shipped something blind — and here’s why every other team is one silent dependency away from the same fate.
And here the room stiffens again. Blame team A? In a blameless culture? How dare you. Never mind that you’re not asking to punish anyone on team A — you’re asking to learn from what team A did, which is supposedly the entire point. Doesn’t matter. The word “blameless” has quietly been upgraded from “we won’t punish you” to “we won’t mention you,” and naming a team as the origin of a systemic lesson now registers as an attack. So the lesson dies at the property line of the team that owns it, out of politeness.
And that’s the whole trap, closing. To generalize the lesson, you have to say “team A broke it.” But the culture now says you’re not allowed to say team A broke it. So the one move that would turn a single outage into an organization-wide improvement is forbidden by the very principle that was supposed to make us better at learning from outages. Blameless, deployed as a synonym for nameless, doesn’t protect people. It protects the diagram. It eats its own tail — the rule written to keep engineers safe becomes the rule that keeps the system from ever examining itself above the level of a single team.
So the action items stay small. Not because the bigger fix is unclear. Because the bigger fix would have to point at something, and pointing is the one thing we’ve outlawed.
And notice I’ve kept this entirely at eye level — you, your peers, team A across the hall. This is the gentle version of the problem, the one where everyone in the room is roughly equal and the only thing getting hurt is a diagram. It gets a great deal worse the moment the thing you’d have to point at is sitting a floor above you. But that’s a climb for part two.
The horoscope
So this is what “blameless” quietly turns into.
We don’t name the person, because naming is uncomfortable — and now the report can’t explain how anything happened. We don’t name the team, because that would be blame — and now the lesson can’t leave the room it was born in. Two erasures, same reflex: protect someone from a word, and lose the fix that word was hiding.
We didn’t pick the true conclusions. We picked the safe ones. And a document full of safe conclusions is a document that offends no one and teaches nothing. That’s not a post-mortem. That’s a horoscope with a severity rating.
That’s meaningless.
And the worst part is that meaningless is the optimistic case. It gets a lot worse once the hierarchy gets involved — but as I said a moment ago, that’s a topic for part two.



