Root Cause Analysis as Ceremonial Flogging
There is a particular stage of a critical production incident when engineering stops being engineering and becomes theater.
The application is failing. Customers are getting errors. Graphs have abandoned the normal range and are now exploring modern art. Someone is restarting pods, someone else is checking database saturation, three people are trying to determine whether the queue is backed up or merely expressing itself, and a developer who has not spoken in twenty minutes suddenly says, “Oh. That’s bad.”
This is the useful part of the incident.
Then the architects arrive.
Not physically, usually. They appear in the war room as a collection of initials.
DG joined the meeting.
RM joined the meeting.
Chief Architect joined the meeting.
This is when you know the technical problem is about to receive governance.
The architects do not begin by asking what they can do. They begin by asking what happened, which is reasonable in principle, except that twelve people have spent the last forty minutes discovering what happened and are currently engaged in the somewhat urgent activity of making it stop happening.
Someone gives the thirty-second version. Traffic increased. The shared service saturated. Retries made it worse. Several callers piled on. Eventually the dependency fell over hard enough that healthy services began failing too.
There is a pause.
Then one of the architects says, “This is exactly the kind of cascading failure we’ve been warning about.”
And there it is.
The first stroke of the ceremonial flogging.
We Told You So
There is a correct time to say “I told you so.” I do not know when it is. I am reasonably confident it is not while another engineer is typing kubectl commands like a bomb disposal technician.
Nevertheless, the architecture contingent is now fully engaged.
“We’ve been talking for years about proper service isolation. What we need to understand is why this path was allowed to remain coupled in the first place, because if you look at the broader—”
Chief Architect left the meeting.
Nobody reacts.
This is normal.
The Chief Architect is extremely busy. His thoughts are therefore delivered in a distributed fashion. You may receive the beginning of one on Tuesday and the conclusion sometime after the next reorganization.
Someone says the downstream pool is at 98%.
Another architect continues. “Right, but this is exactly why standards matter. We cannot keep solving these things tactically. We have to start thinking in terms of architecture rather than symptoms.”
Someone says they are shedding traffic.
Error rate begins to fall.
“We need to get out of this reactive mode.”
At this point the system may actually be recovering, but nobody is entirely sure because the war room is now receiving a keynote.
The interesting part is that everybody in the room knows who designed the system.
The architecture was not discovered in a cave. It did not arrive on stone tablets. These same architects reviewed it. Some of them drew it. One of them almost certainly has the original diagram, complete with pleasantly aligned boxes, carefully spaced arrows, and a database cylinder positioned so confidently that it seems almost rude to question it.
But somehow, during an incident, architecture undergoes a legal transformation. It is no longer something architects helped produce.
It becomes something that happened to the architects.
They warned us. Apparently. Constantly. Heroically. Like villagers watching a volcano while the rest of us insisted on putting a shopping mall in the crater.
The Experts Have Arrived
I am in favor of experts. Experts are useful. But software has developed an unusual category of expert whose primary qualification is that they have achieved a safe distance from software.
They know Java.
Or knew Java.
Java was definitely involved.
There was probably an application server. Perhaps EJBs. Maybe XML configuration so elaborate that archaeologists will someday classify the era by schema version. Then, at some point, these people ascended.
They no longer write code.
They write principles.
The principles are important because code is specific, while principles can survive indefinitely without encountering a compiler.
This creates an interesting dynamic in a war room. The engineer who changed the service last Tuesday is explaining that retry concurrency is multiplying load on a downstream dependency. The architect who last shipped production code when Java was a thing people were still excited about is explaining why the actual problem is insufficient adherence to enterprise architectural discipline.
The developer has metrics.
The architect has rectangles.
The developer can show the failure.
The architect can show where the failure should not have happened.
This is a surprisingly powerful position. Reality has made the tactical mistake of disagreeing with the diagram.
Then, ten minutes later:
Chief Architect joined the meeting.
“Sorry, I had to jump to another critical call. Where were we?”
Someone begins to explain.
“No, no, I’ve got the gist. The core issue here is that we’ve allowed too much local optimization and not enough—”
Chief Architect left the meeting.
You have now received another architectural fragment.
Keep these. Eventually they may form a sentence.
BD: Boxes Drawn Per Minute
Architects have a measurable productivity metric, although the organization has not yet had the courage to formalize it.
BD.
Boxes Drawn per minute.
A junior architect may sustain three or four BD during normal meetings. A principal architect who senses organizational attention can exceed twelve. During a major incident, with executive visibility, the numbers become frightening.
The shared screen changes. Dashboards disappear. The system architecture appears.
A box is drawn around the failing component.
“This is the fundamental issue.”
Another box appears.
“We need an isolation layer here.”
Then another.
“Potentially some form of orchestration.”
An arrow.
“A policy boundary.”
Another arrow.
“Possibly asynchronous decoupling.”
The system is still partially down. Nobody has said who is going to build any of this. That is implementation detail.
We have moved beyond implementation detail.
At eighteen BD, the diagram begins generating its own weather.
The developer who knows the actual codebase watches all this with the expression of a man seeing plans for a second floor added to a submarine.
The Incident Becomes Educational
At some point, the outage ends. This creates a dangerous vacuum.
The people who were actively fixing things stop typing. Someone says traffic looks stable. Someone else says they are going to leave the mitigation in place overnight. A third person has already mentally moved on to food.
This should be the end of the call.
It is instead the beginning of the lecture.
Now that customers are no longer actively being injured, we can focus on the people who remain.
The architects.
For the next hour, they explain how incidents like this are avoided. Not how this incident could have been avoided. That would require uncomfortable specificity. Instead, we receive the general theory of Avoiding Problems.
Architectural rigor. Proper standards. Up-front design. Clear ownership. Service boundaries. Failure isolation. Governance. Someone says “holistic.”
This is usually how you know dinner is no longer happening.
The war room becomes continuing education. The people who have just spent two hours reading logs, checking traces, correlating deployments, reducing concurrency, patching retry behavior, and watching actual production behavior are instructed in the importance of thinking about production behavior.
Then the Chief Architect reappears.
Chief Architect joined the meeting.
“I just want to make one point before I have to drop. We need to stop treating architecture as something that happens after the fact. Architecture is how we prevent these kinds of issues in the first place, and if teams are empowered to make local decisions without—”
He stops.
There is perhaps a half-second of silence.
Chief Architect left the meeting.
Someone asks whether he lost connection.
No.
He had another meeting.
Of course he did.
The Minor Historical Problem
There is, occasionally, an awkward fact.
The service boundary being criticized was approved by architecture. The shared dependency was part of the platform standard. The retry strategy came from a common library. The database topology followed the reference design. The “temporary” coupling exists because an architecture review rejected a smaller solution three years earlier in favor of something more strategic.
Everybody knows this.
Nobody says it.
This is professionalism.
It is possible that someone has the old design document open in another tab. It is possible that the document contains an architect’s name. It is possible that the architect currently explaining the importance of service isolation is that architect.
These things happen.
Organizations have poor memory but excellent document retention.
The temptation to paste the link into the meeting chat can become physically painful.
Do not do this.
You are tired. You have elevated privileges. Nothing good can come from this combination.
The Root Cause
Eventually someone asks for the root cause.
Technically, the root cause might be something like this: a downstream dependency slowed, callers retried aggressively, concurrency increased, the shared resource saturated, latency increased further, more callers timed out, and more retries followed.
Congratulations. You have built a small distributed denial-of-service attack using only approved internal components.
But this explanation is too mechanical. It does not produce enough institutional learning.
So we continue digging.
Why was the retry policy too aggressive? Because it was the shared default.
Why was it the shared default? Because the platform team standardized it.
Why did we use the shared library? Because architecture required the standard client.
Why did architecture require the standard client?
Silence.
This line of inquiry is becoming counterproductive.
Someone suggests we focus on forward-looking actions.
The ceremonial flogging resumes.
Blamelessness
Modern incident management is often described as blameless. This is a good idea. People make mistakes. Systems create conditions in which mistakes become incidents. Punishing individuals discourages reporting and teaches everyone to conceal uncertainty.
So we do not blame people.
We blame categories.
“Insufficient architecture.”
“Lack of governance.”
“Failure to follow standards.”
“Inadequate technical oversight.”
These are excellent blame targets because they sound systemic while still pointing vaguely toward someone else.
The best category is “engineering maturity.”
Nobody knows exactly who owns engineering maturity, so everyone can agree it must improve.
An action item is created.
It will be due next quarter.
The Great Prevention
By the end of the meeting, we have developed a plan to prevent recurrence.
There will be a new architecture review requirement for high-criticality services. The architects will review designs earlier. Teams will document failure modes. Retries will require explicit approval. Critical dependencies will need isolation plans. A working group may be formed.
Possibly a council.
This is all quite sensible.
That is the dangerous part.
Every process added after an incident is sensible in isolation. Enough sensible processes, stacked carefully, will eventually prevent anyone from deploying anything.
This is one of the ways organizations achieve reliability.
Nothing fails because nothing moves.
Somewhere, a developer will eventually propose a small service. The service will call one API and write to one database. Architecture will ask for a failure-mode analysis. The developer will provide one. Architecture will ask whether the design has considered cascading failure. The developer will add a queue. Someone will ask whether the queue is itself a single point of failure.
Another box will appear.
BD will rise.
Months will pass.
The service will finally launch.
Two years later, it will fail because of a configuration value nobody knew existed in the standard client library.
There will be a critical incident.
A war room will open.
The developers will arrive first.
They always do.
They will look at logs, traces, recent deploys, thread pools, queues, saturation, retries, and all the ugly little facts that actual systems produce when diagrams are no longer available to protect them.
Then, after perhaps twenty minutes:
Chief Architect joined the meeting.
A familiar voice will say, “This is exactly the sort of thing architecture has been warning about. If you look at the bigger picture, what we really need to understand is—”
Chief Architect left the meeting.
And somewhere, very quietly, a developer who absolutely knows better than to sign his real name will mute himself.
— Jim Plemon
Senior Software Developer
Available for Architecture Consultation Pending Legal Review