The Tribunal
The last architecture tribunal I attended was about a critical application I wrote in AWS Lambda. The timing was excellent: the team had about a year to ask questions before I shipped it. Nobody did. Then it ran in production for six months. Then we had the tribunal.
This is roughly how architectural governance works in mature organizations. First, you are trusted to design and ship the system. Then, once it is demonstrably working, everyone becomes extremely concerned about the design.
Round one was me in a room defending the decision. Why Lambda? What about scaling? Cold starts? Cost? Deployment? Operational support? All reasonable questions, and all questions I had considered before putting a critical workload into production, because I am a senior engineer and still have enough brain activity to operate heavy machinery.
The biggest reason was scaling. At the time, Kubernetes autoscaling was a deal breaker for this workload. Karpenter could add capacity, but adding nodes could take minutes. Minutes are fine if you are baking bread. They are less ideal when traffic has already arrived and customers are currently discovering your capacity model.
Lambda avoided the problem entirely. We kept the ZIP lean, cold starts were well under 500ms, and it scaled fast enough for the workload without needing anyone to stand nearby holding a wrench. This was not theoretical. By tribunal time, we had six months of production evidence.
Round two happened without me.
Architecture and SRE got together and produced a spreadsheet showing that Lambda was operationally much more expensive than Kubernetes. This was impressive because AWS had already produced a competing spreadsheet. It was called the bill.
We could see what Lambda cost. The tribunal spreadsheet had the advantage of being able to make Kubernetes costs disappear. Platform team? Existing. Cluster? Existing. Nodes? Existing. Operational labor? Existing. Months spent building the Kubernetes deployment path? Apparently part of nature.
Meanwhile, the Lambda pipeline had needed about two weeks of Jenkins work. This was presented as concerning. The Kubernetes deployment had taken months.
Two weeks, however, was evidence that Lambda was weird.
You have to admire the accounting.
Later, SRE improved Kubernetes autoscaling. And to be fair, they did improve it. The solution was clever.
Instead of waiting for new EC2 capacity when workloads needed to scale, they proactively added nodes to the cluster and filled the spare capacity with dummy pods. Then, when a real application needed more room, Kubernetes could evict the dummy pods and replace them with useful ones very quickly.
Sub-second autoscaling.
Innovation.
This is still how it works today, and look, technically, it works. But I struggle a little with the word innovation here. We pre-provisioned machines, filled them with fake work so the scheduler would keep them warm, then murdered the fake work when actual work appeared.
This is not quite the moon landing.
It is more like reserving every table in a restaurant with mannequins so real customers can be seated immediately.
Very clever mannequins. Possibly enterprise-grade mannequins.
The important thing is that nobody has to say, “We keep spare servers sitting around.”
We have dummy pods.
Completely different.
I am not anti-Kubernetes. I am anti pretending the answer was obvious after the fact. At the time we made the decision, the Kubernetes scaling behavior did not meet the requirement. Lambda did. Later, Kubernetes was tuned until it could.
Great. That is engineering.
What is not engineering is retroactively pretending the original decision was foolish because the preferred platform eventually invented a workaround. Especially when that workaround is, essentially, “keep extra computers running.”
We used Lambda because we wanted fast scaling without keeping extra computers running. Kubernetes eventually achieved the same responsiveness by keeping extra computers running.
Then the spreadsheet explained that Lambda was the expensive option.
Perfect.
I think the real problem was not that Lambda was unreliable. It was too reliable in the wrong way. It did not need the royal Kubernetes cluster. It did not need months of deployment work. It did not need a clever pool of sacrificial dummy pods waiting to die for the greater good.
It just ran. For six months. Quietly.
And nothing irritates a standardized platform quite like a workload that appears not to need it.
That is how you end up in a tribunal. Not because the system failed, but because it worked without permission.