Estimated reading time: 8 minutes
Key takeaways:
- Dashboards can look fine while nobody can explain why a change failed. AI has split shipping a change from understanding it.
- Watch for recurring conflict: reopened MRs, rollbacks turning into debates, and decisions escalated that teams should settle themselves.
- Skip the comprehension score. If losing one engineer means rebuilding understanding from old prompts, that’s a risk worth fixing now, not after an SLO breach.
Your services look healthy by every measure you traditionally trust. Software delivery is moving faster than ever before. Rework is within healthy margins and AI-assisted change is routine. You may even have agentic code review and QA processes. From the senior leadership view, this looks like healthy software delivery in the age of AI.
Then a consequential change fails. Rollback is quick and seamless, with minimal customer impact. From your perspective, the event was absorbed well.
In the after-action, you ask the engineering teams why it failed, and they struggle to give you any substantial answers. They can tell you what changed, and when. They can show the approval chain and how it passed QA. No one could explain why the change caused the failure.
The dashboards did not lie, you just asked it questions it was never built to answer. You now have a different question that needs to be answered: does your organization still understand the systems they are changing?
Your inbox, upgraded.
Receive weekly engineering insights to level up your leadership approach.
The error budget accounting we never had to name
In our scenario, there was no impact to our service level objective (SLO) which in this case is 99.9% uptime. This gives us an allowance of 0.1% downtime as our error budget, roughly 43 minutes per month. This provides clear insight into a service’s reliability.
For leadership, this directly shows how much unreliability a service can absorb before we reconsider the pace or risk of further change. This gives us a confident view into service failure and performance.
Error budgets and SLOs were never intended to tell you how well your team understands a service. That knowledge often developed while engineers worked through a change. Writing the code, questioning a decision in review, and investigating a failed test all required time inside the problem.
With AI, a team can produce a working change without doing as much of that work or acquiring the understanding that comes with it.
AI can reduce the cost of producing a change without reducing the cost of understanding it. Code, tests, documentation, reviews, and even parts of QA can now be produced with far less human contact than before. Output scales, but understanding does not necessarily scale with it. The organization can create more change while carrying less comprehension per change.
This is where the unnamed accounting lives. The error budget still tells you how much failure the service has consumed. Now you can no longer infer that the people responsible for that service can explain why a change behaves the way it does, or if they can challenge its assumptions or safely modify it when the expected path breaks. The budget is still pricing failure. What changed is the collateral behind the change.
The work got cheaper
Before frontier models, shipping more software meant spending more time on standard engineering practices. You built a mental model, spent time inside the problem, and resolved it. Oftentimes you would hit walls you could not easily reason your way through. If you were lucky, you had a good mentor on your team who would coach you through it. If not, you were stuck pouring over Stack Exchange, learning discernment over what was a good solution to try and what was a bad one.
Developers now use models to write changes and tests to go with them. They often do not write any code they ship today. Code reviews are increasingly agentic and come back summarized and annotated, with suggested fixes. Documentation is generated before anyone has had to explain the decision in their own words, often in more detail than anyone needs and likely just for someone else’s AI to read. For senior leaders, this looks exactly like what we paid for. On the surface, more productivity.
What gets overlooked is the loss that comes with the disappearance of the effort in producing artifacts. I have sat in enough reviews where everyone could point to the typical evidence. The ticket, who approved it, the test result, and the rollback.
However, the conversation changed as soon as someone asked why the system behaved that way. The evidence of the process was there. The explanation was not, and this reveals a gap in understanding.
That distinction gets easier to miss as the delivery numbers improve. Risk has increased. Less people on your team have a mental model of the system. They can ship more often and recover faster, but fewer people can explain a consequential change without going back to the prompt or the tool that produced it. Nothing on the dashboard has to turn red for that to happen.
If the team responsible for a critical service has to reconstruct a material change from the prompt history before they can explain it, I want to know before we approve the next one. I would rather postpone the next change by an hour or a day a week, than discover during the next failure that we shipped past our own understanding. That is a leadership decision.
More like this
Finding legibility
Talk to the team that owns your primary critical service. Have that conversation without the go-to engineer whose name is on every review, or who is responsible for the last several material changes.
The remaining members should be able to explain what changes have been introduced, what alternatives were considered, and what risks are inherent in the choices made. If they cannot explain those decisions without reaching for AI, or waiting on the key engineer, this risk is already part of the operating model.
This lack of understanding is visible as recurring conflict before it becomes an outage. You have to be open to reading it as telemetry.
An MR or some seemingly trivial decision keeps getting reopened by the same group of engineers who keep going in circles around work they supposedly handed off. A rollback that should be routine turns into a debate, and questions the team should be able to settle amongst themselves get escalated to leadership.
On the surface, none of these events are catastrophic on their own. Taken together, the pattern suggests a team carrying responsibility without understanding the system they are accountable for.
Teams owning critical services need to make consequential decisions with a developed mental model of the system. If they have to rely on a third party, a prompt, or a model to own that, the org chart still says who owns and is accountable for the service. The real knowledge, however, is somewhere else.
Comprehension to change risk
Your first impulse might be to create a comprehension score, which is a very normal reaction. However, it is a bad idea. Plenty of dashboards already treat a number as judgment.
I distill the leadership decision down to this: if a material change depends on concentrated knowledge, single points of failure, or solely on AI models, that change is riskier than a similar change the team can explain and operate without them.
Labeling AI-produced changes with higher categories of risk does not make them bad. Work produced by or with AI is not lesser by any means. The issue is that the assurances we typically use are no longer adequate. They still matter, but are incomplete. I add how much of the knowledge required to operate this change actually exists inside the team that owns it.
If losing one engineer causes the team to flounder, and they have to rebuild comprehension of the service from old prompts before they can touch it safely, that is a single point of failure. No one needs to memorize the codebase. The owning team needs enough comprehension that a single absence or attrition does not turn the next change into archeology.
We have a team rushing to make changes to a critical service and system they do not clearly understand. If you have gone through the exercise above, it should be clear to them what’s missing in relation to comprehension. This team needs the authority to push a release back, and so should the teams supporting its delivery, while the service is healthy. If we insist on an SLO breach, or burning error budget first, we are telling them to come back after something has failed.
As senior leaders, we set the tone. Are we delivering at all cost? Do we listen to the warning signs coming from the people carrying the system? If understanding has thinned enough that the team cannot make the next decision cleanly, slowing down is not resistance to AI adoption. It is the cost of keeping control of the system we are still accountable for.

Berlin • November 9 & 10, 2026
Close the gap between what leadership expects and what’s actually possible at LeadDev Berlin.
AI changed the unit economics of change
In our scenario, the service came back. Customers were not affected, and the error budget survived. By the measures we put in place, the organization handled the event well. Yet we still walked out of the review unable to explain why the change failed. I would not let the green dashboard erase this.
AI has made change cheaper and faster, but that is not so for understanding. We no longer get to assume that knowledge arrives with the artifact simply because it once did. If we want speed, we also have to decide what level of understanding is required before the next change moves. Both are achievable with AI.
However, waiting for the SLO to burn means production becomes the place where we discover what the organization stopped understanding. That is too late, and that decision belongs to us.