Everything breaks.
Reliability engineering is the science of quantifying, understanding, and managing failure rates. Reliability engineering includes planning for the aftermath of failure while working with design to minimize failure and limit its impact. The careful use of redundancy in design, routine maintenance/inspection, and recovery planning are at the heart of the discipline.
The other day I was talking with a colleague about failure rates and error rates (Admit it, you’re jealous.) He said, “Well, people make mistakes, too.” That’s where I lose it in that kind of conversation.
“No. Machine’s are not like people. They don’t have compassion, empathy, a conscience, financial worries, fights with their spouse, crazy kids, elder care concerns, or heart disease.”
“Machines, on the other hand are very predictable. All you have to do is test them adequately. You can know, with great precision, how a machine fails, how often that is likely to happen and where the failure is likely to manifest itself.”
AI is more like a toaster than it is like a human being.
We get blinded to that fact when we anthropomorphize. Pretending that a machine is human (which the AI companies want you to do) puts you in the position of believing that a machine should be judged by comparison to human beings.
Balderdash.
With machines, testing shows the timing and mode of failures. Given our massive compute capacity, exercising an agent 100,000 times to learn where it breaks is not a problem. That’s one test script and one click of a button.
Admittedly, the test plan takes imagination, detailed hard work, and time to set up. The harder challenge is sifting through the results to see what is there. Fortunately, we have LLMs to help with that. Pattern recognition is their superpower.
Failures can take many forms, ranging from slightly off, to completely wrong to dangerous, with errors occurring on a scale from inconsequential to disastrous. The one thing you know for sure about an AI implementation is that it will fail. The responsible thing is to figure out how often and how bad.
Without an understanding of the consequences of failure, it’s impossible to manage risk.
Currently, we are neglecting risk assessment in favor of ‘token maxing’. You can be certain that the ungoverned use of AI will not end well. The difficult exercise of understanding what could go wrong is an argument for a centralized AI control function. In our new world, failure spreads across functions rather than staying in its swim lane. A centralized AI operation would have to be built by consolidating responsibility across functions.
Imagine trying to install that.
It is as if we put a slide rule in the hands of every user but gave no usage instructions. Is it a fancy pea shooter, a desk ornament, a fidget spinner, a nice straight edge, a weapon? The centralized approach is a classic example of closing the gate after the horse left the corral.
And still, the broad understanding of how AI works in our organization needs some kind of home. No agent should be released into the organization without a rigorous testing and evaluation process. Every agent should emerge with onboard testing to monitor drift/failure and to correct it or turn the agent off when required.
The real key to understanding the future of work is going to be the rearrangement of accountability, authority, and responsibility in the organization. AI Governance is at the heart of the question.



