Here is an unpopular position: the biggest problem with AI systems in 2026 isn’t that they fail. It’s that we’ve all quietly agreed to stop asking why.
The mainstream story goes something like this — AI tools are getting smarter, deployments are getting smoother, and the occasional glitch is just growing pains. I don’t buy it. What I see, and what people inside software engineering are saying out loud this year, is something stranger. Failures are happening more often, they’re harder to explain, and the standard response has become a shrug. Not a postmortem. Not an investigation. A shrug.
What an inexplicable failure actually looks like
If you’re not technical, you might picture a failure as a crash. Red text, an error message, something obviously broken. Those are the easy ones. Somebody can trace them.
The failures I’m talking about are quieter. A report runs but the numbers are wrong. An automated process skips a step and nobody notices until a customer calls. An AI assistant that worked fine on Tuesday gives different answers on Wednesday, with no change to anything anyone can point to. Someone restarts it. It works again. Everyone moves on.
That last part is the interesting bit. In an older era of software, “it started working again and we don’t know why” was a five-alarm situation. Today it’s Tuesday.
Why this is happening now
Two pressures are stacking up, and they reinforce each other.
The first is complexity. Modern business software is layers on layers. An ERP system that handles your accounting sits under integrations, which sit under dashboards, which now sit under AI features bolted on top. Each layer was built by different people at different times with different assumptions. When something goes wrong three layers down, the person seeing the symptom has no visibility into the cause.
The second is people. There’s a labor shortage in exactly the roles that would normally chase these problems down — the engineers with deep knowledge of how a specific system actually behaves. Institutional memory is thin. The person who would have known why the batch job fails on the last day of the month left eighteen months ago and nobody wrote it down.
Put those together and you get a predictable outcome. Investigating a weird failure is expensive, slow, and requires expertise nobody has to spare. Restarting the service is free. Guess which one wins.
The moderate failure is the real story
We tend to talk about software disasters — the big, public, embarrassing ones. Those get written up and studied. But the more common pattern in 2026 is the moderate failure. The project that ships eight months late and does two-thirds of what was promised. The digital transformation that technically went live but that nobody trusts enough to stop running the old spreadsheets alongside it.
These aren’t catastrophes. They’re worse in aggregate, because they’re survivable. They cost real time, real money, and real customer goodwill, and because no single one of them is fatal, nobody declares an emergency. The organization absorbs the damage and keeps going. That absorption is the normalization.
The broken door problem
There’s a framing I came across recently that stuck with me. A broken door is funny. Someone pushes when they should pull, everyone laughs, the day continues. A broken AI service is not funny, because you can’t see the mechanism. With the door, you understand immediately what went wrong and you adjust. With the AI service, you learn something different — that the thing cannot be trusted, and you don’t know when it can be.
Trust is the resource being spent here. Not uptime. Not budget. Trust.
What non-technical people can reasonably do about it
You’re probably not going to debug an ERP integration. But you have more standing here than you think.
- Ask why, out loud, and write down the answer. Not to assign blame. To create a record. “We don’t know” is a legitimate answer, and documenting that you don’t know is how you eventually find out.
- Track the small stuff. If the same weird thing happens monthly, that pattern is the most valuable diagnostic information in the building, and you’re the one holding it.
- Treat “we restarted it” as incomplete. It’s a fine short-term fix and a terrible explanation. Ask what would happen if it broke again during your busiest week.
- Be skeptical of adding AI on top of something already shaky. If the underlying system is unreliable, a new layer doesn’t fix that. It hides it.
Where I land on this
I’m not arguing against AI tools. I write about them because I think they’re genuinely useful. But usefulness and trustworthiness are separate things, and right now the industry is shipping a lot of the first while quietly hoping nobody audits the second.
The normalization of unexplained failure isn’t a technology problem. It’s a cultural one, and cultural problems can be reversed by people deciding to behave differently. Asking why is free. It’s also, increasingly, rare — which is exactly what makes it worth doing.
🕒 Published: