\n\n\n\n When Your AI Cannot Have a Bad Day - Agent 101 \n

When Your AI Cannot Have a Bad Day

📖 5 min read•810 words•Updated Sep 24, 2026

Think about the difference between a spellchecker and a parachute. If the spellchecker misses a typo, you shrug and fix it later. If the parachute misses, there is no later. Most of the AI you have used so far lives in spellchecker territory. It drafts an email, suggests a photo caption, guesses what you meant to search. When it flubs, the cost is mild annoyance.

The people speaking on the Real World AI Stage at TechCrunch Disrupt 2026 build the parachute kind. Shield AI chief technology officer Nathan Michael, Waabi founder and CEO Raquel Urtasun, and General Motors director of robotics strategy Mikell Taylor are sitting down to talk about what changes when failure is not an acceptable outcome. Defense systems. Self-driving trucks. Cars on public roads. Three very different products, one shared problem.

Why this conversation is different from the usual AI panel

Most AI discussion right now centers on capability. Can the model do the thing? Can it do the thing faster, cheaper, with fewer humans in the loop? That framing works fine for a chatbot that helps you rewrite a cover letter.

It falls apart the moment your AI operates a two-ton vehicle. The question stops being “can it do the thing” and becomes “how do we prove it does the thing, reliably, in situations nobody thought to test for, and how do we convince a regulator we are right?” That is a fundamentally harder question, and it is the one this panel is set up to address: safety validation, regulatory navigation, and building trust.

Notice that none of those three are engineering problems in the narrow sense. They are engineering problems tangled up with legal problems, public perception problems, and the very human problem of deciding how much certainty is enough.

The GM number worth sitting with

General Motors revealed that nearly 90% of the code created by its autonomous driving team is now AI-generated. If you are not a developer, let me translate what that means in practice.

Code is the set of instructions that tells a machine what to do. Historically, humans wrote it, line by line, and other humans reviewed it. In GM’s autonomous driving work, the overwhelming majority of those instructions are now being produced by AI, with humans in a different role: directing, reviewing, deciding what gets shipped.

That is a big shift in who is doing the typing. It is not the same as a shift in who is accountable. Somebody at GM still has to stand behind the behavior of a car in traffic, and no amount of AI-generated code changes that. This is the tension the panel is walking into. The tooling for building safety-critical systems is getting dramatically faster. The standard for proving those systems are safe has not gotten one bit looser.

Validation is the actual product

For a non-technical reader, here is the mental model I would hold onto. In consumer AI, the demo is the product. You see it work, you like it, you use it. In safety-critical AI, the demo is the easy part. The product is the evidence.

Anyone can film a self-driving truck completing a highway run. Proving that the same system behaves sensibly in fog, at a construction detour, around a driver doing something genuinely stupid, and in the thousands of scenarios nobody filmed, is the real work. That evidence is what a regulator reads. That evidence is what determines whether a company becomes a real business or an expensive footnote.

Waabi has been putting money behind exactly that thesis. The company secured $1 billion in new funding in January 2026 and announced progress on generalization, the ability for a system to handle situations outside its training, in June. Generalization is not a marketing word here. It is the difference between an AI that has memorized a route and one that can reason about a road it has never seen.

What you can take from this if you do not write code

You do not need to follow autonomous vehicle engineering to get value out of this panel. The useful takeaway applies to any AI you are asked to trust at work.

  • Ask what happens when it is wrong, not just how often it is right. The cost of a mistake tells you how much scrutiny the system deserves.
  • Ask who validated it and how. “It performed well in testing” is a claim. Testing methodology is evidence.
  • Ask who is accountable. AI-generated work still needs a human name attached to the decision to ship it.

Shield AI, Waabi, and GM operate in domains where getting these answers wrong is measured in lives rather than customer complaints. The discipline they have built under that pressure is the most transferable thing in AI right now, and it travels well down to the far more ordinary decisions the rest of us make about which tools to trust.

🕒 Published:

🎓
Written by Jake Chen

AI educator passionate about making complex agent technology accessible. Created online courses reaching 10,000+ students.

Learn more →
Browse Topics: Beginner Guides | Explainers | Guides | Opinion | Safety & Ethics
Scroll to Top