What If AI Makes a Mistake?
AI will make a mistake. Not might - will. The question is: when it happens, will you know, will you be able to explain it, and will you be protected?
TL;DR
Three types of AI mistakes: confident hallucinations, biased decisions, and silent drift. Each has a prevention strategy. The real risk isn't AI making mistakes - it's not knowing when it does.
The fear that stops deals
Every executive I talk to has the same nightmare: AI makes a wrong decision, a customer is harmed, the company is sued, and nobody can explain why the AI did what it did.
This fear is rational. AI systems are opaque. When they’re wrong, they’re wrong with confidence. And “the AI did it” is not a legal defense.
But this fear, left unaddressed, becomes paralysis. You don’t deploy. You don’t ship. You watch competitors move while you’re stuck in risk assessment meetings that never end.
The solution isn’t to avoid AI. It’s to build AI with accountability built in.
The three types of AI mistakes
Not all AI errors are the same. Understanding the types determines your prevention strategy:
Type 1: The confident hallucination. The AI states something false with complete confidence. Common in large language models. Example: AI tells a customer their insurance policy covers something it doesn’t. Customer acts on it. Company is liable for the AI’s promise.
Type 2: The biased decision. The AI makes a decision that systematically disadvantages a group. Common in hiring, lending, insurance, and pricing. Example: AI rejects resumes from a specific demographic because the training data favored another. Company faces discrimination claims.
Type 3: The silent drift. The AI was accurate at launch. Over months, data changes, patterns shift, accuracy degrades. Nobody notices until damage is done. Example: fraud detection model trained on 2024 fraud patterns misses a new 2026 fraud type. Losses accumulate for months before anyone detects the problem.
Each type has a different prevention strategy. Knowing which ones apply to your use case is the first step.
How to prevent Type 1: Hallucination
Ground the AI in your data. RAG (Retrieval-Augmented Generation) means the AI retrieves from your documents before answering. It’s not making things up - it’s reading. If the answer isn’t in your documents, it says “I don’t have that information” instead of inventing one.
Confidence thresholds. If the AI’s confidence is below a threshold, it says “I’m not sure” instead of guessing. This requires building confidence scoring into the system - not an afterthought, but an architecture decision.
Human-in-the-loop. For high-stakes answers, the AI drafts, a human reviews. The AI doesn’t send the email - it prepares it. The AI doesn’t approve the claim - it recommends approval. Humans stay in control of consequential actions.
Output validation. Check the AI’s response against known facts before it reaches the user. If it contradicts your knowledge base, flag it. If it makes a claim not supported by retrieved documents, block it.
How to prevent Type 2: Bias
Audit training data for representation. If your training data is 90% from one demographic, your model will be biased. This is a data problem, not a model problem. Fix the data before training.
Test for bias before deployment. Run the model on test cases designed to surface discrimination. Submit synthetic resumes with identical qualifications but different names. Apply for loans with identical financials but different demographics. If the model treats them differently, fix the data, not the model.
Monitor for bias in production. Track outcomes by demographic. If disparities emerge, investigate. Not monthly - continuously. Automated dashboards that flag statistical anomalies.
Document your bias testing. Under the EU AI Act, this is required for high-risk systems. Under common sense, it’s required for not getting sued. If you can show you tested for bias and addressed findings, you have a defense. If you can’t, you don’t.
How to prevent Type 3: Drift
Monitor input distributions. If the data coming in looks different from the data the model was trained on, accuracy is degrading. Statistical tests (population stability index, KL divergence) can detect this automatically.
Monitor output distributions. If the model’s predictions shift significantly - more approvals, more rejections, different categories - something has changed. Either the world changed or the model broke. Both need attention.
Monitor user feedback. If users start reporting errors, the model is drifting. Build a feedback loop into the product. Make it easy to report wrong answers. Track error reports as a leading indicator.
Schedule regular retraining. Not “when we notice a problem” - on a cadence. Monthly, quarterly, or based on drift metrics. Treat retraining as maintenance, not as an emergency response.
Set drift alerts. Automatic notifications when metrics cross thresholds. Don’t wait for a human to notice a dashboard. Have the system tell you when something is wrong.
The accountability framework
Accountability isn’t a feature you add at the end. It’s a framework you build from the start:
Before deployment:
- Risk assessment - what could go wrong, how bad is it, who’s affected
- EU AI Act classification - what risk tier, what requirements
- Decision log - why you chose this approach, what risks you accepted, what mitigations you built
- Bias testing - documented, with results
- Security review - prompt injection, data leakage, adversarial inputs
During operation:
- Audit trail - every AI decision logged: input, output, confidence, timestamp
- Human override - users can always override the AI. Always. No exceptions.
- Escalation path - low confidence or high stakes → human review. The system knows when to ask for help.
After an incident:
- Incident response - what happened, why, what’s the impact, how do we prevent recurrence
- Documentation - you can explain to a regulator, lawyer, or customer exactly what the AI did and why
- The phrase “we don’t know why the AI did that” should never come out of your mouth
The “explainability” requirement
Under the EU AI Act, high-risk AI systems must provide “meaningful information about the logic involved in their decisions.”
This doesn’t mean you need to explain the math. It means you need to explain: what data the AI considered, what factors were most important, and why the output was what it was.
For most systems, this is achievable with the right architecture: feature importance scores, decision trees alongside neural networks, or LLM explanations of reasoning steps.
If your vendor says “AI is a black box, you can’t explain it,” they’re using the wrong architecture or don’t know how to build explainable systems. Explainability is a design choice, not a research problem.
The real risk isn’t AI making mistakes - it’s not knowing when it does
AI will make mistakes. Humans make mistakes. The difference: humans know when they’re unsure. AI doesn’t - unless you build it to.
A well-designed AI system:
- Knows when it’s uncertain and escalates
- Logs every decision for accountability
- Has fallback paths for when it fails
- Has humans in the loop for high-stakes decisions
- Can be explained to a regulator
A poorly designed AI system:
- Makes confident wrong decisions
- Has no audit trail
- Has no override mechanism
- Can’t be explained to anyone
- Discovers its errors when a customer complains
The risk isn’t AI. The risk is bad AI architecture. And that’s a choice, not an inevitability.
FAQ
Can we be sued for AI mistakes? Yes. Under EU AI Act, GDPR Article 22, and general liability law. The question is whether you can demonstrate reasonable precautions. Documentation is your defense. “We tested for bias, we monitored for drift, we had human override, we logged every decision” is a defense. “We didn’t know it would do that” is not.
Should we have insurance for AI? Talk to your insurers. AI liability insurance is emerging. But insurance doesn’t replace good architecture - it supplements it. You insure against residual risk, not against negligence.
What if our vendor says AI mistakes are “unavoidable”? They’re not preventable, but they’re manageable. If your vendor can’t explain how they manage mistakes - fallbacks, confidence thresholds, human override, monitoring - don’t hire them. They’re building a system that will fail in the worst possible way: silently and confidently.
Ready to apply this to your situation?
Book an AI Readiness Call30-min call. No pitch. You leave with one concrete next step - even if it’s not us.
Jacek Trefon
AI engineering leader. 28 years building technology, 4+ years building production AI systems. I help companies assess, architect, build, and deploy AI that actually ships. Based in Spain, working globally.
Keep Reading
All articles →AI Projects I Turn Down
Every consultant says they're honest. Few prove it. Here's my proof: a list of AI projects I've turned down, why I said no, and what I recommended instead.
The CEO's Guide to AI
Most CEOs don't understand AI. They pretend they do in board meetings while secretly Googling 'what is a large language model.' Here's what you actually need to know.
EU AI Act for Mid-Market
The first comprehensive AI regulation. Fines up to €35M or 7% of global revenue. Most content targets enterprise. This is the plain-English guide for 50-500 employee companies.