Your Data Isn't Ready for AI
80% of AI projects fail, and the #1 reason isn't the model - it's the data. Here's how to find out if yours is ready before you spend money.
TL;DR
Five data problems kill AI: siloed, inconsistent, incomplete, unstructured, and stale. Seven questions to assess readiness. Clean data first, build AI second. AI on bad data = confidently wrong answers.
Your data is messier than you think
Every AI project starts with the same assumption: “our data is fine.”
It’s never fine. I’ve never audited a company whose data was as clean as they thought. Not once. Not in four years of building AI systems.
This isn’t a criticism. It’s reality. Data degrades over time. Systems change. People leave. Documentation drifts. Fields that were supposed to be required aren’t. Formats that were supposed to be standard aren’t. Data that was supposed to be in one system is in three.
The companies that succeed with AI don’t have perfect data. They have honest data - they know what’s broken, they know what’s missing, and they fix it before building on it.
The 5 data problems that kill AI
1. Siloed data
Your customer data is in the CRM. Your transaction data is in the ERP. Your support data is in the ticketing system. Your product data is in a separate database. None of these systems talk to each other.
AI needs connected data. If you’re building a customer churn predictor, the AI needs to see CRM data (who they are), transaction data (what they bought), and support data (what they complained about). If these are in separate silos with no common identifier, the AI can’t connect the dots.
Symptom: “We’d need to export from three systems and merge them manually.” Fix: Build a data pipeline that consolidates sources into a single view. Before the AI project.
2. Inconsistent data
The same field is formatted differently across records. Phone numbers: “+34 600 123 456”, “34600123456”, “+34-600-123-456”, “0034 600 123 456”. Company names: “ACME Corp”, “Acme Corporation”, “ACME”, “acme corp”. Dates: “2024-01-15”, “15/01/2024”, “Jan 15, 2024”.
Inconsistency confuses AI. It sees four different phone numbers when they’re the same number. It sees four different companies when they’re the same company. Patterns that should be obvious are hidden by formatting noise.
Symptom: “We have a lot of duplicates but we’re not sure how many.” Fix: Data normalization - standardize formats, deduplicate records, validate against reference data.
3. Incomplete data
Fields are empty. Records are missing. The CRM has a 60% completion rate on the “industry” field. The ERP is missing customer email addresses for 40% of records. The support system doesn’t link tickets to customer accounts.
Incomplete data means the AI is making decisions with partial information. It’s like asking someone to predict the weather with temperature data but no humidity, no pressure, and no wind speed.
Symptom: “We track that… sometimes… when someone remembers.” Fix: Identify which fields are critical for your AI use case. Enrich them. Make them required going forward. Accept that historical data may be permanently incomplete.
4. Unstructured data
Your most valuable data is in documents, emails, PDFs, meeting notes, and Slack messages. It’s not in any database. It’s not searchable. It’s not structured.
AI can work with unstructured data - that’s one of its strengths. But only if you can access it, organize it, and feed it to the AI in a way that’s useful. If your institutional knowledge lives in 10,000 PDFs scattered across SharePoint, that’s a data infrastructure problem before it’s an AI problem.
Symptom: “Our processes are documented… somewhere… by someone… at some point.” Fix: Document inventory, consolidation, and indexing before building AI on top of it.
5. Stale data
Your data was accurate when it was entered. But that was 2 years ago. Customers have changed jobs. Companies have merged. Products have been discontinued. Prices have changed.
AI trained on stale data makes recommendations based on a world that no longer exists. It recommends products you don’t sell. It contacts customers who have left. It classifies transactions using outdated categories.
Symptom: “We updated the CRM… in 2023.” Fix: Establish data refresh cycles. Identify which data changes frequently (pricing, inventory, customer status) and which is relatively stable (company name, industry). Prioritize freshness for the data that feeds your AI.
How to assess data readiness
Seven questions. Answer them honestly:
-
Where does your data live? Can you list every system, database, and document store in one sentence? If it takes a paragraph, you have silos.
-
Who owns it? Is there a person responsible for data quality in each system? Or is it “everybody’s responsibility” - which means nobody’s?
-
How clean is it? What percentage of records are complete? What percentage are duplicates? What percentage have formatting issues? If you don’t know, that’s the answer: you don’t know.
-
How complete is it? What percentage of critical fields are filled? Which fields are routinely left empty? Why?
-
How consistent is it? Are the same fields formatted the same way across records? Across systems? Can you merge records from different systems without manual cleanup?
-
How fresh is it? When was the last update? How often does it change? Is there a refresh process, or does data only change when someone manually edits it?
-
How accessible is it? Can your team get to the data via API? Via database query? Or does it require an export, an email, and a waiting period?
If you can’t answer these questions confidently, you don’t have a data readiness problem. You have a data awareness problem. The audit fixes both.
The data audit checklist
Before building AI, run this checklist:
- Source inventory: List every data source, its format, its owner, and its update frequency
- Ownership: Assign a person to each data source - not a team, a person
- Duplication analysis: How many duplicate records exist? What’s the dedup strategy?
- Completeness scoring: For each critical field, what percentage is populated?
- Consistency check: Sample 100 records. Are formats consistent? Are values valid?
- Access audit: Can you access each source via API or direct query? What authentication is needed?
- Pipeline reliability: If you build a data pipeline, how will you know when it breaks?
What “data ready” looks like
- Single source of truth for each data type (or a documented process for merging sources)
- Documented schema - every field has a name, type, description, and validation rule
- >95% completeness on critical fields
- Consistent formats - same field, same format, every record
- Regular updates - data refreshes on a schedule, not when someone remembers
- Accessible via API - not just via CSV export and email
If you’re not there yet, that’s fine. Most companies aren’t. But fix it before building AI. AI on bad data doesn’t produce “roughly right” results - it produces confidently wrong results that look credible. That’s worse than no AI at all.
Why most consultants skip this
Data work is unglamorous. It doesn’t produce demos. It doesn’t make for impressive presentations. It doesn’t have a “wow” moment.
But it’s the difference between AI that ships and AI that fails. Every time.
A consultant who skips data assessment and jumps to model building is either inexperienced or optimizing for their own timeline, not yours. The model is the fun part. The data is the important part.
FAQ
How long does data cleanup take? 2-8 weeks depending on the number of systems, data volume, and state of disarray. It’s not glamorous, but it’s the foundation everything else stands on.
Can we do AI and data work in parallel? No. AI on bad data = bad AI. You can build the data pipeline while assessing the model approach, but the data must be ready before training or retrieval begins. Otherwise you’re building on sand.
What if we don’t have enough data? Sometimes the answer is “collect more data first.” This is frustrating to hear but better than building AI that doesn’t work. Alternatively, use an API-based approach that doesn’t require your own training data - but you still need clean data for retrieval and context.
Ready to apply this to your situation?
Book an AI Readiness Call30-min call. No pitch. You leave with one concrete next step - even if it’s not us.
Jacek Trefon
AI engineering leader. 28 years building technology, 4+ years building production AI systems. I help companies assess, architect, build, and deploy AI that actually ships. Based in Spain, working globally.
Keep Reading
All articles →AI Projects I Turn Down
Every consultant says they're honest. Few prove it. Here's my proof: a list of AI projects I've turned down, why I said no, and what I recommended instead.
Build vs Buy vs Wrap: The AI Framework
The three options - wrap an API, buy a platform, or build custom - each have clear trade-offs. Here's the decision tree I walk clients through, with real costs.
The CEO's Guide to AI
Most CEOs don't understand AI. They pretend they do in board meetings while secretly Googling 'what is a large language model.' Here's what you actually need to know.