I want to share a conversation that happens, in some version, in nearly every AI engagement we take on. It happens about three months in. And it is almost always a surprise to the client, even though it shouldn’t be.
The project is scoped. The use case is defined. The AI architecture is designed. And then we look at the data — really look at it — and discover that what exists doesn’t match what was assumed to exist. It’s fragmented across systems that don’t communicate. It’s inconsistently formatted. Large portions of it are missing. The definitions used in one system contradict those used in another. Historical data was captured manually and has errors no one ever cleaned. The AI we’ve been asked to build needs clean, structured, accessible data — and the data that exists is none of those things.
This is not a rare edge case. It is the default state of enterprise data in almost every organisation that hasn’t explicitly invested in data infrastructure. And it is the single most common reason AI projects take twice as long and cost twice as much as scoped.
Why This Is Consistently Underestimated
The problem is invisible until you look. When leadership teams assess AI readiness — in the strategic conversations, the Board presentations, the vendor evaluations — data is usually mentioned but rarely examined. The assumption is that the data exists because the systems that generate it exist. The CRM has customer data. The ERP has transaction data. The helpdesk has support data. This is true. The question of whether that data is in a form that can train or power an AI model is a different question entirely — and it’s almost never asked until it becomes a blocker.
The gap between “we have data in a system” and “we have data an AI can use” is not a small technical gap. It is often a three to six month remediation project involving data integration, cleaning, standardisation and governance work that has nothing to do with AI and everything to do with the boring, unsexy infrastructure that should have been built years earlier.
What "AI-Ready Data" Actually Requires
AI-ready data has four properties. Accessibility — it can be queried and retrieved by an AI system without manual extraction. Consistency — the same concept is defined and formatted the same way across all sources. Completeness — the fields and records the AI needs are actually populated, not empty or sparse. Recency — the data reflects current reality rather than being months or years out of date.
In our experience, most enterprise datasets are partially accessible, inconsistently structured, significantly incomplete in the fields that matter most, and of mixed recency. This doesn’t make AI impossible — but it does make data preparation an explicit, scoped, resourced piece of work that must happen before the AI is built.
The Harder Conversation: Legacy Systems
Many organisations carry the legacy of technology decisions made five, ten or fifteen years ago. Core business data lives in systems that were built before modern data standards, before APIs were universal, and before anyone imagined feeding this data into machine learning models. The data is technically there. Getting it out, reliably and in usable form, is a significant engineering challenge.
This is not a problem AI solves. It’s a problem that must be solved before AI can be applied. The businesses that have invested in modern data infrastructure — cloud data warehouses, clean API layers, unified customer data platforms — move much faster on AI initiatives because the prerequisite work is already done. The businesses that haven’t face a choice: do the infrastructure work now, or scope AI initiatives to the data that is accessible today and accept the limitations that come with it.
What to Do About It
The practical answer is to assess your data landscape before you scope your AI initiative — not as an afterthought, but as step one. Understand what data exists, where it lives, what state it’s in and what gaps exist for the specific use case you’re pursuing.
Then scope the data work explicitly. Budget for it. Resource it. Don’t pretend it won’t be needed, because the cost of that pretence is always paid later — with interest.
The organisations that do this honestly tend to complete AI projects on time and on budget. The ones that don’t tend to double the timeline and then blame the technology.
The technology is rarely at fault.