· Mohamed Ben Haddou · AI readiness · 8 min read
Are you AI-ready? Part 2 — Data Foundations: "our data isn’t ready" is half true, and fixable
The most common reason an AI idea dies is a sentence: "our data isn’t ready." It is usually half true — and the half that is true is smaller and cheaper to fix than people fear. The second dimension of AI readiness: what "ready" means, per use case, and why your documents may be your biggest asset.

Second of six articles on the dimensions of the Mentis READY Framework. Part 1 covered Strategy & Value — what AI is for, who owns it, how ideas get ranked. This one is about the dimension that kills more of those ideas than any other.
The sentence that ends most AI projects
“Our data isn’t ready.” We hear it in almost every first conversation, usually from the person who knows the systems best, and usually with relief — because it closes the topic. Nobody has to decide anything if the data isn’t ready.
Here is what twenty years of building models on other people’s data has taught us: the sentence is almost always half true. The half that is true — some data is missing, some is dirty, some is locked in a vendor’s system — is real. The half that is false is the word “our”: it treats the organisation’s data as one undifferentiated lake that is either ready or not. It isn’t. Data is ready for a use case, or it isn’t. And for the use case that matters most, the gap is usually smaller, and cheaper to close, than the sentence implies.
That reframing is the whole point of the Data Foundations dimension. We do not assess whether your data is ready. We assess whether the data each priority use case needs is findable, accessible, good enough, and lawful to use — and what it would take to get there.
What “ready” means, concretely
For a given use case, four questions decide readiness:
- Does it exist? Is the signal actually recorded somewhere — not “we have a CRM” but “the field that would predict churn has been filled in, consistently, for three years”?
- Can we reach it? Is it accessible through something better than a quarterly export from a vendor, with the rights to use it for this purpose?
- Is it good enough? Not perfect — good enough for the decision. Forecasting tolerates noise that credit scoring does not.
- Can we keep it that way? Is there an owner, a quality rule, a refresh path — or will the model decay the month after launch?
Notice what is not on the list: a data lake, a master-data programme, a two-year platform migration. Those may be worth doing. They are not prerequisites for the first use case, and treating them as such is how organisations spend eighteen months “preparing” and ship nothing.
The five maturity levels of Data Foundations
| Level | What it looks like in practice |
|---|---|
| 1 · Ad hoc | Nobody can say what data exists, where, or in what state. Quality is unknown. Documents are scattered. |
| 2 · Experimenting | Key data is locked in silos or vendor systems; known quality problems, never measured; files stored but unsorted. |
| 3 · Structured | Data is accessible with manual effort; quality measured in places; documents partly organised in a DMS. |
| 4 · Managed | Catalogued, accessible through governed platforms; owners and quality rules; documents indexed and searchable. |
| 5 · Optimised | Curated and quality-monitored against standards; unstructured estate structured, access-controlled and AI-ready. |
Two patterns recur across the organisations we assess. First, structured data is usually at level 3 — reachable with effort — while the unstructured estate (contracts, reports, procedures, emails) sits at level 1 or 2, even in firms with excellent databases. Second, the organisations that rate their data worst are often the ones with the most usable data for generative-AI use cases, because their problem is access and organisation, not absence. More on that below.
Three questions that tell you where you are
These are the three the scorecard asks for this dimension.
Can you find and access the data your priority use cases need? “We don’t really know what we have” is level 1. “Locked in silos or with vendors” is 2 — and a more common answer than it should be, because many contracts never secured the right to extract and reuse the data. “Accessible with manual effort” is 3 and perfectly workable for a pilot; the jump to 4 is governed access, and that is a platform investment to make once a use case has proven its value, not before.
How would you describe your data quality? “Unknown” and “known problems, never measured” are 1 and 2, and they share a fix: measure. Completeness, consistency, timeliness of the specific fields a use case needs — one afternoon of profiling, one page of results. Organisations are regularly surprised in both directions.
Your unstructured estate — documents, contracts, emails? “Scattered, no overview” through “indexed and searchable” to “structured, access-controlled, AI-ready”. For many organisations this is the question that matters most, and it is the one they have thought about least.
The overlooked asset: your documents
Most AI conversations of the last two years have been about generative AI, and generative AI in an enterprise runs on documents: procedures, contracts, reports, tickets, correspondence. Retrieval-augmented generation — an assistant that answers from your own documents with citations — is the most frequently shippable use case we see in regulated sectors, precisely because it needs no labelled data and no clean relational model. It needs three things instead: the documents, permission to read them, and a way to keep the “who may see what” intact.
That last point is where data readiness meets governance. A document assistant that ignores access rights is a data breach with a chat interface. A document assistant that respects them — per user, per matter, per patient — is a confidentiality-preserving tool that legal and the DPO can sign. The difference is a data-foundations decision made at design time, and it is why the data readiness scorecard in our assessment has a line for “permissions preserved in the pipeline” next to “documents indexed”.
Regulated organisations in Belgium and Luxembourg add one more requirement: those documents, and the vectors and logs derived from them, may not leave the EU — often may not leave the building. That is an architecture decision (Part 3 of this series), but it starts as a data decision: classify what is sensitive before anyone builds a pipeline.
Where the law lands on data
Two texts shape data readiness for AI in Europe, and they reward the same discipline.
GDPR asks for a lawful basis and purpose limitation: data collected for one reason cannot silently become training data for another. Data minimisation and the rules on automated decisions apply to the model’s inputs and outputs. None of this forbids AI; all of it rewards knowing exactly which personal data a use case touches — the inventory you need anyway.
The EU AI Act goes further for high-risk systems: article 10 requires data governance practices — relevance, representativeness, error checks, bias examination — for training, validation and test data. If a credit model, a recruitment screen or an eligibility system is on your list, “our data isn’t ready” stops being an excuse and becomes an obligation with a deadline. The good news: an organisation at level 4 on this dimension has most of article 10 in place already.
What to do in the next 30 days
If you recognise yourself at level 1 to 3:
- Pick one priority use case — then audit only its data. Existence, access, quality, rights, for the specific fields or documents it needs. One to two weeks, one page of findings. Do not boil the lake.
- Fix access before quality. Data you cannot reach cannot be measured, let alone cleaned. Extracting a table from a vendor system or indexing a document folder is often the single highest-value step.
- Inventory the unstructured estate. Where the documents live, who owns them, what is sensitive, and what the permission model is. You will likely discover your best generative-AI use case in the process.
- Name data owners for the handful of sources that matter. Not a data-governance programme — two or three people accountable for three or four sources.
Do that and the sentence changes from “our data isn’t ready” to “the data for this use case is ready enough, here is what the next one needs” — which is a plan.
Where this sits in the bigger picture
Data Foundations is the second of six dimensions. Next: Architecture & Infrastructure — what your stack needs to run AI in production, and the sovereign, on-premise path for organisations that cannot send data to US clouds. Then Governance & EU AI Act compliance, People & Operating Model, and Security & Trust. Together they produce the maturity radar at the heart of the AI Readiness Assessment, our four-week, fixed-price diagnostic for mid-market organisations in regulated sectors.
Want your own reading? The free AI Readiness Scorecard asks the three data questions above — and fifteen more across the other dimensions — and gives you your maturity radar in four minutes. No account, no sales call attached; if you want a second opinion on your results, leave an email and we will write back within a business day.
Mohamed Ben Haddou is the founder of Mentis Consulting (Brussels, ULB spin-off, since 2005) and an Independent AI Expert for the European Commission.
