All insights
InsightsIA opérationnelledonnéesstart from anywhere

“Our data isn’t ready”: the prerequisite myth

“We’ll do AI once our data is clean.” Yet only 7% of companies say their data is completely ready — and waiting doesn’t move that number. The prerequisite is a myth, and it has a measurable cost.

Published on June 23, 20266 min readexecutive leveldata verified August 12, 2026

TL;DR

  • Only 7% of companies say their data is "completely ready" for AI (HBR Analytic Services for Cloudera, 2025 survey, 2026 report): if readiness were a prerequisite, 93% would never be allowed to start.
  • Gartner predicts that through 2026, 60% of AI projects unsupported by AI-ready data will be abandoned — yet that readiness can only be assessed use case by use case.
  • The readiness illusion is measured: 84% of IT leaders say they are confident in their data, while 18% say it is fully governed (Cloudera Data Readiness Index, 2026).
  • The ISO/IEC 5259-4:2024 standard defines data quality for machine learning as a continuous process, not a "ready / not ready" state — and trust in data rises from 50% to 71% once governance is in place (Drexel LeBow / Precisely, 2026).
  • Our 30+ delivered projects started from seven real starting points, none of which was "ready" — from the Monday-morning Excel file to six systems that never talked to each other.
01

Nobody is ready, and that is good news

"Our data isn't ready for AI." The sentence is careful, sincere — and wrong in its premise: that modern data is the prerequisite of the journey. We argue the opposite: it is the result. Waiting for your data to be ready is the surest way for it never to be.

The numbers used to justify waiting actually say something else. Gartner predicts that 60% of AI projects without adequate data will be abandoned through 2026. But only 7% of companies consider their data completely ready: if "ready" were the entry ticket, 93% of organizations would be disqualified on the spot. Gartner adds that 63% of organizations either lack AI-adequate data management practices or do not know whether they have them — the prerequisite is not just rare, it is unverifiable.

02

"Ready for what?" — the missing question

Research has said it for thirty years: data quality does not exist in the absolute. Neil Lawrence (Data Readiness Levels, 2017) shows that readiness is assessed against a given task. Wang and Strong defined it by usage back in 1996:

"Quality data are data that are fit for use by data consumers." — Wang & Strong, Journal of Management Information Systems, 1996.

The ISO/IEC 5259-4:2024 standard takes this logic to its conclusion: data quality for machine learning is a process framework — assess, improve, repeat — not a state to reach before starting. "Not ready" remains an incomplete sentence until you answer: ready for which use? This is why the schema is the deliverable: the useful data model emerges from real usage.

The measured gap between the 84% who feel confident and the 18% who are actually governed confirms the trap: waiting to be ready means waiting for a state almost nobody reaches.

03

Start narrow, prove, expand

Three independent bodies of work converge. RAND recommends staying focused "on the problem to be solved, not the technology." The 2025 MIT NANDA report (methodology debated, but convergent) observes that organizations succeeding with generative AI start narrow, on a non-critical process, prove the value, then expand. And Drexel LeBow / Precisely measures the trust gain from 50% to 71%: you don't govern your data in order to start, you govern it by starting.

This is the start from anywhere mechanic: begin with what exists, pick a narrow use case, ship, and let quality rise. It sits at the heart of how we build.

What doing nothing costs — 42% of companies report that more than half of their AI projects were delayed, underperformed or failed because of data readiness issues (study commissioned by Fivetran, 2025). Meanwhile, AI adoption among EU enterprises went from 8.1% in 2023 to 20.0% in 2025 (Eurostat). Waiting is not neutral: it is a decision, with a price.

04

Seven starting points, zero data warehouse

Our thirty-odd delivered projects did not start from modern infrastructure, but from seven real situations — an Excel file updated by hand every Monday, a process that existed only in people's heads, six systems that never talked to each other… None of them was waiting for a data warehouse.

The MCP protocol — an open-source standard that connects AI to existing systems, "like a USB-C port for AI applications" — tools exactly this: AI without a data warehouse, querying the systems already in place instead of replacing them. A school ERP with a SOAP API became queryable in natural language this way; a finance team now publishes its dashboards by talking to Claude, in minutes. The logic mirrors the strangler fig applied to legacy systems: you wrap what exists, you don't wait for it.

05

Where to start: seven observed trajectories

Where to start? Where you are. Here are our seven starting points and what each became.

Starting pointConcrete first stepObserved outcome
An Excel file updated by hand every MondayGenerate the exact same table automaticallyA multi-tenant SaaS with row-level isolation in Postgres (RLS)
Spreadsheets used as a databaseStructure the columns into tables and views, in the same tool17 custom Airtable extensions in production, within a year
A legacy business app with a SOAP APIA read-only connector, without touching the applicationA school ERP linked to a living repository, queryable by AI
A CRM nobody pulled data out ofBring calls, emails and SMS into the tool already openA full CRM inside Airtable, adopted immediately
Dashboards sent as email attachmentsPublish the same content behind SSO instead of emailA secure Next.js/Supabase portal + a Claude Desktop extension (MCP)
A process that existed only in people's headsWrite it down, then model it as a shared schemaA documented process, then tooled into a production application
Six systems that never talked to each otherOne first flow between two systems, driven by a mappingA sync engine: 6 CRMs/ERPs, one repository, zero re-keying

Every row started with data that was not ready for AI in the sense of the surveys above, and ended with cleaner data than it began with — because usage demanded it.

06

The limits of this approach

Starting with what you have is not a license to start carelessly. Some uses do impose real prerequisites: personal data requires a legal basis and controlled access before any processing; critical uses — payroll, invoicing, health — deserve serious upfront qualification. Multiplying small starts without ever industrializing produces a collection of prototypes: the first step only has value within a trajectory, with a schema, governance and progressive expansion. Finally, our seven trajectories are observed cases, not a statistical law: they prove what is possible, not what is automatic.

Key takeaways

  • "Ready" is only defined relative to a use case: until the use case is chosen, the question is undecidable.
  • The documented sequence that works: narrow scope, proof of value, expansion — governance gets built along the way.
  • A hand-maintained Excel file is a legitimate starting point: several of our production systems literally started there.

The right question is not "is our data ready?" but "which first use case — narrow and useful — do we choose?". The schema, the quality and the governance are built by walking. Ownward helps companies perform better through technology — and above all, take back control.

Sources

The first-person project facts come from our own work: Ownward internal data, 2026.

Data verified on August 12, 2026.

All trademarks belong to their respective owners. This article is neither sponsored nor endorsed by the vendors mentioned.

Is this on your desk right now?

Tell us where you stand. We reply with concrete elements — what we would do first, in your business.

Talk about your situation

Keep reading

All insights