
Your CFO wants a straight answer to a simple-sounding question. Where does next quarter's AI budget go, into cleaning up the data or into building the AI itself. You have maybe twenty minutes in that meeting, and the uncomfortable truth is that most people in the room have not actually agreed on what problem the AI is supposed to solve in the first place.
That gap gets papered over constantly. Someone says "fix the data first," because it sounds responsible. Someone else says "just start building," because leadership wants visible progress this quarter. Both answers can be defended in a slide, and both skip past the decision that actually determines whether this initiative works.
There is a real answer to which comes first, and most organizations are asking a different question entirely.
Why Data Readiness Is Not the First Question to Answer
That real question starts with an instinct almost everyone in the room already has. Ask ten data leaders whether to fix data readiness before starting an AI initiative, and most will say yes without hesitating. It is the responsible answer. It is also, in a meaningful number of cases, the wrong first move. Data readiness still matters. The problem is that jumping straight to it assumes a decision has already been made that usually has not.
RAND Corporation interviewed 65 experienced data scientists and machine learning engineers to understand why AI projects fail. Their finding cuts against the instinct in that CFO meeting. Eighty-four percent of interviewees named leadership-driven failure as the primary reason projects collapse, ahead of data quality itself. Business leaders and technical teams frequently disagree, sometimes without realizing it, about what problem the AI is actually meant to solve or how success will be measured.
That is a scoping problem, sitting upstream of the data-first-or-AI-first debate entirely. Skip it, and the sequencing decision you make next stops mattering, because you end up optimizing the wrong thing regardless of which path you pick.
What an Unscoped AI Problem Actually Looks Like
That wrong optimization has a shape to it, and it rarely announces itself. Leadership asks for a model that predicts customer churn. The technical team builds one, optimized for prediction accuracy. Three months later, leadership asks why churn has not improved, because what they actually needed was a model that identified which at-risk customers were worth the cost of retaining, a narrower and more useful target than simply predicting who was likely to leave.
Both sides did their jobs, each one correctly by their own definition. The project failed because the target metric and the business objective were never the same thing, and nobody caught the mismatch until the work was already done.
A simple test cuts through this before a single dollar gets spent on data or infrastructure. Can leadership and the technical lead each independently write down, in one sentence, what business outcome this AI initiative is meant to change. That alignment has to happen before the sequencing conversation starts.
Two Different Data Problems Once the Use Case Is Scoped
Once leadership and the technical team agree on what problem the AI is solving, the real sequencing question finally becomes answerable, and it is narrower than most people expect. It comes down to which of two situations you are actually in.
In the first, the data your use case needs already exists somewhere in the organization, but it cannot be trusted. It is inconsistent across systems, duplicated, ungoverned, or simply wrong in ways nobody has caught yet.
In the second, the data does not exist yet in usable form. The specific use case is simply new, and the organization has never needed to capture, structure, or retain that information before now.
These are two different problems, each calling for its own first move.
When Data Exists But Cannot Be Trusted
This is the quality-bottlenecked path, and it is the one most content on data readiness already addresses, for good reason. Picture a customer record that shows three different names for the same account, a product catalog with duplicate entries carrying different prices, or a sales figure that changes depending on which system you pull it from.
RAND's interviews found this pattern repeatedly. Organizations believe they have good data because they get weekly reports, without realizing that the data behind those reports was never built to support a new purpose.
When this is the actual bottleneck, fixing the data first is the correct sequence. The reason has little to do with feeling responsible. Building AI on top of untrustworthy data produces a faster path to a model nobody trusts, and that is a more expensive failure than a delayed start.
When AI Needs Data That Does Not Exist Yet
The availability-bottlenecked path looks different, and it gets far less attention in most data readiness conversations, which is part of why organizations mishandle it. Here, the data is simply incomplete. It does not exist yet, because the use case is asking a question the organization has never needed to answer before.
Picture a manufacturer launching a new product line and wanting AI to flag supplier risk before a shipment gets delayed. The company has years of clean, trustworthy sales data. It has never once tracked supplier lead-time variability at the granularity this use case needs, because nobody needed that number before. Cleaning existing records will not surface a number that was never captured, so the only real fix is building the data collection itself.
This is a real and growing constraint, and a significant one. Stanford HAI's 2026 AI Index flags a broader version of this same problem at industry scale, noting genuine concern among researchers that the supply of high-quality training data itself may be approaching structural limits within the next several years. Synthetic data has not proven a full substitute. If data availability is becoming a binding constraint even at the scale of the entire AI industry, it deserves equal weight to quality inside a single organization, as a different problem in its own right.
When this is the actual bottleneck, spending months on a data governance program before touching the AI initiative solves a problem you do not have. The right move is building the specific data capture and structure the use case needs, alongside building the AI capability itself, on the same timeline instead of a strict sequence.
The Real Cost of Sequencing Data Readiness and AI Wrong
Whichever path actually applies, guessing wrong in either direction carries a real cost, and it shows up differently depending on which default an organization reaches for. Organizations that default to "clean everything first" often spend on a broad governance initiative when the actual use case only needed a narrow, targeted fix, delaying value for months over a problem that was smaller than assumed.
Organizations that default to "build now, fix data later" run into the opposite cost. They discover the trust problem only after the model is in production, when a bad prediction has already reached a customer or a decision, and the fix now requires unwinding both the data and the deployed system at once.
A third pattern is just as common, and often the most expensive of all. Organizations split the budget, funding a partial data cleanup and a partial AI build at the same time, without ever diagnosing which one the use case actually needed.
Both efforts end up underfunded, and neither one finishes. The AI initiative stalls waiting on data that was never fully fixed, and the data initiative loses priority the moment the AI timeline slips, leaving the organization with the appearance of progress on two fronts and real progress on neither.
All three defaults carry real risk. Diagnosing which path actually applies costs far less than guessing wrong in any of them.
How to Know Which Path Applies to Your Organization
The scoping check comes first. Confirm that leadership and the technical team can each state the target business outcome in the same terms, and resolve any gap in that alignment before anything else.
Once scoped, the diagnostic question is direct. Does the data your use case needs already exist somewhere, just inconsistent or untrustworthy, or does it not exist yet in usable form at all. That single distinction determines whether the right next move is a data quality initiative or a parallel build.
This is the exact question a Data & AI Strategy engagement is built to answer, replacing the guesswork with a clear sequencing plan.
Key Takeaways
- RAND's interviews with 65 experienced AI practitioners found leadership-driven misalignment was the most cited primary cause of AI project failure, named by 84% of interviewees, ahead of data quality itself.
- Confirming what problem the AI use case solves, in terms both leadership and the technical team agree on, has to happen before any sequencing decision is meaningful.
- Once scoped, sequencing splits into two distinct paths: fixing data that exists but is not trustworthy, or building toward data that does not exist yet for the specific use case.
- Stanford HAI's 2026 AI Index confirms data availability is a serious, industry-recognized constraint carrying equal weight to quality.
- Splitting the budget between partial data cleanup and a partial AI build, without diagnosing which one the use case actually needs, is common in practice and often costs more than committing to either path deliberately.