There's a hope that shows up constantly, in nearly every industry, usually unspoken but occasionally said out loud in exactly these words: "maybe AI can help us finally sort out our data mess." It's an understandable hope, and it would be nice to say it were true more often than it actually is. Across a fair number of organizations that have tried exactly this, AI doesn't fix fragmented, inconsistent, poorly defined data. It automates whatever is already there, fragmentation included, and it does so faster and with more apparent confidence than the manual process it's replacing, which means an organization with unresolved data problems doesn't get a cleaner picture from an AI initiative. It gets the same confusion it already had, produced more quickly, presented more authoritatively, and considerably harder to catch before it does real damage.

Why AI feels like a shortcut past data work
The appeal of skipping straight to AI, past the unglamorous work of actually fixing data definitions and consistency, is easy to understand, and nobody who's felt it deserves any less credit for it. Data cleanup work is slow, thankless, and rarely produces anything visually impressive to show a leadership team. It's meetings about metric definitions, spreadsheet reconciliation, patient negotiation between departments that have quietly disagreed about a term's meaning for years. AI, by contrast, feels like it might let an organization leapfrog straight past all of that directly to insight: throw the existing data, messy as it is, at a sufficiently sophisticated model, and let the model itself sort out the patterns that human effort hasn't yet managed to untangle.
This hope isn't entirely irrational. Modern machine learning genuinely can find patterns in messy data that a human analyst working by hand might miss. What it can't do is know, on its own, that three plants are using the word "available" to mean three subtly different things, and therefore treat those three data streams as if they were measuring the same underlying reality. A model trained across inconsistently defined data doesn't resolve the inconsistency. It learns from all three definitions simultaneously, as if they were one, and produces a confident-looking output that's actually an unexamined blend of three different, incompatible realities.
What actually happens when you skip the data foundation

What actually happens, in the cases that play out this way, follows a fairly consistent pattern. The model trains successfully. There's rarely a technical failure at this stage, because the underlying algorithms are perfectly capable of finding some pattern in whatever data they're given, regardless of whether that pattern reflects anything real or useful. The model produces output that looks confident and specific, because that's simply what these systems do by default; a forecast or a recommendation rarely arrives hedged with visible uncertainty about the quality of its own inputs unless someone has deliberately built that hedging in.
The output gets used, at least initially, because it looks authoritative and the organization was hoping for exactly this kind of clear, confident answer. And then, at some point, sometimes weeks later, sometimes embarrassingly close to a decision that actually mattered, someone notices the recommendation doesn't match reality in a specific, checkable way, and starts digging into why. What they usually find isn't a flaw in the model's algorithm. It's that the model was trained on data that meant three different things depending on which part of the organization it came from, and nobody had told the model that, because nobody had fully realized it themselves before the AI initiative forced the question.
The minimum data readiness bar worth clearing first
None of this means an organization needs a fully mature, enterprise-wide data platform before attempting any AI initiative. That would be its own overcorrection, and one that delays real value for years chasing a level of data perfection most organizations will never fully reach regardless of how much they invest in it. What it does mean is that there's a minimum, specific bar worth clearing for the particular data a given use case actually depends on, before that use case gets built, rather than a generic aspiration toward "good data" across the whole enterprise.
That minimum bar is narrower than it sounds: a single, explicitly agreed definition for each metric the use case needs, documented somewhere everyone involved can see it; a known, checked level of consistency in how that metric is actually calculated across every site or team that would feed data into the model; and a realistic understanding of the metric's actual quality, its error rate, its gaps, its known quirks, rather than an assumption that it's fine because nobody's complained about it recently. Clearing this bar for one specific use case's specific data needs is a matter of weeks in most organizations, not years, precisely because it's narrow and specific rather than an attempt to fix everything at once.
Sequencing data work and AI work together, not one after the other
The useful sequencing isn't "fix all our data, then start on AI," a sequencing that guarantees years of delay chasing an unattainable standard of universal data cleanliness before any AI value gets delivered at all. It's identifying the specific data a particular, well-scoped use case needs, fixing that specific data to the minimum bar described above, and building the AI use case on that now-solid foundation, all within the same focused timeframe rather than as two entirely separate, sequential programmes. Each new use case brings its own specific data requirements into focus, and each one gets its own narrow, targeted data-readiness pass rather than waiting on a single, enterprise-wide data transformation to complete first.
This approach has a genuinely valuable secondary effect worth naming: the data work done for the first use case often turns out to be reusable for the second and third, because metric definitions, once actually agreed and documented for one purpose, tend to hold up reasonably well for adjacent purposes too. An organization that works this way ends up building its data foundation incrementally, use case by use case, funded by and directly justified by the AI value each one delivers, a far more sustainable pattern than a standalone data initiative that has to justify its own budget in the abstract, disconnected from any specific, visible outcome.
"Garbage in, garbage out," but faster and more convincing
The old software engineering adage "garbage in, garbage out" undersells what actually happens when you point a modern AI system at fragmented data, because it implies the garbage stays roughly as visible as it always was, just processed. What actually happens is closer to garbage in, confident garbage out, delivered faster and with more apparent authority than the manual process ever produced. A spreadsheet built by hand from inconsistent data tends to carry visible seams, a footnote, a caveat, a column someone flagged as "needs review," because the human building it was aware, at some level, of where the shakiness lived. A model trained on the same inconsistent data carries no such seams by default. It outputs a clean number with no visible hedging, because nothing in its design forces it to represent its own uncertainty about data it was never told to distrust.
This is precisely why AI on top of fragmented data is, in a real sense, worse than the manual process it replaces, rather than merely no better. The manual process's flaws were at least partially visible to the people producing it. The automated process's flaws are hidden behind a veneer of algorithmic confidence that makes them considerably harder to catch before someone acts on them, and considerably more embarrassing to unwind once they've already influenced a real decision.
How to spot this failure before it ships, not after
There's a specific, practical check worth running before any AI initiative goes live, and it costs an afternoon rather than months: take the exact fields the model will depend on, and for each one, ask whether it's calculated identically everywhere it will be sourced from, not whether it's called the same thing, but whether the actual calculation logic, the timing, and the inclusion or exclusion rules are genuinely identical. This is a more demanding question than it sounds, because two teams can use an identical field name while calculating it in meaningfully different ways, and that mismatch is invisible until someone actually traces the calculation logic rather than trusting the label.
A second useful check: pull a small, manageable sample of the raw source data from each site or system the model will draw from, and have someone who actually understands the operational reality, not just the data schema, eyeball it for anything that looks locally idiosyncratic. This kind of manual spot-check feels almost embarrassingly low-tech next to the sophistication of the model being built on top of it, but it reliably catches exactly the kind of quiet, locally reasonable divergence that no amount of statistical validation against the model's own training data will ever surface, because the model has no independent way of knowing what the data was supposed to mean in the first place.
The forecast that was wrong three different ways for three different reasons
Imagine building a forecasting model using data pooled from three plants, all apparently reporting the same metric: on-time delivery.
At the start, that sounds straightforward. The data appears comparable, the model trains successfully, and the resulting forecasts look reasonable. Nothing in the technical diagnostics suggests that there is a fundamental problem.
But now imagine that the three plants do not actually mean the same thing by “on time.”
At one plant, expedited orders are excluded from the calculation. At another, performance is measured against the originally promised delivery date. At a third, teams routinely work against a revised delivery date that has gradually become the practical operating standard.
Then add another complication: one plant has several months of missing history following a system migration, but that context never reaches the team building the model.
Individually, none of these issues may look dramatic. Together, they create three different versions of what appears to be one common metric.
The forecasting model does not create those inconsistencies. They already exist in the underlying processes, definitions, and data.
What the model does is combine them.
And because the result emerges as a single, polished forecast with an appearance of mathematical confidence, the underlying inconsistencies can become even harder to see.
That is where the risk lies.
A manual process may be slower and less sophisticated, but its uncertainty is often visible. An AI model can take fragmented definitions, incomplete histories, and locally accepted workarounds and compress them into one authoritative-looking output.
The lesson is not simply that data must be “clean” before AI can be applied. It is more fundamental: before asking whether the model is accurate, organizations need to ask whether the business reality represented by the data is actually consistent.
If three plants use three different meanings for the same KPI, the model may become very good at learning the inconsistency rather than resolving it.
Fix the definition before you fix the model

Fixing the metric definition, in every case like this, turns out to be both cheaper and more foundational than fixing the model itself, because the model was never actually broken. It was working exactly as designed, faithfully learning from whatever data it was given, and the data was the part that needed attention. Before your next AI initiative, name the specific data it will depend on, and clear the minimum readiness bar for that specific data, a single agreed definition, checked consistency, an honest understanding of quality, before building anything on top of it. Fix the definition before you fix the model. In nearly every case like this, the model was rarely the actual problem to begin with. The data it was faithfully learning from was.
Disclaimer
Industry situations in this chapter are composite illustrations unless explicitly attributed to a public source. They are not claims about any particular company, plant, vendor, or incident. External standards, research, and public case studies should be verified before publication. Implementations must be validated against local safety, quality, cybersecurity, regulatory, contractual, labour, privacy, and data-governance requirements. AI recommendations and autonomous actions should remain within clearly defined human authority, operational controls, and tested recovery procedures.
#DataQuality #ManufacturingAI #SteelIndustry #GarbageInGarbageOut #DigitalTransformation #SupplyChainAI #Industry40 #DataGovernance #OperationsExcellence #EnterpriseAI
Further reading
- Why AI Data Quality Is Key To AI Success
IBM, January 15, 2026
- Garbage In, Garbage Out: Making Datasets Ready for AI
Architecture & Governance Magazine, March 3, 2026
- Garbage In, Garbage Out: Why Data Quality Defines AI
V2Solutions, updated July 1, 2026
- Garbage In, Garbage Out — Why Bad Data Destroys Automation ROI
Parseur, updated September 8, 2026
- Why AI Data Quality Is the Top Business Problem in 2026
Sombra, updated May 25, 2026

