AI That Doesn't Require a Data Lake First: A Steel Guide

11 min read

Share this page

Choose where to share this page.

AI That Doesn't Require a Data Lake First: A Steel Guide

You don't need a multi-year data platform to start with AI. Here's how to build a real first use case on data your plant already trusts.

dattarajsandur.com

Share via

AI That Doesn't Require a Data Lake First: A Steel Guide
AI That Doesn't Require a Data Lake First: A Steel Guide

Description

You don't need a multi-year data platform to start with AI. Here's how to build a real first use case on data your plant already trusts.

Accordion controls

There's a specific sentence that shows up often enough, in almost identical form, across enough different steel and metals plants and adjacent heavy-industry operations, that it's recognizable as a distinct pattern rather than an isolated opinion: "we can't really do anything meaningful with AI until we get our enterprise data platform built." It's usually said with a certain resigned confidence, as though it's simply an established fact of the technology landscape rather than an assumption worth questioning. And it's usually wrong, in a way that costs organizations years of delay chasing a data foundation that a genuinely useful first AI use case never actually needed in the first place.

The Trusted Logbook
The Trusted Logbook - AI Generated

Meaningful AI value can start on a narrow, well-scoped dataset the organization already trusts, well before any enterprise-wide data platform exists, and waiting for the platform first is, in most of the cases seen, a slower and more expensive path to the same eventual destination than starting narrow and building the broader foundation alongside the AI work rather than strictly before it.

Why "we need a data lake first" delays value for years

The logic behind "we need a data lake first" is intuitively appealing: AI needs good data, a data lake is where an organization consolidates and cleans its data, therefore build the data lake first and the AI will follow naturally once the foundation is solid. The problem with this logic isn't that it's wrong in principle. A mature, well-governed enterprise data platform genuinely is valuable, for many purposes well beyond AI. The problem is the sequencing, and the multi-year timeline that a genuine enterprise data platform typically requires, standing entirely between the organization and any AI value at all during that entire build period.

Enterprise data platforms are large, complex, genuinely multi-year undertakings almost by their nature. They involve dozens of source systems, competing priorities across business functions, and a scope that tends to expand as more stakeholders discover requirements they'd like included. Treating this large undertaking as a prerequisite for any AI value at all means the organization spends years producing nothing tangible from its AI ambitions, while competitors who started narrower are already several use cases into a compounding track record of real, delivered value. The data lake, once finished, may well be a genuinely valuable asset, but "once finished" is doing an enormous amount of work in that sentence, and it's rarely a short wait, and it's very often a wait measured in years rather than the months most leadership teams initially budget for it.

Picking a use case scoped to data you already trust

Cross-Checking the List
Cross-Checking the List - AI Generated

The alternative isn't to abandon the goal of a solid, mature data foundation. It's to stop treating it as a precondition for starting, and instead pick a first AI use case scoped specifically to data the organization already trusts today, however narrow that trusted dataset happens to be. Every plant, without exception, has at least one dataset that's genuinely well-maintained, consistently defined, and actively relied upon by the people who use it daily, a maintenance log, a quality inspection record, a specific production parameter that's been tracked reliably for years because it matters to someone's daily job.

That trusted dataset, narrow as it is, is a perfectly legitimate foundation for a first AI use case, provided the use case is scoped to what that specific data can actually support rather than stretched to cover a broader ambition the data was never built to serve. This is a smaller, less glamorous starting point than an enterprise-wide initiative, and that's precisely the point. It's achievable now, on data that's already proven itself trustworthy through years of actual daily use, rather than on a future, aspirational dataset that doesn't yet exist in any usable form.

What a narrow first success actually proves

A narrow first AI success, built on data the organization already trusts, proves something considerably more valuable than the specific use case itself might suggest on paper. It proves, concretely and visibly, that the organization can actually take an AI initiative from idea to working, adopted tool, a capability that sounds obvious in the abstract but that a meaningful number of organizations, more often than not, have never actually demonstrated to themselves before their first attempt. It surfaces, in a low-stakes setting, exactly what collaboration between the technical team, the data owners, and the end users actually requires in this specific organization's culture, lessons that are far cheaper to learn on a narrow, forgiving first use case than on an ambitious one where the stakes of getting the collaboration wrong are considerably higher.

And it builds, gradually and concretely, the specific organizational muscle, trust between technical and operational teams, a shared vocabulary for talking about AI recommendations, a track record people can point to, that every subsequent, more ambitious use case will need to draw on. None of that muscle gets built by waiting for a data platform. All of it gets built by actually shipping something real, however modest, and learning from what happens next.

Growing the data foundation alongside the AI programme, not before it

The genuinely productive sequencing runs the two workstreams in parallel rather than treating one as a strict prerequisite for the other. The narrow first use case ships on the trusted dataset that already exists. Separately, and without blocking that first use case, the organization can begin the broader, more deliberate work of building out its data foundation, informed, usefully, by exactly what the first few AI use cases actually needed, rather than built speculatively against a generic, abstract notion of what "good enterprise data" should look like.

This sequencing has a real, practical advantage over the platform-first approach: each new AI use case that comes along naturally expands the scope of data that needs to be brought up to a trusted standard, and that expansion happens driven by demonstrated, specific need rather than by a theoretical completeness argument that tends to make enterprise data projects balloon in scope over time. The data foundation still gets built, and often ends up considerably more fit-for-purpose than a foundation built speculatively in advance, but it gets built as a series of concrete, justified expansions rather than as a single, multi-year prerequisite standing between the organization and any AI value at all.

How to find your own trusted dataset this week

If you want to identify your own organization's version of that trusted maintenance dataset, a short exercise helps more than an abstract inventory of every system in your landscape. Ask a handful of experienced people, across different functions, one specific question: which piece of data in your daily work do you personally check, rely on, and trust without a second thought, because you've watched it prove itself reliable over years of actual use? The answers tend to cluster quickly around a small number of genuinely well-maintained datasets, often not the ones that show up on an official enterprise data inventory as "strategic," but the ones a specific team has quietly kept clean because their own daily work depends on it being accurate.

Once you have a short list of candidate trusted datasets, apply a simple filter before committing to one as a first use case foundation: is it consistently defined and captured the same way regardless of which shift or which person is entering it, is it available at a frequency that matches how often the target decision actually needs to be made, and is there a real, specific decision nearby that this data could plausibly inform well. A dataset that passes all three is very likely a strong foundation for a first, narrow AI use case, and finding one usually takes days of conversation, not months of enterprise data cataloguing.

What this approach doesn't solve, and why that's fine

It's worth being honest about the limits of this narrow, trust-first approach, so it doesn't get oversold as a substitute for eventually doing the broader data work. A use case built on one trusted, narrow dataset won't, on its own, solve cross-plant reporting inconsistency, won't unify master data across systems, and won't give leadership the kind of enterprise-wide visibility a mature data platform eventually can provide. None of that is the goal of the narrow first step, and treating it as a failure of the approach misunderstands what it was meant to accomplish. Its job is to deliver real value quickly, build organizational muscle and trust, and generate concrete, specific lessons that make the eventual broader data investment better targeted when the organization does undertake it, not to replace that broader investment entirely.

The plant that chose a maintenance log over a two-year platform build

Imagine a plant debating two very different paths to AI.

The first is the familiar enterprise route: spend perhaps two years building a comprehensive data platform, harmonising systems, standardising definitions, integrating sources, and only then begin experimenting seriously with AI.

The second is much narrower: identify one business problem where useful, trusted data already exists and see whether AI can create value there first.

Now imagine the plant choosing the second path.

Instead of searching for the most ambitious AI opportunity, the team starts with a maintenance dataset that engineers have been maintaining consistently and using in their daily work for years. It may not be the richest dataset in the enterprise, but there is one important advantage: everyone who works with it already trusts it.

A focused maintenance-prioritisation use case is built around that data.

Within a quarter, the maintenance team has something genuinely useful in its hands, a recommendation that helps determine which pending maintenance activities deserve attention first.

It is not an enterprise-wide transformation.

It does not solve every data problem.

And it certainly does not remove the need for a broader data foundation.

But it creates something the larger platform programme cannot create by itself: evidence of value in actual operations.

The maintenance team can see what AI is doing with data they already understand. The organization begins learning which data attributes actually matter, where quality gaps become important, and what information needs to be available at the moment a decision is made.

Those lessons can then feed back into the broader data-platform programme.

The enterprise data conversation does not have to stop. It can continue in parallel, but now with considerably more clarity.

Instead of asking, “What data might AI need someday?”, the organization can start asking, “What did this real decision actually require?”

That distinction matters.

A large data foundation built entirely in anticipation of future use cases risks becoming an exercise in theoretical completeness. A narrow AI use case built on trusted data can help reveal which parts of that foundation genuinely deserve priority.

The lesson is not “skip the enterprise data platform.”

It is: do not assume that every organization must finish building the perfect data foundation before it is allowed to create its first useful AI outcome.

Sometimes the better path is to start with one trusted dataset, improve one meaningful decision, learn from it—and let real operational experience inform the broader data journey.

Start with the data you already trust, not the data platform you don't have yet

Two Timelines to Value
Two Timelines to Value - AI Generated

Meaningful AI value doesn't require waiting for a mature enterprise data platform to materialize first. It requires finding the narrow, well-scoped dataset your organization already trusts, and building a genuinely useful first use case on exactly that foundation, however modest it seems next to the scale of an eventual enterprise data ambition. Start with the data you already trust, not the data platform you don't have yet. The broader foundation can, and should, grow alongside the AI programme, informed by what it actually needs, not built speculatively, for years, before any real value gets delivered at all to anyone waiting on it.

Disclaimer

Industry situations in this chapter are composite illustrations unless explicitly attributed to a public source. They are not claims about any particular company, plant, vendor, or incident. External standards, research, and public case studies should be verified before publication. Implementations must be validated against local safety, quality, cybersecurity, regulatory, contractual, labour, privacy, and data-governance requirements. AI recommendations and autonomous actions should remain within clearly defined human authority, operational controls, and tested recovery procedures.

#SteelIndustry #ManufacturingAI #DataStrategy #DigitalTransformation #SupplyChainAI #Industry40 #DataLake #OperationsExcellence #EnterpriseAI #ManufacturingLeadership

Further reading

Continue reading