Why Your First AI Pilot Should Be Boring

11 min read

Share this page

Choose where to share this page.

Why Your First AI Pilot Should Be Boring

A modest, well-scoped first AI pilot builds more durable trust than an ambitious flagship one. Here's how to choose the right first target.

dattarajsandur.com

Share via

Why Your First AI Pilot Should Be Boring
Why Your First AI Pilot Should Be Boring

Description

A modest, well-scoped first AI pilot builds more durable trust than an ambitious flagship one. Here's how to choose the right first target.

Accordion controls

If you asked most leadership teams to design their organization's first AI pilot from a blank sheet of paper, an unusually high number of them would gravitate toward something genuinely ambitious: a cross-plant scheduling optimizer, an end-to-end demand-to-supply recommendation engine, something with enough scope to make a real, visible splash in the next board update. The appeal is completely understandable, and enough of these planning conversations, across enough different industries, make exactly how it happens easy to see. The first pilot is usually the moment an organization gets to demonstrate, to itself and to everyone watching, that it's genuinely serious about AI. A narrow, modest pilot doesn't feel like it's demonstrating much of anything. Put plainly, it's usually perceived in the room as a little underwhelming.

The Unglamorous Shift
The Unglamorous Shift - AI Generated

This instinct is almost always a mistake, and the organizations who resist it, who deliberately choose a boring first pilot over an ambitious one, end up scaling AI faster and more durably than the ones who go big first. Not because ambition is bad. Because trust, once it's damaged by a visible, high-profile early failure, is extraordinarily expensive to rebuild, and an ambitious first pilot carries a meaningfully higher risk of exactly that kind of failure than a modest one does.

The appeal (and danger) of the flagship pilot

The flagship pilot is seductive for reasons that have very little to do with whether it's actually the right technical or organizational choice. It photographs well in a leadership deck. It gives a sponsoring executive a genuinely compelling story to tell upward, to the board, about the scale of ambition the organization is bringing to AI. And it often targets a problem that's been frustrating the organization for years, which makes solving it, or appearing to be on the verge of solving it, feel like a much more satisfying use of a first AI initiative than something narrower and less dramatic.

The danger sits directly underneath that appeal, and it's structural rather than incidental. An ambitious first pilot, almost by definition, touches more systems, more stakeholders, more edge cases, and more organizational complexity than a narrow one, which means there are simply more places for something to go wrong, and less institutional experience, at this early stage, to catch those problems before they become visible. When an ambitious first pilot fails, and a meaningful share of them do, publicly, in front of exactly the leadership audience that was watching most closely, the damage isn't contained to that one initiative. It becomes the story the organization tells itself about AI in general: "we tried it, and it didn't really work here." That story is remarkably durable, and remarkably hard to argue against once it's taken hold, however unfair or oversimplified a summary it is of what actually happened.

What makes a pilot "boring" and why that's good

Trusting the Small Thing
Trusting the Small Thing - AI Generated

A "boring" pilot, in the sense meant here, isn't one that lacks ambition in the long run. It's one that's deliberately scoped narrow enough, on a decision small enough and well-understood enough, that it has a genuinely high probability of quietly working the first time. It touches one decision, owned by one clear person, using data the organization already trusts, with a low, contained cost if the recommendation turns out to be wrong in some early edge case nobody anticipated.

Boring, in this specific sense, is a feature rather than a limitation. A boring pilot doesn't need to convince skeptics with the scale of its ambition, because it isn't trying to win the argument about AI's potential in the abstract. It's trying to establish, concretely and undeniably, that this particular organization can build something that works, reliably, on real data, in a real decision, without drama. That's a much lower bar to clear than "transform how we schedule across all our plants," and clearing it convincingly is worth far more, in terms of durable organizational trust, than an ambitious attempt that stalls halfway and leaves everyone arguing about whether it was really a failure or just needed more time.

Picking a decision small enough to prove, big enough to matter

The art of choosing a genuinely good first pilot lies in finding a decision that's simultaneously small enough to prove quickly and cleanly, and consequential enough that succeeding at it actually matters to someone with real influence in the organization. A decision that's small but genuinely trivial, one nobody cares about either way, produces a technical success that persuades nobody, because succeeding at something inconsequential doesn't build the kind of credibility a second, larger pilot will need to borrow against.

A useful filter: does this decision currently get made by one identifiable person, with data that's already trusted rather than data that would need to be cleaned up first, at a frequency high enough that you'll get meaningful feedback within weeks rather than months? And critically, if the recommendation is wrong in some early, unanticipated way, is the resulting cost contained and recoverable, rather than something that shows up as a serious operational or financial problem the whole organization notices? A maintenance-alert prioritization, a single-line quality-exception flag, a narrow reorder-point recommendation for one well-understood product family, these tend to satisfy all four conditions simultaneously, which is exactly why they make unglamorous, genuinely excellent first pilots.

From boring pilot to confident scale-up

The real payoff of a boring first pilot isn't the pilot itself. It's the credibility and the specific, hard-won organizational learning it generates for everything that follows. A pilot that quietly works, on real data, in front of the specific person who has to trust its output every day, produces something an ambitious pilot rarely manages to produce cleanly: a genuine, first-hand testimonial from someone with no institutional incentive to oversell the result, because they were skeptical going in and are now describing what actually happened rather than what leadership hoped would happen.

That credibility becomes the currency the second, more ambitious pilot gets to spend. "This is the same approach that worked for the maintenance team last quarter" is a vastly more persuasive opening line than "trust us, this new system will handle something considerably more complex than anything we've attempted before." The boring pilot isn't a smaller version of the ambitious one. It's the foundation the ambitious one needs in order to be believed.

How to know when you've earned the right to be ambitious

A reasonable question at this point is how long an organization should stay in "boring pilot" mode before attempting something more ambitious, and the honest answer is that it's less about elapsed time and more about specific signals worth watching for. The clearest signal is spontaneous advocacy: does the person who uses the first pilot's output regularly mention it, unprompted, to colleagues, or does someone have to keep reminding them it exists? Genuine trust shows up as the tool becoming part of someone's normal routine without active promotion, not as a static adoption number on a dashboard that nobody actually cross-checks against real behavior.

A second signal worth watching is whether the organization can now answer, with real confidence, questions it couldn't have answered honestly before the first pilot: how does this specific team actually respond when a recommendation turns out to be wrong? What did it take, concretely, to get IT, the business team, and the data function to collaborate smoothly enough to ship something on a reasonable timeline? Those questions are unanswerable in the abstract and very answerable once you've actually run one modest pilot to completion, and the answers materially reduce the risk of the second, more ambitious attempt, because you're no longer guessing at how your own organization behaves under this kind of pressure.

A third, more practical signal: has the first pilot survived contact with a genuinely unexpected input or edge case, and did the organization's response to that moment build confidence rather than erode it? A tool that's only ever been tested against clean, well-behaved data hasn't really been tested yet. One that's hit a messy real-world edge case and been quietly corrected, without drama, has demonstrated something a clean pilot never gets the chance to demonstrate: that the organization can maintain and improve the thing it built, not just launch it once and hope.

The scheduling flagship that failed loudly, and the maintenance alert that succeeded quietly

Consider two hypothetical manufacturing organizations beginning their AI journeys from broadly similar positions.

The first decides to start with an ambitious cross-plant scheduling problem, exactly the kind of complex, high-value decision that naturally attracts leadership attention and appears well suited to AI. A cross-functional team is formed, senior sponsors speak enthusiastically about the initiative, and expectations rise quickly.

As the pilot progresses, however, an important weakness emerges. Experienced planners rely on several informal scheduling constraints, rules of thumb, exceptions, sequencing preferences, and practical trade-offs, that have never been formally documented. Because these constraints are absent from the available data and business rules, the AI system occasionally produces recommendations that look mathematically reasonable but are operationally impractical.

A few highly visible errors are enough to undermine confidence. The pilot is eventually discontinued, but its influence lasts much longer than the project itself. The organization is left with a simplified internal narrative: “AI was tried for scheduling, and it did not work.” That perception can make later AI initiatives harder to introduce, even when the subsequent use cases are very different.

Now consider a second hypothetical organization that takes a more modest first step.

Instead of beginning with a highly complex planning decision, it selects a narrow maintenance recommendation use case. The system helps identify which pending maintenance items may deserve greater attention because they show a higher likelihood of contributing to an unplanned stoppage within the following two weeks. Importantly, the recommendation is built on data the maintenance team already understands, trusts, and uses in its daily work.

There is no major launch event and little organisational fanfare. Over the course of a quarter, maintenance engineers simply begin observing that the recommendations are often useful. Trust develops gradually through repeated operational experience rather than through presentations or executive sponsorship.

That modest success creates something strategically valuable: credibility.

When the organization later proposes a more ambitious AI use case, the conversation begins from a very different position. Instead of asking whether AI can be trusted at all, people are more willing to explore where else it might help.

The contrast illustrates an important principle in industrial AI adoption: the best first AI use case is not necessarily the most impressive one. Many times, a narrow problem with trusted data, visible usefulness, and manageable decision complexity can create more long-term value than an ambitious flagship initiative attempted before the organization is ready.

Save the ambitious use case for your second pilot. Let the first one just work

Two Paths to Scale
Two Paths to Scale - AI Generated

The instinct to make a first AI pilot ambitious is completely understandable, and no leadership team that feels it deserves any less credit for it. It comes from a genuine desire to show real seriousness about the initiative. But trust, once damaged by a visible early failure, is one of the most expensive things to rebuild in any organization, and an ambitious first pilot simply carries more of the specific risks that damage trust than a narrow one does. Save the ambitious use case for your second pilot, once you have a quiet, genuine success behind you and a specific person willing to vouch for it in their own words, unprompted, to a colleague who wasn't in the room when it launched. Let the first one just work.

Disclaimer

Industry situations in this chapter are composite illustrations unless explicitly attributed to a public source. They are not claims about any particular company, plant, vendor, or incident. External standards, research, and public case studies should be verified before publication. Implementations must be validated against local safety, quality, cybersecurity, regulatory, contractual, labour, privacy, and data-governance requirements. AI recommendations and autonomous actions should remain within clearly defined human authority, operational controls, and tested recovery procedures.

#AIPilot #ManufacturingAI #SteelIndustry #DigitalTransformation #ChangeManagement #Industry40 #SupplyChainAI #OperationsExcellence #EnterpriseAI #ManufacturingLeadership

Further reading

Continue reading