Production Losses Aren't a Mystery - Here's the Fix

10 min read

Share this page

Choose where to share this page.

Production Losses Aren't a Mystery - Here's the Fix

Most "unexplained" production losses are just unstructured questions. Here's how a simple issue tree turns them into fixable, ranked causes.

dattarajsandur.com

Share via

Production Losses Aren't a Mystery — Here's the Fix
Production Losses Aren't a Mystery — Here's the Fix

Description

Most "unexplained" production losses are just unstructured questions. Here's how a simple issue tree turns them into fixable, ranked causes.

Accordion controls

"We don't really know why we lost the output." This sentence, or some close variation of it, comes up in production reviews on nearly every continent. It's one of the more common things said in that room, and typically, one of the least often actually true. It's usually delivered with a small shrug, almost an apology, as though the loss is simply a fact of nature rather than something that could, with the right questions, be traced back to its actual origin. It's usually said in good faith, by people who have genuinely looked at the numbers and genuinely don't see an obvious explanation staring back at them. But the loss usually isn't unexplainable at all. What's actually happened, almost every time the question gets dug into, is that it was never structured well enough to produce a real answer in the first place.

The Jam That Wasn't the Cause
The Jam That Wasn't the Cause - AI Generated

Why "we don't know why we lost the output" is rarely true

That shrug is best treated as a signal in its own right. Not evidence that the loss is genuinely mysterious, but evidence that nobody has yet spent the ninety minutes it usually takes to make it stop being mysterious.

When a team says a loss is unexplained, what's usually happened, on closer inspection, is that the loss got logged as a single, catch-all category, "unplanned downtime," "yield variance," "quality hold," and nobody ever broke it down any further than that first, broad label. The category itself is entirely real. Nobody's disputing that the downtime happened or that the yield came in low. The explanation inside that category, though, has simply never been separated out into its component parts, because doing so takes a bit of deliberate structure that a monthly report, by its nature, rarely provides on its own.

Ask "why did we lose the output" as one big, undifferentiated question, and you will reliably get one big, unsatisfying answer back: "unplanned downtime happened." That's true, and it tells you almost nothing you can act on. Ask the same underlying question instead as a structured sequence of progressively narrower questions, and the vague category starts splitting apart, almost on its own, into specific, addressable causes that were there all along, simply uncounted.

Building an issue tree instead of a blame list

The tool for this is refreshingly simple, and it doesn't require any special software or a data science team to run. It's an issue tree, built with nothing more than a whiteboard, a group of the right people in the room, and about ninety focused minutes. Start with the loss category exactly as it currently exists, and ask what could plausibly cause it, listing every branch that comes up without judging or filtering any of them yet. The judging comes later. Then, for each of those branches, ask the same question one level deeper: what could cause that. Within two or three levels of this, typically, most loss categories that had felt like an unsolvable, permanent mystery start turning into a short, specific list of testable causes. A particular changeover step that's always been a little inconsistent. A specific shift pattern that behaves differently from the others. A single piece of equipment that quietly behaves differently after it's run for longer than a certain number of hours.

This is meaningfully different from a blame list, which is the far more common default in a room under pressure, and which tends to stop at categories like "operator error" or "equipment issue" and go no further, because those categories feel satisfyingly complete even though, on reflection, they explain almost nothing useful about what to actually do next. An issue tree keeps splitting, deliberately and a little stubbornly, until the branches are specific enough that someone in the room could, in principle, go test one of them tomorrow morning.

Building the Tree
Building the Tree - AI Generated

Separating root cause from symptom

A common trap along the way, one teams fall into constantly, is mistaking a symptom for a root cause and stopping there, satisfied. "The line stopped because of a jam" is a symptom. True, observable, and almost entirely unhelpful on its own. Why the jam happened in the first place, whether it was a specific variation in an incoming raw material, a part that had quietly worn past its tolerance, or an operator working around a completely different upstream problem that nobody had flagged, is much closer to the actual root, and it's usually two or three "whys" further down the tree than where most teams naturally stop.

The issue tree forces this distinction structurally, almost against the natural inclination of the room, by simply refusing to accept the first plausible sounding answer as final. Every branch gets one more "why" applied to it, until the team genuinely runs out of further explanations. That point, where nobody can push the question any further, is usually exactly where the actual, fixable cause lives, waiting to be acted on rather than merely described.

Turning causes into an opportunity register

Once the tree is fully built out, the branches that recur across multiple separate loss incidents, or that trace back cleanly to a single upstream cause responsible for a disproportionate share of the total loss, become what is best thought of as the opportunity register. A short, explicitly ranked list of what's actually worth fixing first, ordered by how much loss each item is genuinely responsible for, based on the evidence gathered rather than on whoever argues most persuasively in the room.

This is a fundamentally different document from a line that simply reads "unplanned downtime: eight percent of the month," repeated in the same form every reporting period without ever getting smaller. The opportunity register tells you exactly where to spend the next improvement effort, with real specificity, rather than leaving the entire category sitting there as one large, undifferentiated block that everyone has quietly stopped expecting to actually shrink.

The changeover step nobody had reviewed in years

A plant taken on for this kind of work had logged "unplanned downtime" as a single catch-all category for what turned out to be well over a year, without any further breakdown. A number reviewed diligently every single week in the production meeting, discussed at some length, and never fully explained by anyone in the room, month after month. It had become, in a quiet way, simply accepted as an unavoidable cost of doing business, the kind of number everyone had stopped expecting to actually move.

A single issue-tree workshop, using roughly two months of existing incident logs that nobody had previously analyzed in this way, split that one broad category into eleven genuinely distinct branches within a single afternoon. One of those eleven branches, tracing back cleanly to a specific changeover procedure that, as it turned out, hadn't been formally reviewed or updated in several years, accounted for roughly sixty percent of the entire total loss on its own. A single, specific, fixable cause hiding in plain sight inside a number everyone had assumed was too diffuse to meaningfully attack.

The loss was never actually a mystery. It had simply never been asked about with enough structure to reveal, clearly, where the bulk of it actually lived. Once that structure was applied, for the cost of a single afternoon workshop, the path forward became almost obvious.

Running the workshop: who to invite and how to keep it honest

The mechanics of an issue-tree workshop matter more than they might seem to at first glance, because the same tool run badly produces a blame list dressed up as an issue tree, and the same tool run well produces genuine clarity in ninety minutes. A few things make the difference. Invite the people who are actually closest to the process, operators and shift leads, not just their managers, because the branches that matter most are usually things only the people doing the work every day would think to mention. Keep management in the room, but keep them listening rather than leading. The moment a plant manager starts suggesting branches, the room quietly starts building the tree the manager wants to see rather than the one that's actually true.

Start with real incident data if you have it, even two months' worth of imperfect logs, rather than working purely from memory. Memory is a surprisingly unreliable narrator for how often something actually happens versus how memorable the one dramatic instance of it was. And resist the urge to solve anything during the workshop itself. The issue tree's job is to produce the ranked list of causes. The fixing happens afterward, with the right smaller group, once everyone agrees on where the loss actually lives. Mixing diagnosis and solutioning in the same ninety minutes is one of the most common ways these workshops run long and produce a weaker result than they should.

When the loss splits four ways instead of one

Not every issue tree produces a single dominant branch as cleanly as the changeover example above, and it's worth being honest about that, because the method is still valuable even when the answer is messier. At another plant, a persistent quality-hold category, roughly five percent of monthly output, consistently, for over a year, turned out, once the tree was built, to have no single dominant cause at all. Instead, it split fairly evenly across four separate branches: a raw material variation from one particular supplier, a measurement point that two different shifts interpreted slightly differently, a tooling wear pattern that showed up mainly on older equipment, and a genuine training gap among a group of more recently hired operators.

That result was, in its own way, just as useful as finding one dominant cause, because it told the plant something equally important. This wasn't a single fixable problem being mislabeled as five percent of loss. It was actually five percent worth of four genuinely different, independent issues, each requiring its own separate fix, each worth roughly a quarter of the available improvement effort. Knowing that changed the entire improvement plan. Instead of one workshop chasing one root cause, the plant ran four much narrower, much faster fixes in parallel, each owned by a different person, and each measurably closed within the following quarter.

Structure the question before you accept it as unexplainable

An unexplained production loss is, typically, very rarely genuinely unexplainable in any deep sense. It's usually a question that's been asked at the wrong level of abstraction, one large, undifferentiated category presented month after month, instead of a structured tree of progressively narrower ones that actually gets built out and interrogated. Before accepting, in your own next production review, that a loss "just happens" or writing it off as background noise that everyone has learned to live with, spend a single workshop, ninety minutes, a whiteboard, the right five or six people in the room, actually building the issue tree behind it.

From One Category to Eleven Causes
From One Category to Eleven Causes - AI Generated

The mystery tends to dissolve considerably faster than most rooms expect going in, and what's left, once the tree is built, is a short, specific, and genuinely fixable list rather than a vague and permanent feeling category. If your organization has a "we've always had about this much unplanned downtime" number that nobody has questioned in years, that number, more than almost any other, is worth the ninety minutes. More often than not, it rarely survives the exercise in the same shape it went in, and even when it doesn't collapse into one dominant cause, you'll walk out with something considerably more useful than what you walked in with: a specific, ranked, and testable map of exactly where the loss lives, instead of a single tired category that everyone has quietly agreed to stop questioning.

Disclaimer

Industry situations in this chapter are composite illustrations unless explicitly attributed to a public source. They are not claims about any particular company, plant, vendor, or incident. External standards, research, and public case studies should be verified before publication. Implementations must be validated against local safety, quality, cybersecurity, regulatory, contractual, labour, privacy, and data-governance requirements. AI recommendations and autonomous actions should remain within clearly defined human authority, operational controls, and tested recovery procedures.


#RootCauseAnalysis #ManufacturingLeadership #ProductionLosses #SteelIndustry #OperationsExcellence #ContinuousImprovement #PlantManagement #ProcessImprovement #Industry40 #LeanManufacturing

Further reading

Continue reading