Chapter 02 : Beyond OTIF: Why One Supply Chain Metric Is Never Enough

16 min read

Share this page

Choose where to share this page.

Chapter 02 : Beyond OTIF: Why One Supply Chain Metric Is Never Enough

OTIF can look healthy while margin, inventory, quality, resilience, and cash exposure worsen. Learn how complementary supply-chain indices improve decisions.

dattarajsandur.com

Share via

Chapter 02 : Beyond OTIF: Why One Supply Chain Metric Is Never Enough
AI Generated

Description

OTIF can look healthy while margin, inventory, quality, resilience, and cash exposure worsen. Learn how complementary supply-chain indices improve decisions.

Accordion controls

At the end of the month, the operations review begins with good news.

On-time in-full delivery is at 97.8 percent, two points above target. The chart is green. The team has protected the customer promise, kept the largest accounts supplied, and avoided the kind of service miss that gets discussed in the boardroom.

Then finance asks why gross margin is down.

The answer is buried in the recovery effort. A customer order was rescued with premium freight. The plant ran overtime to recover a sequence change. Inventory was pulled from a lower-volume customer and will need to be replenished at a higher cost. A quality hold delayed the original shipment, and the replacement was rushed before the root cause was understood. Working capital has increased because several uncertain orders were built early “just in case.”

The OTIF result is real. So is the margin leakage.

This is the uncomfortable truth about supply-chain performance: a strong headline metric can coexist with serious exposure. The number may describe what customers received while saying very little about what the organization had to sacrifice to deliver it, what risk it carried forward, or whether the next promise is defensible.

OTIF is useful. It is not sufficient.

The decision-ready supply chain therefore needs a family of complementary lenses. Each lens makes a different kind of exposure visible. Together, they help leaders understand not only whether the network delivered, but how it delivered, what it consumed, what it put at risk, and what decision should come next.

The green headline metric
The green headline metric - AI Generated

When the headline metric looks healthy

Every organization needs a small set of measures that create focus. OTIF has earned its place because it is easy to understand and closely connected to the customer experience. Did the order arrive when promised, and did it contain the agreed quantity?

That clarity is valuable. It is also dangerous when the measure becomes a proxy for overall health.

Imagine a manufacturer with a fixed delivery commitment to a strategic customer. Three days before dispatch, a component shortage threatens the order. The planner has four choices: delay the shipment, partially ship it, substitute material, or recover the promise through premium freight and overtime.

The team chooses recovery. The plant resequences production. A carrier is upgraded. A supervisor approves overtime. Customer service keeps the original delivery date. The order arrives complete and on time.

At the next review, the OTIF dashboard records a success.

But several other facts are now true:

- The cost to serve the order has increased.

- The plant has created fatigue and schedule instability.

- Another customer’s allocation has been weakened.

- Inventory has been consumed from a less visible location.

- The shortage’s root cause may still be unresolved.

- Cash may be tied up in replacement or buffer stock.

- The organization may have learned to rescue promises rather than improve the system.

None of these facts invalidates the OTIF result. They explain its price.

The leadership challenge is not to punish the team for protecting the customer. It is to see the entire decision. A delivery metric should trigger a better question: “What did it take to achieve this service result, and what future exposure did the recovery create?”

The limits of OTIF and local KPIs

Metrics become misleading when they are asked to answer questions they were not designed to answer.

OTIF is a service outcome. It does not directly measure resilience, inventory usefulness, quality stability, demand confidence, cost-to-serve, or cash conversion. A high OTIF score can mean that the network is healthy. It can also mean that people are working very hard to hide instability from the customer.

The same issue appears across functions.

Procurement may report purchase-price savings while supplier lead-time variability is increasing. Manufacturing may report schedule adherence while the schedule is being protected through excessive finished-goods inventory. Logistics may report freight utilization while urgent orders are waiting for the next available departure. Finance may celebrate inventory reduction while service risk is moving into lost sales and expedites.

These measures are not wrong. They are partial.

The limits of local KPIs
The limits of local KPIs - AI Generated

There are four common traps.

The outcome trap

An outcome metric tells us what happened, but not necessarily whether the process is becoming more reliable. A good month can be produced by favorable demand, heroic intervention, or temporary inventory. Without leading indicators, the organization cannot tell which explanation is true.

The boundary trap

Most metrics have an owner and a boundary. OTIF may be calculated by order line, shipment, customer, plant, or region. A metric can improve when the calculation boundary changes, even if the customer experience does not.

The incentive trap

When a metric is tied strongly to targets or bonuses, teams rationally optimize it. They may prioritize easy orders, ship partial quantities under favorable definitions, build inventory early, or use expensive recovery actions. The metric improves while the system becomes less balanced.

The averaging trap

An aggregate number can hide concentration. An OTIF score of 98 percent does not tell a leader whether the two percent of misses affected one strategic launch, ten low-value orders, or a regulated customer. Average performance is not the same as acceptable exposure.

The answer is not to abandon OTIF. It is to place it in a wider decision framework.

The ten lenses in the framework

The ten complementary lenses
The ten complementary lenses - AI Generated

The indices in this series are complementary lenses, not competing scoreboards. Each one is designed to surface a different leadership question.

1. OTIF Network Reliability Index

The On-Time and In-Full Network Reliability Index, or ONRI, asks whether the organization is delivering the customer promise as agreed.

ONRI remains essential because service is where the supply chain becomes visible to the customer. But reliability should be interpreted with context. A shipment that arrived on time through repeated expedites is not equivalent to one that flowed through a stable process.

Useful supporting signals include promise-date changes, partial shipments, premium freight, manual overrides, and the share of orders fulfilled from planned rather than emergency inventory.

2. Margin Integrity Index

The Margin Integrity Index, or MII, asks whether the service outcome is economically sound.

MII can include freight premiums, overtime, rework, substitutions, discounts, inventory write-down exposure, and other costs required to protect the order. It does not argue that every expedite is bad. It makes the trade-off visible.

A high ONRI with falling MII is a warning that the network is buying service with margin.

3. Working Capital Velocity Index

The Working Capital Velocity Index, or WCVI, examines how efficiently cash moves through inventory, receivables, and payables without weakening service.

Inventory is not automatically healthy or unhealthy. The useful question is whether it is positioned, current, and usable for the demand that matters. A finished-goods mountain may support OTIF while concealing obsolete, misallocated, or slow-moving stock.

WCVI adds the cash and time dimension to service discussions.

4. Supply Resilience Index

The Supply Resilience Index asks how well the network can absorb disruption and recover.

It may consider supplier concentration, alternate sources, recovery time, capacity flexibility, logistics options, critical-material exposure, and the quality of contingency plans. A network can have excellent current service and weak resilience if it depends on one supplier, one port, or one constrained production asset.

5. Quality Stability Index

The Quality Stability Index connects quality performance to flow and customer risk.

A quality hold can affect delivery, inventory, production sequence, customer confidence, and regulatory exposure. Quality should not be treated as a separate departmental topic. It is part of the decision about whether material is genuinely available and whether a promise is safe to make.

6. Demand Confidence Index

The Demand Confidence Index asks how much trust the organization should place in the demand signal behind a plan.

Forecast accuracy alone is not enough. Leaders need to understand forecast bias, volatility, forecast value-add, customer commitment, order-pattern changes, and the difference between statistical demand and commercial intent.

Low demand confidence should influence inventory, capacity, procurement, and promise decisions.

7. Flow Efficiency Index

The Flow Efficiency Index looks at how smoothly material moves through the network.

It can include queue time, changeover disruption, rework, handoff delay, waiting for quality release, transport dwell, and the ratio of value-adding time to total elapsed time. A network that meets delivery through buffers and intervention may show weaker flow efficiency even when OTIF remains strong.

8. Cost-to-Serve Exposure Index

The Cost-to-Serve Exposure Index identifies where service commitments are becoming disproportionately expensive.

The issue may be a customer segment, product family, geography, order profile, packaging requirement, or service promise. Cost-to-serve exposure is not a reason to reduce service blindly. It is a reason to understand which commitments need redesign, repricing, segmentation, or a different operating model.

9. Decision Latency Index

The Decision Latency Index measures how long it takes to recognize a meaningful change, interpret it, choose a response, and execute that response.

This index is especially useful because delay is often created between systems and functions. A supplier may report a risk at 09:00, but the planner may not see it until 11:00, the commercial team may not be consulted until 14:00, and the transport decision may miss the cut-off at 16:00.

The organization may have all the necessary data and still lose the available options through slow coordination.

10. Recovery Readiness Index

The Recovery Readiness Index asks whether the organization is prepared to respond to a foreseeable disruption.

Readiness includes named owners, approved alternatives, usable inventory, tested playbooks, decision rights, communication paths, and the ability to monitor whether the intervention worked. A risk register without a feasible recovery path is an inventory of concerns, not resilience.

How indices interact and conflict

The value of an index family appears when the indices are read together.

Interacting indices
Interacting indices - AI Generated

Consider four possible situations:

High ONRI, low MII

The customer promise is being protected, but the cost of service is rising. The likely causes include premium freight, overtime, rework, low-volume production runs, and commercial concessions. The right response may be to segment promises, redesign recovery rules, or address the root constraint.

High ONRI, low WCVI

Service is strong, but cash is becoming trapped in inventory or receivables. The organization should examine whether it is building too early, holding the wrong stock, or using inventory to compensate for unreliable planning.

High ONRI, low resilience

Current execution looks reliable, but the network is exposed to a shock. This is common when a business has stable supply from a concentrated source. Leadership should test alternatives before the disruption forces a decision.

Low ONRI, high recovery readiness

The current service result is weak, but the organization may have the capability to improve quickly. It has owners, alternatives, and playbooks; execution or data quality may be the limiting issue. This is different from a network that has poor service and no credible recovery path.

No combination automatically determines the answer. The indices create a structured conversation about what is happening and what trade-off is acceptable.

Avoiding local optimization

Local optimization begins with a reasonable goal. A plant wants to reduce changeovers. Procurement wants to consolidate suppliers. Logistics wants to fill vehicles. Finance wants to reduce inventory. Customer service wants to protect the promise.

The problem begins when each goal is optimized without an end-to-end consequence check.

For example, manufacturing may increase batch sizes to improve equipment utilization. The decision lowers local cost but creates more finished-goods inventory, reduces product mix flexibility, and increases the chance that demand changes before the material is used. OTIF may remain high for a time because the inventory is available. WCVI, demand confidence, and resilience may deteriorate.

The solution is not to maximize every index simultaneously. That would be impossible. The solution is to make the conflict explicit and govern it.

An executive operating review should ask:

- Which index is improving?

- Which indices are deteriorating at the same time?

- Is the movement a deliberate trade-off or an unintended consequence?

- Who benefits from the decision, and who carries the risk?

- How reversible is the decision?

- What signal would tell us to change course?

This is where composite indices and decision narratives are more useful than a wall of dashboards. They preserve the relationships between numbers.

A practical operating review

The executive operating review
The executive operating review - AI Generated

A balanced review can begin with the customer outcome and then widen the lens.

Start with ONRI by segment, product family, customer priority, and promise type. Do not begin with the aggregate alone. Identify where service is fragile, not only where it is missed.

Then review MII and WCVI for the same population. Ask whether service protection is consuming disproportionate margin or cash. Look specifically for premium freight, overtime, late changes, emergency buys, and inventory that is physically present but not practically usable.

Next, examine resilience, quality, demand confidence, and flow. These are the conditions that explain whether the current result is repeatable. A high service result built on unstable quality release or uncertain demand should not be treated as a stable baseline.

Finally, review decision latency and recovery readiness. Where did the organization lose time? Which action could have been taken earlier? Which playbook, permission, or data connection would have preserved optionality?

A strong review ends with decisions, not observations. Each material exposure should have an owner, an action, a timing expectation, and a measure of whether the action improved the situation.

Questions for an executive operating review

Leaders can use the following questions to move beyond metric inspection.

  • “Where is OTIF being protected through expensive or unsustainable intervention?”
  • “Which customers or products have good service results but poor margin integrity?”
  • “Where is inventory supporting service, and where is it merely hiding uncertainty?”
  • “Which current success depends on a supplier, asset, route, or individual that has no credible alternative?”
  • “Which quality conditions make nominal inventory unavailable for customer use?”
  • “Where has demand confidence changed, and what decisions have not yet caught up?”
  • “Which bottlenecks create the most decision latency between signal and action?”
  • “If the largest current constraint failed tomorrow, could we name the owner, options, authority, and recovery path?”
  • “Which metric conflict is deliberate, and which one is accidental?”
  • “What should we stop doing because it improves a local KPI while weakening the end-to-end outcome?”

These questions do not require a perfect data platform to begin. They require a willingness to look at the cost and consequence behind the headline number.

Start with one decision, not ten dashboards

One decision, many consequences
One decision, many consequences - AI Generated

An organization does not need to implement all ten indices at once. Start with a recurring decision where the tension is already visible.

Customer promise recovery is a useful candidate. Define the decision: when an order becomes fragile, should the team reallocate, expedite, resequence, substitute, partially ship, or reset the promise?

Then connect the minimum evidence: ONRI, MII, WCVI, quality status, available alternatives, and decision latency. Record the chosen action and the outcome. After several cycles, add resilience, demand confidence, flow, or recovery readiness where they improve the decision.

This approach keeps the framework practical. It also creates a learning loop. The organization can see whether a high-risk order was correctly identified, whether the recommendation was feasible, whether the decision was made in time, and what the intervention cost.

The first version can be deliberately simple. A weekly review may use a shared worksheet with one row per decision episode and a small set of evidence fields: customer promise, available quantity, quality status, planned versus actual cost, inventory consequence, owner, action, and outcome. That worksheet will often reveal more than a polished dashboard because it forces the team to record the choice and its consequence together. Once the pattern is understood, the organization can automate the evidence collection and introduce more sophisticated scoring. Technology should follow the decision design, not substitute for it.

This also changes the tone of performance management. Instead of asking which function caused the miss, the review can ask which relationship between metrics was misunderstood. Was inventory considered available before quality release? Was the customer promise protected without a margin threshold? Did a procurement saving increase resilience exposure? These questions move the organization from blame toward better operating rules while preserving accountability for the action taken.

Over time, the record becomes a practical memory of the network: which interventions worked, which promises were fragile, and which metrics reliably warned of trouble before the customer saw it.

That memory is what turns a collection of metrics into an improving decision system.

It makes tomorrow’s review more informed than today’s.

That is how measurement becomes an operating capability rather than a reporting ritual.

The aim is not to create a more impressive dashboard. It is to create a better operating conversation.

No single number can carry the supply chain

OTIF deserves to remain visible. Customers care about reliable delivery, and organizations need a clear service standard. But a supply chain is not a single promise. It is a system of promises, constraints, trade-offs, and consequences.

One number cannot represent all of that.

The decision-ready supply chain uses metrics as lenses rather than verdicts. ONRI shows the customer outcome. MII shows the economic integrity of the outcome. WCVI shows the cash and inventory consequence. Resilience, quality, demand, flow, cost-to-serve, decision latency, and recovery readiness explain whether the result is sustainable and what may happen next.

The most important insight is NOT that one index is better than another. It is that their relationships contain the meaning.

A green OTIF result may be a sign of health, a sign of heroic recovery, or a sign that risk is accumulating somewhere less visible. Leaders should be able to tell the difference.

Decision quality improves when the organization can see what each number says, what it leaves out, which numbers are moving together, and which trade-offs need a conscious choice.

No single number can represent a supply chain. Better decisions come from understanding the relationships between numbers, and acting while those relationships can still be changed.

Disclaimer

Industry situations in this chapter are composite illustrations unless explicitly attributed to a public source. They are not claims about any particular company, plant, vendor, or incident. External standards, research, and public case studies should be verified before publication. Implementations must be validated against local safety, quality, cybersecurity, regulatory, contractual, labour, privacy, and data-governance requirements. AI recommendations and autonomous actions should remain within clearly defined human authority, operational controls, and tested recovery procedures.


#SupplyChain #OTIF #SupplyChainMetrics #SupplyChainManagement #SupplyChainResilience #SupplyChainAnalytics #WorkingCapital #OperationsManagement #Logistics #DecisionIntelligence

Takeaways

Table with 11 rows and 2 columns.

Excerpt

Practical point / context

“The OTIF result is real. So is the margin leakage.”

A service success can still carry a significant economic cost.

“OTIF is useful. It is not sufficient.”

OTIF should remain important without becoming a proxy for total supply-chain health.

“These measures are not wrong. They are partial.”

Local KPIs provide useful evidence but cannot explain the entire end-to-end outcome.

“The value of an index family appears when the indices are read together.”

The relationships between metrics reveal trade-offs and hidden exposure.

“A high ONRI with falling MII is a warning that the network is buying service with margin.”

Connect customer service performance with cost-to-serve and profitability.

“No combination automatically determines the answer.”

Indices should support judgment, not replace it with a mechanical score.

“The solution is not to maximize every index simultaneously.”

Balanced performance requires deliberate trade-offs rather than impossible optimization.

“An alert without a playbook is just a new source of anxiety.”

Metrics become useful when they lead to clear ownership and action.

“Start with one decision, not ten dashboards.”

Begin with a recurring decision and build the minimum evidence needed to improve it.

“No single number can represent a supply chain.”

The article’s central takeaway: decision quality depends on understanding metric relationships.

Further reading

Continue reading