Part IV: AI-Native Projects · Chapter 13 · Human-in-the-Loop Is Not Enough

22 min read

Share this page

Choose where to share this page.

Part IV: AI-Native Projects · Chapter 13 · Human-in-the-Loop Is Not Enough

Discover how and why human presence is not sufficient for responsible AI unless people have time, authority, context, and confidence to intervene meaningfully.

dattarajsandur.com

Share via

Meaningful Human Control
AI Generated

Description

Discover how and why human presence is not sufficient for responsible AI unless people have time, authority, context, and confidence to intervene meaningfully.

There is a comforting phrase in the world of industrial AI: human-in-the-loop (in short HITL).

It appears in architecture diagrams, governance papers, procurement documents, and executive presentations. It suggests that people remain in control while artificial intelligence performs the repetitive work around them. The phrase reassures leaders that no machine will make an important decision without a person somewhere in the process.

But “somewhere in the process” is doing a great deal of work.

A quality engineer who receives an alert after the production slot has passed is technically in the loop but practically powerless. A planner who can click reject but must explain every disagreement, may be encouraged to accept by default. A shift in-charge who sees a recommendation without the plant context, may approve it simply because the screen presents it with confidence. A manager who is held accountable for an automated decision but cannot see its assumptions is not in control; that person is carrying responsibility without having the means to exercise it.

The important question is therefore not whether a human appears in the workflow. The important question is whether that human has the time, authority, context, confidence, and practical ability to change the outcome.

That is the difference between human presence and meaningful human control.

FAQ targets

By the end of this chapter, you must have answers to these questions:

- What does human-in-the-loop is not enough mean for manufacturing leaders?

- How can organisations apply this idea in manufacturing?

- What role do people and governance play?

- How can success be measured?


The approval button nobody trusts

Imagine an integrated steel plant preparing a set of coils for dispatch. The order is important: a customer has a planned shutdown, the material is linked to a project milestone, and missing the delivery window will create consequences far beyond a late truck. The planning system identifies a likely conflict. One coil is available, but its quality release is not complete. Another coil meets the quality requirements, but it is committed to a different customer. A third option requires a changeover that will reduce the night shift’s available production time.

An AI capability reviews the situation and recommends reallocating the second coil. The recommendation is displayed to the planner with a green status indicator and a sentence saying that the alternative “minimizes expected service risk.” There are two buttons: Approve and Reject.

At first glance, this appears to be a human-controlled process.

Now look more closely.

The planner has less than five minutes before the dispatch sequence is locked. The system does not show that the alternative customer has a contractual priority clause. It does not show that the coil has a pending surface-quality observation from the finishing line. The planner knows both facts because they exist in an email and in a shift log, but the recommendation engine does not. Rejecting the recommendation requires a written explanation, while approving it requires one click. The planner is also aware that previous overrides have been described as “manual exceptions” during performance reviews.

The human is present. The human is not meaningfully in control.

This pattern is common because organizations often design AI around the model’s output rather than around the person’s decision. The model is treated as the centre of the system, and the human is added near the end as a safety feature. In practice, a human who arrives too late, sees too little, or has too little authority is not a safety feature. That person is an approval ritual.


The central argument

Putting a human somewhere in an AI workflow does not automatically make the system responsible. Meaningful control is a designed capability, not a checkbox.

A responsible decision capability must help a named person make a better choice inside a real decision window. It must reveal the assumptions that matter, distinguish facts from estimates, show feasible alternatives, make uncertainty visible, and provide a legitimate route for disagreement. It must also preserve the organization’s ability to learn from what happened after the decision.

The distinctive lens of this chapter is simple:

Meaningful human control requires time, authority, context, confidence, and consequence.

These five conditions are connected. Time without authority merely gives someone longer to watch an outcome they cannot change. Authority without context turns intervention into guesswork. Context without confidence creates hesitation. Confidence without consequence produces unexamined automation. And consequence without a learning record turns accountability into blame.

The question to carry through this chapter is:

What must be designed around the decision so that people can act with confidence without losing accountability?


Presence is not control

It is useful to separate three ideas that are often mixed together.

Human-in-the-loop usually means that a person reviews, approves, or confirms something during the workflow. This can be appropriate for many decisions, but it says little about the quality of the review.

Human-on-the-loop usually means that a person supervises a system operating with some autonomy and intervenes when needed. This may be more efficient, but it creates a higher bar: the person must be able to notice the problem and intervene before the system creates an unacceptable consequence.

Meaningful human control asks a different question. It asks whether the person has a practical right and ability to influence the result, given the speed, complexity, authority structure, and consequences of the decision.

Consider a furnace control recommendation. If the system proposes reducing temperature and the operator has ten minutes to review it, the operator may be able to compare the recommendation with furnace condition, material chemistry, refractory status, and downstream requirements. If the same recommendation arrives when the furnace is already in a transient state and must be acted on within thirty seconds, a nominal approval process may be unsafe. The correct design may be a pre-approved bounded response with clear escalation rules, not a slow manual approval screen.

Meaningful control is therefore not always slower. Sometimes it means placing human judgment earlier, when choices are still open. Sometimes it means giving an operator the authority to stop an action immediately. Sometimes it means allowing the system to execute a low-risk action automatically while preserving a clear intervention path for a person.

The design must follow the decision’s timing and consequence, not a fashionable label.


The five conditions of meaningful control

The five conditions of meaningful control
The five conditions of meaningful control - AI Generated

1. Time to intervene

The first condition is a real decision window.

Many AI projects measure alert latency but do not measure intervention latency. An alert may be technically generated in seconds, yet reach the right person only after the practical opportunity has disappeared. A customer service team may receive a delay prediction after the truck has left. A quality team may see a defect probability after the batch has moved to the next process. A project manager may learn about a supplier risk after the installation crew is already mobilized.

The useful measure is not simply “How quickly did the model respond?” It is “How much meaningful time remained for the owner to choose among viable options?

This changes design priorities. A system should distinguish between an early weak signal and a late strong signal. It should show when the decision window will narrow. It should escalate not merely because confidence is high, but because the cost of waiting is increasing.

In a steel plant, the planner may not need a perfect prediction that a customer order will miss its promised date. The planner may need an imperfect signal three days earlier, while production sequence, allocation, transport, and customer communication can still be changed. A highly accurate signal received after the last feasible intervention is operationally less valuable.

2. Authority to intervene

The second condition is legitimate authority.

A person may have a technical button to reject a recommendation but lack organizational permission to use it. This happens when overrides are treated as failures, when escalation paths are unclear, or when responsibility is assigned to one role while decision rights remain with another.

Authority must be explicit. Who can reject the recommendation? Who can change the threshold? Who can stop execution? Who can release an exception? Who must be informed? What happens when the authorized person is unavailable during a night shift?

For a quality release recommendation, the quality authority may be non-negotiable. For a production sequencing suggestion, the planner may be able to change the order within defined constraints. For a line-stop recommendation, the operator may need immediate stop authority, even if the subsequent investigation belongs to maintenance or engineering.

The principle is not that every person should be able to change everything. The principle is that every consequential decision should have a clear, accessible, and respected route for intervention.

3. Context to understand the recommendation

The third condition is context.

A recommendation without context is an instruction wearing the costume of intelligence. People need to know what the system observed, what it assumed, what it did not know, which constraints were applied, and which alternatives were considered.

This does not require exposing every model parameter. It requires making the decision legible.

For example, a dispatch recommendation might say:

- The promise is at risk because the assigned coil is awaiting final inspection.

- The recommended substitute meets grade and dimension requirements.

- It is currently allocated to Customer B, whose contractual priority is lower for this date.

- Reallocation would create a new risk for Customer B in four days.

- The recommendation assumes transport capacity is available on the proposed route.

- The system has not incorporated an informal commitment recorded in a planner’s shift note.

This is far more useful than a confidence score of 87 percent. A confidence score may describe the model’s internal belief, but it does not tell the decision-maker whether the recommendation fits the actual operating situation.

Context also includes the human realities of work: which team is on shift, whether a crane is available, whether a maintenance permit is active, whether a customer has a special tolerance, whether a previous workaround has already consumed flexibility. Much of this context is not neatly stored in one database. A mature decision capability acknowledges that limitation instead of pretending that the system sees the whole plant.

4. Confidence to challenge or accept

The fourth condition is confidence… not confidence that the AI is always right, but confidence that the person can judge when it may be wrong.

People do not need an AI system to be infallible. They need it to be honest about uncertainty and predictable about its behaviour. If the system explains what it knows, what it is estimating, and when it is outside its experience, people can develop calibrated trust.

Uncalibrated trust has two forms. Automation bias occurs when people accept a recommendation because it came from the system. Algorithm aversion occurs when people reject useful recommendations because earlier failures damaged their confidence. Both are design problems.

Confidence grows when the system shows relevant evidence, presents alternatives, acknowledges missing context, and records outcomes. It also grows when people see that thoughtful overrides are respected rather than punished.

In an integrated steel plant, a shift in-charge may trust a casting sequence recommendation when it aligns with known mould constraints, tundish availability, chemistry requirements, and the condition of the downstream rolling schedule. The same person may distrust it when it ignores a recently changed maintenance restriction. The issue is not whether the algorithm is “good” in general. The issue is whether its reasoning is credible in this situation.

5. Consequence and learning

The fifth condition is consequence.

After a decision, the organization must know what was recommended, what the person chose, why the choice was made, what happened, and whether the result changed the future system.

Without this record, disagreement disappears into anecdote. The organization cannot tell whether people are overriding the system because the model is poor, because the data is incomplete, because local constraints are invisible, or because the operating policy has changed.

An override should not be treated automatically as a defect. It can be a valuable observation. A planner who rejects a recommendation may be exposing a hidden constraint. An operator who stops an automated action may be identifying a safety condition not represented in the system. A customer service representative who changes a suggested promise date may know that a customer’s production campaign has moved.

Learning does not mean turning every override into a new training label without review. It means creating a disciplined way to examine patterns and decide what should change: data, policy, interface, authority, model, or process.


Human stories hidden inside industrial decisions

The phrase “the user” is too abstract for responsible design. Real users have names, roles, pressures, habits, and consequences.

The planner is balancing customer promises, furnace campaigns, inventory, changeovers, transport, and the credibility of tomorrow’s plan. The scheduler is trying to create a sequence that is mathematically feasible and practically survivable. The operator is listening to equipment, watching process behaviour, and dealing with conditions that may not be visible in a dashboard. The quality engineer is protecting both the customer and the plant from a release that cannot be defended. The shift in-charge is making trade-offs with incomplete information while being expected to maintain safety and output.

The customer service representative experiences the decision from another direction. A system may suggest a reassuring message, but the representative knows that the customer will ask whether the revised date is genuinely achievable. Marketing may have promised responsiveness, flexibility, or premium service. Internal logistics may know that the proposed movement competes with another urgent load. Management may see the aggregate service metric but not the human effort required to preserve it.

If an AI capability ignores these different experiences, it may optimize a number while weakening the system around it.


From prediction to action

Prediction is often the easiest place to start. A model can identify a probable delay, quality concern, capacity conflict, demand change, or project risk. But a prediction creates value only when it changes what someone can do.

Progression from Prediction to Responsible Action
Progression from Prediction to Responsible Action - AI Generated

The progression from prediction to responsible action can be described as a ladder:

1. Prediction: What is likely to happen?

2. Explanation: What evidence and assumptions support that view?

3. Comparison: Which feasible alternatives are available?

4. Recommendation: Which option best balances the stated objectives?

5. Preparation: What work can the system prepare for a person?

6. Approval: Which person must authorize the action?

7. Bounded execution: What may the system execute within explicit limits?

8. Learning: What happened, and what should change?

Not every decision should climb to bounded execution. The right level depends on consequence, reversibility, confidence, speed, and feedback quality.

A low-risk inventory classification may be automated. A customer promise change may require review. A quality disposition may require a designated technical authority. An emergency equipment action may need immediate operator control followed by formal review. The objective is not to maximize autonomy. The objective is to match autonomy to the decision’s risk and reversibility.

A simple rule may be more valuable than a sophisticated model if people can understand it, use it, and improve it. A system that generates beautiful recommendations but cannot connect them to work execution is not decision intelligence. It is a commentary layer.


When AI recommendations fail

When a recommendation is wrong, teams often ask, “Why did the model fail?” That is sometimes the correct question, but it is not the first question.

The more useful diagnostic sequence is:

- Did the signal arrive too late?

- Was the decision window incorrectly understood?

- Was relevant context missing, stale, contradictory, or trapped in an informal channel?

- Did the system confuse a soft preference with a hard constraint?

- Were the alternatives technically feasible but impossible for the shift to execute?

- Did the recommendation optimize one metric while damaging a more important promise?

- Was the explanation too vague for the human to evaluate?

- Was authority unclear or culturally discouraged?

- Did the action fail to reach the correct execution system?

- Was the outcome recorded in a way that permits learning?

This sequence prevents an organization from blaming the model for a process-design failure. It also prevents the opposite mistake: blaming people for overriding a system that did not provide the information or authority they needed.


Architecture without abstraction

The architecture should follow the decision rather than the organizational chart.

Relevant events may come from ERP, MES, APS, quality systems, logistics platforms, sensors, documents, project systems, email, shift logs, and conversations. These sources should not be collapsed carelessly into a single “truth.” They should be connected with enough semantic and temporal context to support the decision being made.

A practical architecture has several layers:

Evidence and context. Bring together the facts, constraints, recent events, local notes, and source timestamps that affect the decision. Identify what is known, estimated, missing, or disputed.

Decision logic. Compare alternatives against explicit objectives and constraints. Preserve policy rules and hard limits separately from learned patterns so that a model cannot quietly override a safety or quality boundary.

Human decision workspace. Present the recommendation, reasons, alternatives, uncertainty, and consequences in a form suited to the role. A planner, operator, and executive should not see the same screen simply because they are looking at the same event.

Authority and workflow. Route the decision to the person with the right to act. Define approval, rejection, escalation, delegation, and emergency intervention. Make the path work during weekends, night shifts, and system outages.

Execution guardrails. If an agent can call a tool or change a system, constrain its identity, permissions, scope, transaction values, timing, and rollback path. The agent should be a bounded participant in the operating system, not an unaccountable substitute for one.

Outcome and learning. Record the recommendation, the decision, the reason, the actual action, the result, and any subsequent correction. This is the foundation for improving both the AI and the operating model.

The architecture should also make failure visible. If a recommendation cannot be generated because a critical source is stale, the system should say so. If the action cannot be executed because a permission is missing, the system should not present the process as complete. Honest incompleteness is safer than simulated completion.


Designing the intervention itself

Intervention is often treated as a negative event: the human interrupts automation. A better design treats intervention as a first-class operation.

An intervention should have a clear reason category, such as missing context, policy exception, operational infeasibility, safety concern, customer commitment, data error, or model disagreement. The person should be able to add a short explanation without writing an essay. The system should show what will happen after intervention and who will receive the next responsibility.

There should also be different kinds of intervention. A person may pause an action, modify a parameter, choose another alternative, request more evidence, transfer the decision, or stop the workflow entirely. Treating all these actions as “reject” loses useful meaning.

For example, a scheduler may not reject an AI-generated plan completely. The scheduler may preserve the recommended sequence but change the start time because a crane inspection is pending. An operator may not reject an automated response; the operator may place the equipment in a safe state and ask the system to re-evaluate. A customer service representative may accept the operational recommendation but change the communication because the customer needs a different explanation.

The interface should reflect these real choices.


A practical ninety-day starting sequence

Organizations do not need to solve meaningful human control for every decision before beginning. They can learn through one bounded use case.

Days 1–30: Observe the decision in its natural habitat

Start with a recurring decision that people already make under pressure. Interview the decision-makers where the work happens: control room, planning office, quality laboratory, dispatch bay, maintenance area, customer service desk, or project room.

Ask them to describe the last time the decision went well and the last time it went badly. Ask what they knew at the time, what they wished they had known, who could approve a change, what they were afraid to do, and when the opportunity to act disappeared.

Do not begin by asking whether they want AI. Begin by asking where the decision becomes difficult.

Document the decision window, options, constraints, authority, handoffs, informal workarounds, and consequences. Pay particular attention to information that exists only in conversations, shift books, spreadsheets, or personal memory.

Days 31–60: Design the human decision contract

Create a one-page decision contract. It should state:

- the decision owner;

- the decisions the capability supports;

- the evidence it will use;

- the assumptions it will expose;

- the alternatives it will compare;

- the actions a person may take;

- the actions the system may take;

- the escalation path;

- the safe failure mode;

- the outcome that will be measured.

Design the recommendation screen with the decision-maker, not only with an analytics team. Test whether the person can understand the recommendation during a realistic shift, not in a quiet demonstration.

Agree in advance that responsible overrides will be studied, not automatically penalized. Without this agreement, the pilot will produce compliance theatre instead of learning.

Days 61–90: Run in shadow mode

Let the capability make recommendations without changing production decisions. Compare its recommendations with human choices and actual outcomes. Ask where the system would have helped and where it would have created a new risk.

Use disagreements as structured evidence. If planners repeatedly ignore the same recommendation, investigate whether the model misses a constraint or whether the planners are using an undocumented policy. If operators cannot explain why the system is recommending an action, improve the context display before increasing automation.

At the end of the shadow period, select a narrow set of low-risk, reversible actions for a controlled trial. Define the stop conditions before enabling execution.

After day 90: Expand authority carefully

Increase authority only when the evidence supports it. The evidence should include not only model accuracy, but also intervention time, override quality, user comprehension, execution reliability, safety and quality outcomes, and the organization’s ability to explain what happened.

A system that performs well in a stable week may not be ready for an outage, campaign change, demand spike, or unusual material condition. Expansion should therefore be tied to operating conditions, not merely to a calendar milestone.


Questions for leaders

Before approving a human-supervised AI capability, leaders should ask:

- What recurring decision is this capability intended to improve?

- Who experiences the consequences most directly?

- How much time remains for intervention when the recommendation appears?

- Who has the authority to reject, modify, pause, or stop the action?

- What context will the person see, and what context will remain outside the system?

- Which assumptions must be visible?

- What is the safe failure mode when data, models, connectivity, or permissions fail?

- Will a thoughtful override be respected, investigated, and learned from?

- What will be recorded when a person disagrees?

- What evidence would justify giving the system more authority later?

- Which decisions should remain human-owned even if supporting work becomes automated?

These questions move the conversation from “Does the model work?” to “Can this organization operate responsibly with the capability?


Conclusion: design the human who remains in control

Human control is meaningful only when the human can still make a difference.

That sounds obvious, but it is easy to lose in an architecture review. The diagram shows a person. The workflow contains an approval step. The governance document says “human oversight.” Everyone feels reassured, while the real operator has thirty seconds, incomplete context, unclear authority, and a cultural penalty for disagreeing.

The answer is not to remove AI from the process. The answer is to design the relationship more honestly.

Human-AI Partnership
Human-AI Partnership - AI Generated

Give people time while choices are still open. Give them authority that matches the consequences they carry. Give them context rather than decorative explanations. Give them confidence through visible assumptions and calibrated uncertainty. Give the organization a learning record that treats intervention as intelligence rather than disobedience.

The best industrial AI will not make people feel less important. It will help them spend their judgment where it matters most. It will prepare evidence, reveal trade-offs, surface risks, and carry out bounded work without pretending that every decision can be reduced to a score.

The work begins with one recurring decision, one group of people, and one honest learning loop. That is how an organization moves from attractive language about AI to a practical operating capability—and from human presence to meaningful human control.


Disclaimer

Industry situations in this chapter are composite illustrations unless explicitly attributed to a public source. They are not claims about any particular company, plant, vendor, or incident. The ideas are intended for editorial and strategic discussion. Any implementation must be validated against local safety, quality, cybersecurity, regulatory, contractual, labour, data-governance, and operational requirements. Human oversight must not be used as a substitute for formal engineering controls, statutory responsibilities, professional certification, or established plant procedures.

#HumanInTheLoop #HumanOversight #ResponsibleAI #ManufacturingAI #HumanOnTheLoop #AIGovernance #DecisionIntelligence #IndustrialAutomation #SmartManufacturing #OperationalSafety

Continue reading