Human Decision Systems
Learning objectives
By the end of this chapter, you will be able to:
- distinguish a prediction from an operational decision;
- assign decision authority across people and automated components;
- design confidence bands, review queues, overrides, and escalation paths;
- record the evidence needed for audit, contestability, and learning; and
- evaluate the combined performance of the model, workflow, and human reviewers.
From predictions to decisions
A deployed model returns an estimate. A decision system uses that estimate, together with rules, context, capacity, and human judgement, to choose an action. These are not the same object.
For the guide’s decision-support service, let (p_i) be the model’s estimated probability that case (i) needs intervention. The service can expose (p_i), but the organization still has to define what happens next. A simple policy might be
[ a_i = \[\begin{cases} \text{routine}, & p_i < t_L \\ \text{human review}, & t_L \le p_i < t_H \\ \text{priority}, & p_i \ge t_H. \end{cases}\]]
The thresholds (t_L) and (t_H) encode consequences and operational constraints. They should not be chosen from model metrics alone. A false negative may delay help, a false positive may consume scarce capacity, and an unnecessarily large review band may overwhelm reviewers.
| Layer | Output | Example question |
|---|---|---|
| Model | score or class | How likely is the target outcome? |
| Policy | proposed action | Which score ranges require review? |
| Workflow | routed case | Who receives it, and by when? |
| Human | confirmed or changed action | Is relevant context missing from the model? |
| Organization | realized outcome | Was the action timely, fair, safe, and useful? |
This separation makes change safer. A model can be retrained without silently changing decision authority, and a policy threshold can be revised without pretending that the underlying model has changed.
Interfaces, explanations, and uncertainty
A decision interface should help a reviewer make a better decision, not merely display a model output. It should show:
- the proposed action and calibrated probability;
- the thresholds and policy version that produced the route;
- the most relevant case information and its source time;
- missing, stale, or out-of-range inputs;
- a concise explanation of influential factors and known limitations;
- the expected consequence of each available action; and
- clear controls to confirm, override, escalate, or defer.
Showing a probability to many decimal places creates false precision. Use an appropriate resolution, plain-language bands, and context such as “near the review threshold.” Explanations should describe what influenced the model; they do not establish causality or replace domain evidence.
Good interface behaviour also reduces automation bias. Reviewers should be able to form an initial assessment from case evidence, see why the case was routed, and record a reason when they disagree. Interface guidance for human-AI interaction emphasizes communicating what the system can do, supporting efficient correction, and making changes understandable (Amershi et al. 2019).
Overrides, escalation, and contestability
An override is not automatically a model failure. It may reveal information unavailable to the model, a legitimate exception, a poor interface, a policy mismatch, or reviewer inconsistency. Every override therefore needs structured context.
A useful decision event contains:
{
"case_id": "case-00427",
"model_version": "decision-model-1.4.0",
"policy_version": "routing-policy-2.1",
"score": 0.63,
"proposed_action": "human_review",
"final_action": "priority",
"actor_role": "trained_reviewer",
"reason_code": "verified_context_not_available_to_model",
"explanation": "Recent event increases urgency.",
"decided_at": "2026-08-05T09:14:00Z",
"review_due_at": "2026-08-05T10:14:00Z"
}Free text can supplement a reason code, but it should not be the only record. Structured codes make it possible to detect repeated policy exceptions and distinguish productive overrides from noise.
Contestability requires a path for an affected person or authorized representative to question a consequential decision. The process should identify the decision, preserve the evidence and versions used, assign an independent reviewer when appropriate, communicate the result, and correct downstream records. Logging without a review mechanism is auditability, not contestability.
Escalation triggers should include:
- a high-consequence action with low evidence quality;
- a score close to a threshold when small input changes alter the route;
- disagreement between model output, rules, and verified case context;
- repeated overrides for the same reason or subgroup;
- a missed review deadline; and
- a challenge or complaint about a previous decision.
Measuring system-level outcomes
Model discrimination remains useful, but it does not reveal whether the decision workflow functions well. Evaluate the complete chain:
| Dimension | Example measure | Why it matters |
|---|---|---|
| Predictive | recall, precision, calibration | Checks the quality of estimates |
| Decision | false-negative and false-positive action rates | Connects errors to actions |
| Workflow | review rate, queue time, SLA attainment | Reveals operational feasibility |
| Human | override and escalation rates, reviewer agreement | Detects ambiguity and policy gaps |
| Equity | outcome and error rates by relevant group | Finds uneven system effects |
| Impact | intervention benefit, harm, cost, delay | Tests whether the system improves outcomes |
Metrics need denominators and slices. “Twenty overrides” is uninterpretable without the number reviewed, the reason distribution, and the affected groups. Low override rates are not necessarily good: they may indicate appropriate automation, but they can also indicate rubber-stamping or a difficult interface.
Outcome labels may arrive later than decisions. Keep separate timestamps for prediction, review, action, and outcome so that monitoring windows do not compare incomplete cohorts. Where interventions affect the outcome, naive accuracy comparisons can also become misleading: the decision changes what later becomes observable.
Case-study implementation
The chapter program simulates 5,000 decision events. It generates a latent outcome, an imperfect model score, a two-threshold routing policy, human review in the uncertainty band, and specialist escalation for a subset of difficult cases.
Run it from the repository root:
Terminal
bash scripts/bash/12-simulate-human-decision-system.shThe program writes:
results/12-human-decision-summary.csv, containing system-level indicators;results/12-human-decision-events.csv, containing auditable case-level events; andresults/figures/12-human-decision-system.png, showing the routing funnel and the effect of review.
The simulation is a teaching instrument, not evidence that a particular threshold or review process is safe. Its value is that it makes the policy assumptions measurable. Try changing LOW_THRESHOLD, HIGH_THRESHOLD, the review-capacity limit, or the simulated reviewer reliability, then inspect the trade-offs among workload, false negatives, false positives, and escalation.
Before deployment, the case-study team should approve explicit acceptance criteria. Examples include a maximum false-negative action rate, a minimum review-SLA attainment rate, a bounded queue size, no unexplained subgroup disparity, and complete decision-event logging. The criteria should be evaluated together; optimizing only one can transfer harm elsewhere.
Chapter summary
A model score becomes valuable only through a governed decision process. That process must separate prediction from policy, specify who has authority, communicate uncertainty without false precision, make overrides and escalation usable, and support contestability. Evaluation must then move beyond model accuracy to workload, timeliness, human behaviour, equity, and realized outcomes.
The next chapter combines delivery, observability, monitoring, lifecycle controls, governance, and human decision design into the end-to-end systems case study.
Exercises
- Change the review band from 0.35–0.70 to 0.25–0.80. Which error rates improve, and what happens to review workload?
- Add a daily review-capacity constraint. Define how excess cases are prioritized and what safe fallback applies.
- Propose five override reason codes that are specific enough to support learning but broad enough to use consistently.
- Define a contestability workflow for one consequential decision in your domain, including evidence retention and response deadlines.
- Identify one outcome metric that the model can influence through intervention. Explain why a simple post-deployment accuracy comparison could be biased.