A model becomes valuable only when the surrounding system can turn its predictions into reliable, reviewable, and improvable decisions. This chapter completes the continuous decision-service case study by operating the service as a whole system. The model remains important, but it is now one component among deployment controls, service-level objectives, monitoring, feedback, retraining rules, governance checks, and human oversight.
The case study uses a binary risk score produced by decision-api. A higher score indicates a case that may require intervention. The system must decide whether to automate a low-risk outcome, route a case to a human reviewer, or trigger an urgent intervention. The objective is not to maximize the number of automated decisions. It is to make appropriate decisions while preserving reliability, traceability, and a safe path for uncertain or consequential cases.
By the end of the chapter, you will be able to:
connect the operational and model layers of a production system;
distinguish request-level signals from decision-quality signals;
implement explicit automation and human-review policies;
combine service, drift, outcome, and governance evidence in one release decision;
generate an auditable system-readiness report; and
identify the next action when the system is healthy, degraded, or unsafe.
The complete decision pathway
The deployed service is only one stage in the pathway:
Code
flowchart TD A[Incoming case] --> B[Validate input] B --> C[Score with model] C --> D{Decision policy} D -->|Low risk| E[Automated outcome] D -->|Uncertain| F[Human review] D -->|High risk| G[Urgent intervention] E --> H[Record outcome] F --> H G --> H H --> I[Monitor and learn]
flowchart TD
A[Incoming case] --> B[Validate input]
B --> C[Score with model]
C --> D{Decision policy}
D -->|Low risk| E[Automated outcome]
D -->|Uncertain| F[Human review]
D -->|High risk| G[Urgent intervention]
E --> H[Record outcome]
F --> H
G --> H
H --> I[Monitor and learn]
Each transition introduces a possible failure. Valid data may arrive late. The model service may be available while a dependency is unavailable. A statistically stable score distribution may still hide deteriorating outcomes. A technically strong replacement model may fail a governance check. Human review may become a bottleneck. End-to-end operation therefore requires evidence from several layers.
Layer
Primary question
Example evidence
Delivery
Can the candidate be released safely?
tests, immutable artifact, staged rollout
Service
Is the API dependable?
availability, latency, error rate
Data and model
Is scoring behaviour still credible?
missingness, drift, score distribution
Decision
Are actions producing acceptable outcomes?
precision, recall, false-negative rate
Human
Are uncertain cases reviewed in time?
review rate, queue size, turnaround time
Governance
Can the system be explained and audited?
lineage, approval, policy and incident records
No single metric proves that the system is healthy. The operational decision must combine the evidence and retain the reason for the result.
Define the operating policy first
The case study uses two thresholds:
scores below 0.35 receive the routine automated outcome;
scores from 0.35 through 0.74 enter human review; and
scores at or above 0.75 trigger urgent intervention.
These thresholds are policy parameters, not properties of the model. They express the cost of errors, available review capacity, and the organisation’s tolerance for automation. They should be versioned separately from model parameters so that a policy change does not masquerade as a model change.
The policy also defines system gates:
Gate
Target
Purpose
Availability
at least 99.0%
prevent decisions from depending on an unreliable service
p95 latency
at most 250 ms
keep the synchronous pathway responsive
Error rate
at most 1.0%
detect failed requests
Population stability index
below 0.20
identify material score drift
False-negative rate
at most 15.0%
constrain harmful missed cases
Human-review rate
at most 45.0%
keep the review queue operationally feasible
Governance record
complete
require model, data, policy, and approval identifiers
A gate is not automatically a universal best practice. It is a declared operating constraint for this example. Production targets should be negotiated from user needs, harm severity, workload, and recovery capability.
Run the capstone workflow
The executable case study simulates a reference window and a current production window. It then evaluates requests, decisions, outcomes, review demand, distribution drift, and governance completeness.
Run it from the repository root:
bash scripts/bash/13-run-end-to-end-system.sh
The wrapper runs scripts/python/13-run-end-to-end-system.py and writes:
results/13-system-readiness-summary.csv, containing one row per system gate;
results/13-system-readiness-report.json, containing policy, metric, gate, and decision evidence; and
results/figures/13-end-to-end-system-readiness.png, summarising the decision pathways and readiness gates.
The simulation uses a fixed random seed, so the teaching result is reproducible. A real implementation would replace simulated arrays with telemetry, prediction logs, delayed outcomes, review-queue events, and registry metadata.
Interpret the readiness evidence
The output distinguishes three concepts that are often incorrectly collapsed:
Metric: a measured property such as p95 latency.
Gate: a test that compares a metric with an operating target.
Decision: the system action produced from all required gates.
The overall decision is GO only when every required gate passes. A failed gate produces HOLD, even when most metrics look strong. This conservative rule is suitable for a release or continued-operation checkpoint in which the gates represent minimum safety conditions.
The plot has two panels. The first shows how current cases were distributed across routine automation, human review, and urgent intervention. The second shows the normalized margin for each gate. Positive margins pass; negative margins fail. Normalization allows unlike units—milliseconds, percentages, drift indices, and completeness—to appear in one readiness view without pretending that they are the same measurement.
Figure 14.1: End-to-end system readiness: decision pathways and gate margins.
The readiness report is more important than the visual alone. It records:
the model version and policy version;
the data window and simulation seed;
the thresholds used for action routing;
the measured metric values;
the direction and target of every gate;
the pass or fail result; and
the overall operational decision.
This structure makes the decision reproducible. An operator can tell not only that a release was held, but which requirement caused the hold and what evidence was used.
Respond according to the failed layer
An end-to-end checkpoint should lead to a specific response rather than a generic alert.
Failed evidence
Immediate response
Longer-term investigation
Availability, latency, or errors
stop rollout or shift traffic to a stable version
capacity, dependency, deployment, and timeout analysis
Score drift
intensify monitoring and validate input changes
data pipeline, population, and feature-contract analysis
Outcome quality
restrict automation or roll back the model/policy
labels, subgroup errors, calibration, and retraining
Review capacity
narrow the review band or add reviewer capacity safely
workflow design, prioritisation, and decision thresholds
Governance completeness
block release
ownership, lineage, approval, and documentation repair
Retraining is only one possible response. It does not repair API failures, missing lineage, a broken upstream contract, or an overloaded review team. The correct response follows from the layer that failed.
Apply the workflow to a candidate release
Suppose a newly trained model performs better offline. The release pathway should still proceed through controlled stages:
Register the candidate with its training data, code revision, metrics, and intended use.
Run unit, contract, integration, and model-quality tests in continuous integration.
Build one immutable artifact and promote the same artifact between environments.
Deploy to staging and verify service behaviour with production-like contracts.
Use a canary or shadow deployment to compare operational and prediction behaviour.
Confirm that the policy and governance gates remain valid for the candidate.
Expand traffic only while the required gates pass.
Record the final approval, policy version, artifact identity, and rollback target.
This approach separates model selection from production promotion. Better offline performance creates a candidate; it does not grant automatic production authority.
Close the feedback loop safely
After deployment, outcomes arrive later than predictions. The monitoring design must preserve the identifiers needed to join a prediction with its eventual outcome without placing sensitive raw features in general-purpose logs. A minimal record might include a pseudonymous case identifier, event time, model version, policy version, score, action, reviewer status, and outcome availability.
Once outcomes mature, the system can evaluate decision quality and compare automated and reviewed pathways. A retraining trigger should produce a candidate workflow:
Code
flowchart TD A[Monitoring signal] --> B[Investigate cause] B --> C{Model change needed?} C -->|No| D[Repair system or policy] C -->|Yes| E[Train candidate] E --> F[Validate and approve] F --> G[Controlled promotion]
flowchart TD
A[Monitoring signal] --> B[Investigate cause]
B --> C{Model change needed?}
C -->|No| D[Repair system or policy]
C -->|Yes| E[Train candidate]
E --> F[Validate and approve]
F --> G[Controlled promotion]
The feedback loop remains governed. Production outcomes become training evidence only after their definitions, quality, timing, and permissible use have been checked. Reviewer decisions may contain inconsistency or historical bias; they should not be treated as unquestionable ground truth.
Conduct the operational review
A recurring operational review can use the generated report as a compact agenda:
Reliability: Did the service meet its availability, latency, and error objectives?
Behaviour: Did input and score distributions remain within expected bounds?
Outcomes: Did decision quality remain acceptable after labels matured?
Workload: Could reviewers handle the cases routed to them?
Change: Which model, data, policy, or infrastructure versions changed?
Risk: Were incidents, overrides, or complaints concentrated in a subgroup or pathway?
Action: Should the team continue, constrain automation, investigate, retrain, roll back, or retire the system?
The review should name an owner and a deadline for every action. Dashboards create visibility; ownership creates response.
System completion checklist
The case study is operationally complete when the team can answer yes to all of the following:
Is one immutable service artifact promoted through the delivery pipeline?
Are service-level objectives connected to alerts and runbooks?
Are model inputs, scores, decisions, and outcomes monitorable with appropriate privacy controls?
Are automation, review, and intervention thresholds explicit and versioned?
Can every production decision be linked to a model and policy version?
Does a failed gate lead to a defined response and owner?
Does retraining create a candidate rather than silently replacing production?
Can the system roll back the model, policy, or deployment independently?
Are human reviewers supported by queue, escalation, and override mechanisms?
Is retirement possible when value declines or risk becomes unacceptable?
Conclusion
The capstone changes the unit of success. The question is no longer whether a model predicts well or whether an endpoint returns a valid response. The question is whether the complete socio-technical system can deliver dependable, monitored, governable, and appropriately supervised decisions over time.
The executable workflow demonstrates the core operating pattern: define policy, collect evidence, evaluate explicit gates, record the decision, and respond at the layer that failed. That pattern connects DevOps, MLOps, governance, and human decision design into one production discipline.
The model is deployed once. The system must keep earning the right to operate.