Chapter Beyond Model Deployment established that a deployed model service is only one component of a decision system. This chapter makes that distinction concrete by introducing the case study used throughout the rest of the guide.
The starting point is the validated breast-cancer prediction API built in the Model Deployment guide. It loads a fitted preprocessing-and-classification pipeline, validates eight numeric measurements, and returns the predicted class together with the estimated probability of malignancy. Here, that service becomes one component in a larger diagnostic-support workflow.
The objective is not to design a deployable medical device. It is to create a realistic educational architecture in which operational reliability, model monitoring, feedback, governance, and human judgment can be studied together.
Educational use only
The Wisconsin Diagnostic Breast Cancer dataset is a small teaching dataset with 569 observations and 30 numeric features (scikit-learn developers 2026). The case study is not clinically validated and must not be used to diagnose patients, choose treatment, or replace professional judgment.
Learning objectives
By the end of this chapter, you should be able to:
define the purpose, users, inputs, outputs, and boundaries of a model-enabled system;
distinguish the online prediction pathway from asynchronous operational and learning pathways;
assign responsibilities to system components rather than treating the model API as the whole application;
identify failure modes at interfaces between components;
translate a decision workflow into measurable service and data requirements;
explain why prediction records require model, request, and outcome identifiers; and
use an architecture diagram as a living operational contract.
The continuous case study
The fictional organization in this guide operates several diagnostic clinics. A laboratory system records measurements derived from a digitized image of a breast-mass sample. A clinician reviews the measurements, the model output, and the wider clinical evidence before determining the next appropriate step.
We will call the complete application the Diagnostic Support System, or DSS.
Intended purpose
The DSS prioritizes cases for clinician review. It does not make a diagnosis. A high estimated probability of malignancy should make a case visible to the reviewing clinician sooner; it must not automatically label a patient, cancel another assessment, or determine treatment.
Primary users
User
Need
Authority
Laboratory professional
submit a complete, valid measurement record
may correct or resubmit source measurements
Reviewing clinician
see the risk estimate with context and limitations
accepts, rejects, or overrides the suggested priority
Operations team
keep the technical pathway available and recoverable
deploys, rolls back, scales, and responds to incidents
Model owner
assess model behavior and approve model changes
proposes or rejects model promotion
Governance owner
verify approved use, access, auditability, and review
can pause the system or restrict its use
The decision and the outcome
The immediate decision is review priority, not diagnosis. The system returns a risk estimate and a suggested queue category. A clinician makes the final prioritization decision.
The eventual outcome is the confirmed diagnosis recorded after the appropriate clinical process. That outcome may arrive days or weeks after the prediction. This delay is crucial: operational monitoring can detect a failing API immediately, while model-performance monitoring cannot calculate real-world false negatives until verified outcomes are available.
Define the system boundary
A useful system boundary states what the team operates and what remains an external dependency.
For this guide, the DSS begins when the laboratory system submits a measurement event and ends when the confirmed outcome is linked to the original prediction. The laboratory instrument, the procedures used to obtain the sample, and the downstream diagnostic process remain outside the technical boundary, but their assumptions and failures still affect the system.
Inside the managed boundary
External dependency or context
API gateway and authentication
laboratory instrument and measurement procedure
input validation and feature adapter
identity and access provider
prediction service and model artifact
clinic network connectivity
decision rules and review queue
wider clinical evidence and clinician judgment
prediction and audit records
diagnostic confirmation process
metrics, logs, traces, and alerts
organizational policies and regulatory obligations
outcome-linking and monitoring jobs
source-system ownership and data correction
Being outside the boundary does not make a dependency unimportant. It means the team must define an interface, an owner, an observable expectation, and a response when the expectation is violated.
The reference architecture
The architecture contains three connected pathways:
the online pathway validates a request and returns decision support;
the operational pathway records telemetry and detects technical failure; and
the learning pathway links delayed outcomes to earlier predictions and evaluates continuing model fitness.
Code
flowchart TD A["Laboratory system"] --> B["Gateway and validation"] B --> C["Prediction service"] C --> D["Decision rules"] D --> E["Clinician review queue"] E --> F["Confirmed outcome"] B -. "telemetry" .-> G["Observability platform"] C -. "prediction record" .-> H["Audit and prediction store"] F -. "delayed label" .-> H H --> I["Monitoring and evaluation"] I -. "approved change" .-> C
flowchart TD
A["Laboratory system"] --> B["Gateway and validation"]
B --> C["Prediction service"]
C --> D["Decision rules"]
D --> E["Clinician review queue"]
E --> F["Confirmed outcome"]
B -. "telemetry" .-> G["Observability platform"]
C -. "prediction record" .-> H["Audit and prediction store"]
F -. "delayed label" .-> H
H --> I["Monitoring and evaluation"]
I -. "approved change" .-> C
The dashed connections are not secondary in importance. They are separated because they should not unnecessarily delay the clinician-facing response.
Online prediction pathway
The online pathway handles the latency-sensitive work:
request -> authenticate -> validate -> adapt features -> predict
-> apply decision rule -> record event -> return response
Every step has a distinct responsibility.
Component
Responsibility
Example failure
Gateway
authenticate, limit traffic, assign request ID
unauthorized request reaches the service
Validator
enforce types, ranges, required fields, and schema version
missing value passes silently
Feature adapter
map the source representation to the model contract
measurement units are misinterpreted
Prediction service
load the approved artifact and run inference
wrong model version is loaded
Decision rules
convert probability into a review suggestion
threshold changes without approval
Review queue
present context and preserve human authority
urgent case is hidden or duplicated
Prediction store
retain the event needed for audit and evaluation
outcome cannot be linked to its prediction
The system should fail safely. If required measurements are missing, the model should not invent them. If the prediction service is unavailable, the clinical workflow must continue through an explicit fallback rather than quietly treating failure as low risk.
Operational pathway
Metrics, logs, and traces describe different aspects of the same request:
metrics show aggregate rates, latency, resource use, and saturation;
logs record structured events and relevant diagnostic context; and
traces connect time spent across distributed components.
These signals should share a request_id. Technical telemetry should avoid unnecessary clinical or personally identifying content. Chapter 06 develops the observability and incident-response design.
Learning pathway
Each prediction must be linkable to a later outcome without relying on feature values or patient names as identifiers. A minimal record includes:
Field
Purpose
prediction_id
stable identifier for the prediction event
request_id
correlation across the online technical pathway
case_id
controlled link to the source case
event_time
when the prediction was produced
model_version
exact model artifact used
schema_version
input contract used by the request
malignant_probability
model output retained for later evaluation
suggested_priority
policy output shown to the reviewer
reviewer_action
accepted, changed, or deferred action
outcome and outcome_time
confirmed label and when it became available
The outcome is not available during the original request. It is attached later through a controlled reconciliation job. Chapter 09 develops this feedback and retraining loop.
Separate model output from system action
The model returns a probability. The system applies a policy to produce a suggested review category. Keeping those artifacts separate prevents a policy change from being confused with a model change.
def suggested_priority(probability_malignant: float) ->str:"""Illustrative policy only; thresholds are not clinical guidance."""if probability_malignant >=0.80:return"priority_review"if probability_malignant >=0.40:return"standard_review_with_flag"return"standard_review"
The thresholds are deliberately illustrative. In a real setting they would require evidence, risk analysis, domain approval, version control, validation, and continuing review. The response should identify both versions:
This response communicates a recommendation and its provenance. It does not claim a diagnosis.
Interface contracts
Most production failures occur at boundaries: one component changes while another continues to assume the old behavior. The architecture therefore needs explicit contracts.
Input contract
The input contract should specify:
field names, types, units, and allowable ranges;
required and optional fields;
schema version and compatibility rules;
treatment of missing or duplicated events; and
the error returned for every invalid condition.
Schema validity is necessary but not sufficient. A numeric value can be correctly typed yet semantically wrong because its unit changed or the upstream measurement procedure drifted.
Prediction contract
The prediction service must expose the model version, expected feature order, class mapping, preprocessing version, and response schema. It must refuse to start when the approved artifact or metadata are inconsistent.
Outcome contract
The outcome pathway must define what counts as a verified label, who can record it, when it becomes final, how corrections are handled, and how it links to the original prediction. Otherwise, retraining may optimize against provisional or incorrectly matched labels.
Workload and timing assumptions
Architecture choices should follow an explicit workload rather than the largest technology stack available. The initial teaching assumptions are:
Characteristic
Initial assumption
Architectural implication
average volume
about 50 cases per day
a small service can handle normal throughput
peak arrival
short clinic-opening and post-lunch bursts
queueing and tail latency still need measurement
response need
interactive review support
online pathway should return promptly
availability need
workflow must degrade visibly and safely
documented manual fallback is required
outcome delay
commonly several days
feedback processing is asynchronous
change rate
code changes more often than approved models
code, model, schema, and policy versions are independent
These are design assumptions, not measured service-level objectives. Later chapters will convert observed behavior and decision needs into targets.
Generate the Chapter 02 figures
The executable program creates a reproducible synthetic workload. It illustrates two architectural facts: demand is bursty even when daily volume is modest, and outcome labels arrive much later than predictions.
Figure 3.1: Synthetic case arrivals by hour and the delayed availability of confirmed outcomes.
The simulation does not represent clinical prevalence or validate a clinical workflow. Its purpose is architectural: the request pathway and the feedback pathway operate on different timescales and should be designed accordingly.
Failure analysis across the architecture
A component inventory becomes operationally useful when each failure has a detectable signal and a defined response.
Failure
Observable signal
Safe response
Owner
source fields missing
validation-error rate by field
reject request and request correction
source-data owner
feature semantics changed
distribution shift or data-quality rule failure
stop affected predictions and investigate
data and model owners
model service unavailable
availability alert and failed trace span
expose failure and use documented manual workflow
operations
latency exceeds workflow need
p95/p99 latency and queue depth
shed nonessential work or scale safely
operations
wrong model loaded
startup integrity check or version mismatch
prevent startup or roll back
release owner
prediction not persisted
reconciliation mismatch
retry idempotently; never create a second decision event
application owner
outcome link missing
low label-join coverage
investigate reconciliation; do not infer labels
data owner
clinician routinely overrides
override rate with documented reasons
review policy, usability, and model fitness
decision owner
This table is the beginning of a control plan. Later chapters add automated tests, delivery gates, dashboards, alerts, runbooks, and governance evidence.
Architecture decisions and trade-offs
Synchronous versus asynchronous work
Authentication, validation, inference, and the response belong in the synchronous path because the clinician is waiting for them. Aggregating monitoring data, linking confirmed outcomes, and calculating drift metrics can run asynchronously.
Persisting the prediction event is more subtle. If the response is returned before durable recording, the system may produce an untraceable recommendation. If recording blocks the request, a storage outage can stop the workflow. A production design must choose deliberately—for example, a durable event queue with explicit delivery guarantees—and test the chosen failure behavior.
Modular components versus unnecessary distribution
The logical components in the diagram do not all need to be separate network services. For the initial implementation, validation, feature adaptation, prediction, and policy logic can remain modules within one application. Clear code boundaries preserve testability without creating avoidable network dependencies.
Human review versus automatic action
Human review is not a decorative confirmation button. The interface must show sufficient context, allow correction and override, record reasons appropriately, and avoid automation bias. The clinician remains responsible for interpreting the complete evidence, while the organization remains responsible for the safety and usefulness of the system it provides.
A staged implementation plan
The guide will evolve this architecture without pretending that every component must appear on day one.
Stage
Capability added
Chapters
baseline
containerized prediction service from the preceding guide
data and model monitoring, feedback, retraining, lifecycle controls
08–10
responsible decision system
security, governance, human decision design
11–12
integrated evidence
complete system implementation and review
13
The architecture is therefore a roadmap as well as a diagram. Every later chapter should improve one or more components, interfaces, or controls without changing the intended purpose silently.
Chapter summary
The Diagnostic Support System extends the deployed breast-cancer model service into a complete, human-reviewed workflow. The important architectural lessons are:
the model estimates risk; a separately versioned policy suggests review priority;
the clinician, not the model, determines the appropriate action;
online prediction, operational observation, and delayed learning are different pathways;
every component boundary needs a contract, an owner, a signal, and a safe failure response;
request, prediction, model, schema, policy, and outcome identifiers provide essential traceability;
logical modularity does not require premature microservices; and
architecture must evolve through controlled, observable stages.
Chapter 03 applies DevOps principles to this architecture and establishes how code, configuration, infrastructure, and operational responsibility move together.
Review questions
Why is review priority a safer system output than an automated diagnosis in this case study?
Which work belongs in the online pathway, and which work can occur asynchronously?
Why must model_version and policy_version be recorded separately?
What is the difference between a syntactically valid feature and a semantically valid feature?
Why does a modest daily request count not eliminate the need to examine peak load?
What evidence would show that outcome linkage is incomplete?
Practical exercise
Adapt the reference architecture to a non-medical model service, such as fraud review, equipment maintenance, customer retention, or crop-disease screening.
State the decision being supported and the action that must remain under human control.
Define the system boundary and list three external dependencies.
Separate the online, operational, and learning pathways.
Create a minimal prediction record containing the identifiers needed for audit and later evaluation.
Identify five failure modes and assign a signal, safe response, and owner to each.
Run the Chapter 02 simulation, change the hourly arrival pattern or outcome delay, and explain which architectural assumption should change.