System Case Study and Architecture

Published

Aug 2026

Introduction

Chapter Beyond Model Deployment established that a deployed model service is only one component of a decision system. This chapter makes that distinction concrete by introducing the case study used throughout the rest of the guide.

The starting point is the validated breast-cancer prediction API built in the Model Deployment guide. It loads a fitted preprocessing-and-classification pipeline, validates eight numeric measurements, and returns the predicted class together with the estimated probability of malignancy. Here, that service becomes one component in a larger diagnostic-support workflow.

The objective is not to design a deployable medical device. It is to create a realistic educational architecture in which operational reliability, model monitoring, feedback, governance, and human judgment can be studied together.

Educational use only

The Wisconsin Diagnostic Breast Cancer dataset is a small teaching dataset with 569 observations and 30 numeric features (scikit-learn developers 2026). The case study is not clinically validated and must not be used to diagnose patients, choose treatment, or replace professional judgment.

Learning objectives

By the end of this chapter, you should be able to:

  • define the purpose, users, inputs, outputs, and boundaries of a model-enabled system;
  • distinguish the online prediction pathway from asynchronous operational and learning pathways;
  • assign responsibilities to system components rather than treating the model API as the whole application;
  • identify failure modes at interfaces between components;
  • translate a decision workflow into measurable service and data requirements;
  • explain why prediction records require model, request, and outcome identifiers; and
  • use an architecture diagram as a living operational contract.

The continuous case study

The fictional organization in this guide operates several diagnostic clinics. A laboratory system records measurements derived from a digitized image of a breast-mass sample. A clinician reviews the measurements, the model output, and the wider clinical evidence before determining the next appropriate step.

We will call the complete application the Diagnostic Support System, or DSS.

Intended purpose

The DSS prioritizes cases for clinician review. It does not make a diagnosis. A high estimated probability of malignancy should make a case visible to the reviewing clinician sooner; it must not automatically label a patient, cancel another assessment, or determine treatment.

Primary users

User Need Authority
Laboratory professional submit a complete, valid measurement record may correct or resubmit source measurements
Reviewing clinician see the risk estimate with context and limitations accepts, rejects, or overrides the suggested priority
Operations team keep the technical pathway available and recoverable deploys, rolls back, scales, and responds to incidents
Model owner assess model behavior and approve model changes proposes or rejects model promotion
Governance owner verify approved use, access, auditability, and review can pause the system or restrict its use

The decision and the outcome

The immediate decision is review priority, not diagnosis. The system returns a risk estimate and a suggested queue category. A clinician makes the final prioritization decision.

The eventual outcome is the confirmed diagnosis recorded after the appropriate clinical process. That outcome may arrive days or weeks after the prediction. This delay is crucial: operational monitoring can detect a failing API immediately, while model-performance monitoring cannot calculate real-world false negatives until verified outcomes are available.

Define the system boundary

A useful system boundary states what the team operates and what remains an external dependency.

For this guide, the DSS begins when the laboratory system submits a measurement event and ends when the confirmed outcome is linked to the original prediction. The laboratory instrument, the procedures used to obtain the sample, and the downstream diagnostic process remain outside the technical boundary, but their assumptions and failures still affect the system.

Inside the managed boundary External dependency or context
API gateway and authentication laboratory instrument and measurement procedure
input validation and feature adapter identity and access provider
prediction service and model artifact clinic network connectivity
decision rules and review queue wider clinical evidence and clinician judgment
prediction and audit records diagnostic confirmation process
metrics, logs, traces, and alerts organizational policies and regulatory obligations
outcome-linking and monitoring jobs source-system ownership and data correction

Being outside the boundary does not make a dependency unimportant. It means the team must define an interface, an owner, an observable expectation, and a response when the expectation is violated.

The reference architecture

The architecture contains three connected pathways:

  1. the online pathway validates a request and returns decision support;
  2. the operational pathway records telemetry and detects technical failure; and
  3. the learning pathway links delayed outcomes to earlier predictions and evaluates continuing model fitness.
Code
flowchart TD
    A["Laboratory system"] --> B["Gateway and validation"]
    B --> C["Prediction service"]
    C --> D["Decision rules"]
    D --> E["Clinician review queue"]
    E --> F["Confirmed outcome"]
    B -. "telemetry" .-> G["Observability platform"]
    C -. "prediction record" .-> H["Audit and prediction store"]
    F -. "delayed label" .-> H
    H --> I["Monitoring and evaluation"]
    I -. "approved change" .-> C

flowchart TD
    A["Laboratory system"] --> B["Gateway and validation"]
    B --> C["Prediction service"]
    C --> D["Decision rules"]
    D --> E["Clinician review queue"]
    E --> F["Confirmed outcome"]
    B -. "telemetry" .-> G["Observability platform"]
    C -. "prediction record" .-> H["Audit and prediction store"]
    F -. "delayed label" .-> H
    H --> I["Monitoring and evaluation"]
    I -. "approved change" .-> C

The dashed connections are not secondary in importance. They are separated because they should not unnecessarily delay the clinician-facing response.

Online prediction pathway

The online pathway handles the latency-sensitive work:

request -> authenticate -> validate -> adapt features -> predict
        -> apply decision rule -> record event -> return response

Every step has a distinct responsibility.

Component Responsibility Example failure
Gateway authenticate, limit traffic, assign request ID unauthorized request reaches the service
Validator enforce types, ranges, required fields, and schema version missing value passes silently
Feature adapter map the source representation to the model contract measurement units are misinterpreted
Prediction service load the approved artifact and run inference wrong model version is loaded
Decision rules convert probability into a review suggestion threshold changes without approval
Review queue present context and preserve human authority urgent case is hidden or duplicated
Prediction store retain the event needed for audit and evaluation outcome cannot be linked to its prediction

The system should fail safely. If required measurements are missing, the model should not invent them. If the prediction service is unavailable, the clinical workflow must continue through an explicit fallback rather than quietly treating failure as low risk.

Operational pathway

Metrics, logs, and traces describe different aspects of the same request:

  • metrics show aggregate rates, latency, resource use, and saturation;
  • logs record structured events and relevant diagnostic context; and
  • traces connect time spent across distributed components.

These signals should share a request_id. Technical telemetry should avoid unnecessary clinical or personally identifying content. Chapter 06 develops the observability and incident-response design.

Learning pathway

Each prediction must be linkable to a later outcome without relying on feature values or patient names as identifiers. A minimal record includes:

Field Purpose
prediction_id stable identifier for the prediction event
request_id correlation across the online technical pathway
case_id controlled link to the source case
event_time when the prediction was produced
model_version exact model artifact used
schema_version input contract used by the request
malignant_probability model output retained for later evaluation
suggested_priority policy output shown to the reviewer
reviewer_action accepted, changed, or deferred action
outcome and outcome_time confirmed label and when it became available

The outcome is not available during the original request. It is attached later through a controlled reconciliation job. Chapter 09 develops this feedback and retraining loop.

Separate model output from system action

The model returns a probability. The system applies a policy to produce a suggested review category. Keeping those artifacts separate prevents a policy change from being confused with a model change.

def suggested_priority(probability_malignant: float) -> str:
    """Illustrative policy only; thresholds are not clinical guidance."""
    if probability_malignant >= 0.80:
        return "priority_review"
    if probability_malignant >= 0.40:
        return "standard_review_with_flag"
    return "standard_review"

The thresholds are deliberately illustrative. In a real setting they would require evidence, risk analysis, domain approval, version control, validation, and continuing review. The response should identify both versions:

{
  "prediction_id": "pred-01J...",
  "model_version": "breast-cancer-logistic-1.0.0",
  "policy_version": "review-priority-1.0.0",
  "malignant_probability": 0.87,
  "suggested_priority": "priority_review",
  "requires_clinician_review": true
}

This response communicates a recommendation and its provenance. It does not claim a diagnosis.

Interface contracts

Most production failures occur at boundaries: one component changes while another continues to assume the old behavior. The architecture therefore needs explicit contracts.

Input contract

The input contract should specify:

  • field names, types, units, and allowable ranges;
  • required and optional fields;
  • schema version and compatibility rules;
  • treatment of missing or duplicated events; and
  • the error returned for every invalid condition.

Schema validity is necessary but not sufficient. A numeric value can be correctly typed yet semantically wrong because its unit changed or the upstream measurement procedure drifted.

Prediction contract

The prediction service must expose the model version, expected feature order, class mapping, preprocessing version, and response schema. It must refuse to start when the approved artifact or metadata are inconsistent.

Outcome contract

The outcome pathway must define what counts as a verified label, who can record it, when it becomes final, how corrections are handled, and how it links to the original prediction. Otherwise, retraining may optimize against provisional or incorrectly matched labels.

Workload and timing assumptions

Architecture choices should follow an explicit workload rather than the largest technology stack available. The initial teaching assumptions are:

Characteristic Initial assumption Architectural implication
average volume about 50 cases per day a small service can handle normal throughput
peak arrival short clinic-opening and post-lunch bursts queueing and tail latency still need measurement
response need interactive review support online pathway should return promptly
availability need workflow must degrade visibly and safely documented manual fallback is required
outcome delay commonly several days feedback processing is asynchronous
change rate code changes more often than approved models code, model, schema, and policy versions are independent

These are design assumptions, not measured service-level objectives. Later chapters will convert observed behavior and decision needs into targets.

Generate the Chapter 02 figures

The executable program creates a reproducible synthetic workload. It illustrates two architectural facts: demand is bursty even when daily volume is modest, and outcome labels arrive much later than predictions.

Run it from the project root:

python scripts/python/02-simulate-system-workload.py

or use the Bash wrapper:

bash scripts/bash/02-simulate-system-workload.sh

The program writes:

  • results/02-system-workload-summary.csv;
  • results/02-synthetic-case-events.csv; and
  • results/figures/02-system-workload-and-feedback.png.
Two-panel chart. Hourly case arrivals show morning and afternoon peaks. A cumulative distribution shows that confirmed outcomes become available over several days rather than during the prediction request.
Figure 3.1: Synthetic case arrivals by hour and the delayed availability of confirmed outcomes.

The simulation does not represent clinical prevalence or validate a clinical workflow. Its purpose is architectural: the request pathway and the feedback pathway operate on different timescales and should be designed accordingly.

Failure analysis across the architecture

A component inventory becomes operationally useful when each failure has a detectable signal and a defined response.

Failure Observable signal Safe response Owner
source fields missing validation-error rate by field reject request and request correction source-data owner
feature semantics changed distribution shift or data-quality rule failure stop affected predictions and investigate data and model owners
model service unavailable availability alert and failed trace span expose failure and use documented manual workflow operations
latency exceeds workflow need p95/p99 latency and queue depth shed nonessential work or scale safely operations
wrong model loaded startup integrity check or version mismatch prevent startup or roll back release owner
prediction not persisted reconciliation mismatch retry idempotently; never create a second decision event application owner
outcome link missing low label-join coverage investigate reconciliation; do not infer labels data owner
clinician routinely overrides override rate with documented reasons review policy, usability, and model fitness decision owner

This table is the beginning of a control plan. Later chapters add automated tests, delivery gates, dashboards, alerts, runbooks, and governance evidence.

Architecture decisions and trade-offs

Synchronous versus asynchronous work

Authentication, validation, inference, and the response belong in the synchronous path because the clinician is waiting for them. Aggregating monitoring data, linking confirmed outcomes, and calculating drift metrics can run asynchronously.

Persisting the prediction event is more subtle. If the response is returned before durable recording, the system may produce an untraceable recommendation. If recording blocks the request, a storage outage can stop the workflow. A production design must choose deliberately—for example, a durable event queue with explicit delivery guarantees—and test the chosen failure behavior.

Modular components versus unnecessary distribution

The logical components in the diagram do not all need to be separate network services. For the initial implementation, validation, feature adaptation, prediction, and policy logic can remain modules within one application. Clear code boundaries preserve testability without creating avoidable network dependencies.

Human review versus automatic action

Human review is not a decorative confirmation button. The interface must show sufficient context, allow correction and override, record reasons appropriately, and avoid automation bias. The clinician remains responsible for interpreting the complete evidence, while the organization remains responsible for the safety and usefulness of the system it provides.

A staged implementation plan

The guide will evolve this architecture without pretending that every component must appear on day one.

Stage Capability added Chapters
baseline containerized prediction service from the preceding guide starting point
repeatable operation DevOps practices, CI/CD, environments, controlled releases 03–05
dependable runtime observability, incident response, scaling, reliability, cost 06–07
learning system data and model monitoring, feedback, retraining, lifecycle controls 08–10
responsible decision system security, governance, human decision design 11–12
integrated evidence complete system implementation and review 13

The architecture is therefore a roadmap as well as a diagram. Every later chapter should improve one or more components, interfaces, or controls without changing the intended purpose silently.

Chapter summary

The Diagnostic Support System extends the deployed breast-cancer model service into a complete, human-reviewed workflow. The important architectural lessons are:

  • the model estimates risk; a separately versioned policy suggests review priority;
  • the clinician, not the model, determines the appropriate action;
  • online prediction, operational observation, and delayed learning are different pathways;
  • every component boundary needs a contract, an owner, a signal, and a safe failure response;
  • request, prediction, model, schema, policy, and outcome identifiers provide essential traceability;
  • logical modularity does not require premature microservices; and
  • architecture must evolve through controlled, observable stages.

Chapter 03 applies DevOps principles to this architecture and establishes how code, configuration, infrastructure, and operational responsibility move together.

Review questions

  1. Why is review priority a safer system output than an automated diagnosis in this case study?
  2. Which work belongs in the online pathway, and which work can occur asynchronously?
  3. Why must model_version and policy_version be recorded separately?
  4. What is the difference between a syntactically valid feature and a semantically valid feature?
  5. Why does a modest daily request count not eliminate the need to examine peak load?
  6. What evidence would show that outcome linkage is incomplete?

Practical exercise

Adapt the reference architecture to a non-medical model service, such as fraud review, equipment maintenance, customer retention, or crop-disease screening.

  1. State the decision being supported and the action that must remain under human control.
  2. Define the system boundary and list three external dependencies.
  3. Separate the online, operational, and learning pathways.
  4. Create a minimal prediction record containing the identifiers needed for audit and later evaluation.
  5. Identify five failure modes and assign a signal, safe response, and owner to each.
  6. Run the Chapter 02 simulation, change the hourly arrival pattern or outcome delay, and explain which architectural assumption should change.