Appendix A — Appendix

Published

Aug 2026

How to use this appendix

This appendix is the operational reference for From Models to Systems. The main chapters explain why production model systems require DevOps, MLOps, observability, governance, and human oversight. This appendix collects the commands, conventions, templates, formulas, and checklists needed to apply those ideas consistently.

Use it when you need to:

  • reproduce the guide environment;
  • locate a chapter’s executable artifacts;
  • run the complete verification workflow;
  • interpret reliability and model-monitoring measures;
  • prepare a release, incident response, retraining decision, or governance review; or
  • diagnose common local, CI, container, and Quarto problems.

The examples assume that commands are run from the repository root on macOS or Linux. Windows users can run the Bash wrappers through Git Bash or Windows Subsystem for Linux, or invoke the corresponding Python programs directly.

Repository conventions

The guide follows one naming rule: executable and generated artifacts begin with the number of the chapter that introduces them. This keeps the narrative, code, evidence, and outputs traceable.

models-to-systems/
├── .github/workflows/       # Continuous integration and delivery
├── app/                     # Minimal decision API
├── data/                    # Raw, processed, and reference data
├── library/                 # Bibliography
├── models/                  # Versioned local model artifacts
├── reports/                 # Human-readable generated reports
├── results/                 # Machine-readable outputs
│   └── figures/             # Generated plots
├── scripts/
│   ├── bash/                # Reproducible command wrappers
│   └── python/              # Executable Python programs
├── tests/                   # Automated tests
├── 00-preface.qmd
├── 01-beyond-model-deployment.qmd
├── ...
├── 13-end-to-end-systems-case-study.qmd
├── 999-appendix.qmd
├── 999-references.qmd
├── Dockerfile
├── requirements.txt
└── _quarto.yml

The main conventions are:

Item Convention Example
Python program scripts/python/<chapter>-<task>.py scripts/python/13-run-end-to-end-system.py
Bash wrapper scripts/bash/<chapter>-<task>.sh scripts/bash/13-run-end-to-end-system.sh
Tabular result results/<chapter>-<name>.csv results/13-system-readiness-summary.csv
Structured report results/<chapter>-<name>.json results/13-system-readiness-report.json
Figure results/figures/<chapter>-<name>.png results/figures/13-end-to-end-system-readiness.png
Chapter source <chapter>-<topic>.qmd 07-scaling-reliability-and-cost.qmd

Generated artifacts are evidence, not hand-edited inputs. When a script changes, rerun it and review the regenerated output before committing both the program and the intended results.

Environment setup

Create and activate the virtual environment

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Confirm that the active interpreter belongs to the repository:

which python
python --version
python -m pip --version

The interpreter path should include models-to-systems/.venv. Activate the environment again whenever a new terminal session begins.

Verify the core toolchain

python -c "import fastapi, matplotlib, numpy, pandas, sklearn; print('Python dependencies OK')"
python -m pytest --version
quarto --version
docker --version

Docker is needed only for container-based exercises. Quarto is installed separately from Python and is required to render the guide.

Reproduce dependencies

Install from the committed dependency file rather than adding packages ad hoc:

python -m pip install -r requirements.txt

When an intentional dependency change is made, update requirements.txt, recreate or refresh the environment, run the tests and chapter workflows, and render the book before committing.

Command reference

Run automated tests

python -m pytest -q

Using python -m pytest ensures that the test runner uses the same interpreter as the active environment.

Start the decision API locally

python -m uvicorn app.main:app --host 127.0.0.1 --port 8000 --reload

In a second terminal, inspect the health endpoint:

curl --fail --silent http://127.0.0.1:8000/health

The interactive API documentation is available at http://127.0.0.1:8000/docs while the service is running.

Build and run the container

docker build --tag decision-api:local .
docker run --detach \
  --name decision-api-local \
  --publish 8000:8000 \
  decision-api:local
curl --fail --silent http://127.0.0.1:8000/health

Inspect and stop the local container:

docker logs decision-api-local
docker stop decision-api-local
docker rm decision-api-local

Use an immutable identifier such as a Git commit SHA for release candidates. A movable tag such as latest is convenient for experimentation but should not be the only identity recorded in deployment evidence.

Run the end-to-end systems checkpoint

bash scripts/bash/13-run-end-to-end-system.sh

The workflow writes:

  • results/13-system-readiness-summary.csv;
  • results/13-system-readiness-report.json; and
  • results/figures/13-end-to-end-system-readiness.png.

Its deterministic teaching scenario should report GO with all seven readiness gates passing. If the result differs, inspect the changed program, policy, dependencies, and generated report instead of editing the report directly.

Render the complete Quarto book

quarto render

Preview during writing:

quarto preview

The configured book output is written to docs/. A clean render completes all chapters without missing-file, cross-reference, or citation warnings.

Configuration templates

The following compact templates illustrate the minimum structure of common operational artifacts. Adapt names and targets to the service rather than copying thresholds without analysis.

Service-level objective

service: decision-api
window: 30d
objectives:
  availability:
    target: 0.990
    indicator: successful_requests / valid_requests
  latency:
    target: 0.950
    indicator: requests_completed_within_250_ms / successful_requests
alerting:
  fast_burn:
    owner: on_call
    response: page
  slow_burn:
    owner: service_team
    response: ticket

Model and policy manifest

model_version: decision-model-v1
data_version: production-window-2026-08
policy_version: decision-policy-v1
code_revision: <git-sha>
artifact_digest: <sha256-digest>
intended_use: risk-based decision support
review_threshold: 0.35
intervention_threshold: 0.75
approval_record: <approval-id>
rollback_target: <previous-artifact-digest>

Monitoring contract

window: 24h
service:
  availability_min: 0.990
  p95_latency_ms_max: 250
  error_rate_max: 0.010
model:
  score_psi_max: 0.200
  false_negative_rate_max: 0.150
human_review:
  review_rate_max: 0.450
required_dimensions:
  - model_version
  - policy_version
  - deployment_environment

Incident record

incident_id: <id>
detected_at: <iso-8601-time>
severity: <severity>
affected_service: decision-api
model_version: <version>
policy_version: <version>
customer_effect: <observed-effect>
immediate_action: <mitigation>
owner: <role-or-team>
status: investigating
rollback_target: <version-or-not-applicable>
follow_up_due: <date>

Do not place secrets, access tokens, raw sensitive features, or personally identifying values in these manifests. Store sensitive configuration in an approved secrets system and reference it by name.

Metric definitions

Metrics should have an explicit population, time window, unit, and aggregation rule. The following definitions match the guide’s operating vocabulary.

Availability

For a defined measurement window:

\[ \text{Availability} = \frac{\text{valid requests served successfully}}{\text{valid requests received}} \]

The denominator must be documented. Excluding failed dependencies or selected error classes can make an unreliable service appear healthy.

Error rate

\[ \text{Error rate} = \frac{\text{failed requests}}{\text{requests received}} \]

Separate client validation failures, server failures, timeouts, and dependency failures when those categories require different owners or responses.

Latency percentile

The p95 latency is the value at or below which 95% of measured request latencies fall. Percentiles reveal tail behaviour that an arithmetic mean can hide. State whether the measurement covers model inference alone or the complete request pathway.

End-to-end availability

When independent components must all succeed, a first approximation is:

\[ A_{\text{system}} = \prod_{i=1}^{n} A_i \]

This product is a useful warning that individually reliable components can form a less reliable pathway. Real dependencies may not fail independently, so production analysis should also examine correlated failures and shared infrastructure.

Error budget

For an availability objective \(SLO\) over \(N\) eligible events:

\[ \text{Error budget} = (1 - SLO)N \]

Budget consumption connects reliability evidence to delivery pace. Rapid consumption should constrain risky changes and prioritize stabilization.

Population stability index

For reference proportion \(r_i\) and current proportion \(c_i\) in bin \(i\):

\[ PSI = \sum_i (c_i-r_i)\ln\left(\frac{c_i}{r_i}\right) \]

PSI signals a distribution change; it does not explain the cause or prove that model quality declined. Investigate binning, sample size, seasonality, pipeline changes, and outcome evidence before selecting a response.

False-negative rate

\[ \text{False-negative rate} = \frac{FN}{TP + FN} \]

In a decision system, define what counts as a positive action. In the Chapter 13 policy, human review and urgent intervention both prevent a positive case from being treated as a routine automated outcome. The calculation therefore follows the deployed decision policy, not only a generic classifier threshold.

Unit cost

\[ \text{Unit cost} = \frac{\text{total attributable operating cost}}{\text{completed decision units}} \]

Define the decision unit and included costs. Infrastructure cost alone can omit review labour, external APIs, observability, and incident response.

Readiness checklists

Before merging a change

  • The change has a clear owner and purpose.
  • Unit, contract, integration, and relevant model tests pass.
  • No secret or sensitive record is committed.
  • Dependency and schema changes are explicit.
  • Generated evidence was recreated from source programs.
  • Documentation and runbooks reflect changed behaviour.
  • A rollback or safe-disable path is known.

Before promoting a release

  • The candidate artifact is immutable and identified by digest or commit.
  • The same artifact passed the preceding environment.
  • Configuration differences between environments are reviewed.
  • Health, readiness, and prediction contracts pass.
  • Service, data, model, outcome, and human-workload gates pass.
  • Model, data, code, policy, and approval versions are linked.
  • Dashboards, alerts, owners, and runbooks are active.
  • The rollback target has been verified.

During an incident

  • Confirm customer and decision impact before optimizing the diagnosis.
  • Name an incident lead and record a timeline.
  • Stabilize the system by rollback, traffic shift, policy restriction, or safe shutdown.
  • Preserve evidence without placing sensitive values in broad-access logs.
  • Identify whether the failure is in delivery, infrastructure, data, model, policy, human workflow, or governance.
  • Communicate status at a defined cadence.
  • Record follow-up actions with owners and deadlines.

Before initiating retraining

  • The triggering signal is reproducible and investigated.
  • Data or outcome quality is sufficient for training and evaluation.
  • The problem cannot be solved more directly by repairing infrastructure, data contracts, or policy.
  • Training data use is permitted and lineage is recorded.
  • Subgroup and temporal evaluation plans are defined.
  • Acceptance criteria are specified before candidate results are inspected.
  • Retraining creates a candidate; it does not overwrite production automatically.

Before retiring a system

  • The replacement, manual process, or service shutdown path is approved.
  • Upstream producers and downstream consumers are identified.
  • Traffic, scheduled jobs, alerts, and retraining workflows are disabled safely.
  • Required evidence and audit records are retained according to policy.
  • Models, datasets, secrets, and infrastructure are archived or removed appropriately.
  • Owners and users are notified, and the retirement decision is documented.

Troubleshooting

Symptom Likely cause Corrective action
ModuleNotFoundError environment inactive or dependencies missing activate .venv; run python -m pip install -r requirements.txt
pytest uses unexpected packages executable comes from another Python installation use python -m pytest -q; verify which python
API health check cannot connect server not started, wrong port, or startup failure inspect the Uvicorn terminal; confirm host and port; read the traceback
Port 8000 is already allocated another service or container owns the port stop that process or publish the container on another local port
Docker daemon connection error Docker Desktop or engine is unavailable start Docker and rerun docker version
Container starts but health fails incorrect command, dependency, model path, or bind address inspect docker logs; ensure the app binds to 0.0.0.0 in the container
Bash wrapper uses the wrong interpreter python resolves outside .venv activate .venv or set PYTHON_BIN to the intended interpreter
Matplotlib cache warning default cache directory is not writable set MPLCONFIGDIR to a writable temporary directory
Generated figure is missing chapter program did not complete or ran from an unexpected location run its Bash wrapper from the repository root and inspect the error
Citeproc reports a missing key citation key is absent or misspelled in library/references.bib match the QMD key and BibTeX identifier exactly, then render again
Quarto cross-reference is unresolved label is missing or does not use the expected prefix give sections sec- labels and figures fig- labels; verify the reference
CI passes locally but fails remotely environment, path, version, or uncommitted-file difference reproduce the CI version; inspect the workflow log and repository status
Readiness decision changes unexpectedly source, seed, threshold, or dependency changed compare the program and manifest; regenerate all Chapter 13 outputs

Avoid fixing generated CSV, JSON, HTML, or PNG files by hand. Correct the source, configuration, or environment and regenerate the evidence.

Evidence and versioning matrix

A production decision should be reconstructable from linked identities.

Evidence Minimum identity Why it matters
Source code commit SHA identifies implementation
Service artifact image digest identifies deployed bytes
Model registry version or artifact digest identifies scoring behaviour
Training data immutable dataset or snapshot version establishes lineage
Feature/schema contract contract version identifies accepted inputs
Decision policy policy version explains action routing
Deployment environment and release identifier locates operational change
Monitoring window start, end, and query/version makes metrics reproducible
Approval approver role, record ID, and time establishes authority
Incident or override record ID and owner preserves exceptions and response

Version the model and decision policy independently. A threshold or routing change can alter real decisions even when the model artifact is unchanged.

Glossary

Alert
A notification that a defined condition requires attention. An alert should have an owner and response; not every dashboard movement should page a person.

Artifact
An immutable output such as a container image, trained model, package, or report that can be identified and promoted.

Canary deployment
A release strategy that sends a limited portion of live traffic to a candidate before broader promotion.

Concept drift
A change in the relationship between inputs and the outcome the system is intended to predict.

Data drift
A change in the distribution or characteristics of model inputs or scores relative to a reference.

Decision policy
Versioned rules that translate model outputs and contextual constraints into actions such as automation, review, or intervention.

Error budget
The amount of failure permitted by a service-level objective during a defined window.

Human-in-the-loop
A system design in which people perform specified review, judgment, escalation, or override functions rather than serving as an undefined fallback.

Immutable promotion
Moving the same identified artifact through environments instead of rebuilding a potentially different artifact at each stage.

Incident
An event that degrades service, decision quality, safety, security, compliance, or user outcomes and requires coordinated response.

Model registry
A controlled record of model candidates, versions, metadata, evaluation evidence, stages, and approvals.

Observability
The ability to investigate a system’s internal state from telemetry such as metrics, logs, traces, and domain events.

Rollback
Restoring a previously verified model, policy, configuration, or deployment after a change causes unacceptable behaviour.

Runbook
An operational procedure that links a known condition to diagnostic, mitigation, escalation, and verification steps.

Service-level indicator (SLI)
A measured property of service behaviour, such as the proportion of successful requests.

Service-level objective (SLO)
A target for an SLI over a defined period, based on the reliability users need.

Shadow deployment
Running a candidate on copied production inputs without allowing its output to control live decisions.

Telemetry
Operational evidence emitted by a system, including metrics, logs, traces, events, and model-specific signals.

Final reproducibility check

Before declaring the repository complete, run:

source .venv/bin/activate
python -m pip install -r requirements.txt
python -m pytest -q
bash scripts/bash/13-run-end-to-end-system.sh
quarto render
git status --short

Review the test output, the Chapter 13 readiness report and figure, the Quarto render log, and every changed file reported by Git. A successful run should leave the guide with passing tests, reproducible system evidence, a clean render, and only intentional repository changes.

This is the final operating lesson of the guide: production readiness is not a single test or deployment event. It is a repeatable chain of evidence connecting code, infrastructure, data, models, policies, people, and outcomes.