Appendix A — Appendix
How to use this appendix
This appendix is the operational reference for From Models to Systems. The main chapters explain why production model systems require DevOps, MLOps, observability, governance, and human oversight. This appendix collects the commands, conventions, templates, formulas, and checklists needed to apply those ideas consistently.
Use it when you need to:
- reproduce the guide environment;
- locate a chapter’s executable artifacts;
- run the complete verification workflow;
- interpret reliability and model-monitoring measures;
- prepare a release, incident response, retraining decision, or governance review; or
- diagnose common local, CI, container, and Quarto problems.
The examples assume that commands are run from the repository root on macOS or Linux. Windows users can run the Bash wrappers through Git Bash or Windows Subsystem for Linux, or invoke the corresponding Python programs directly.
Repository conventions
The guide follows one naming rule: executable and generated artifacts begin with the number of the chapter that introduces them. This keeps the narrative, code, evidence, and outputs traceable.
models-to-systems/
├── .github/workflows/ # Continuous integration and delivery
├── app/ # Minimal decision API
├── data/ # Raw, processed, and reference data
├── library/ # Bibliography
├── models/ # Versioned local model artifacts
├── reports/ # Human-readable generated reports
├── results/ # Machine-readable outputs
│ └── figures/ # Generated plots
├── scripts/
│ ├── bash/ # Reproducible command wrappers
│ └── python/ # Executable Python programs
├── tests/ # Automated tests
├── 00-preface.qmd
├── 01-beyond-model-deployment.qmd
├── ...
├── 13-end-to-end-systems-case-study.qmd
├── 999-appendix.qmd
├── 999-references.qmd
├── Dockerfile
├── requirements.txt
└── _quarto.yml
The main conventions are:
| Item | Convention | Example |
|---|---|---|
| Python program | scripts/python/<chapter>-<task>.py |
scripts/python/13-run-end-to-end-system.py |
| Bash wrapper | scripts/bash/<chapter>-<task>.sh |
scripts/bash/13-run-end-to-end-system.sh |
| Tabular result | results/<chapter>-<name>.csv |
results/13-system-readiness-summary.csv |
| Structured report | results/<chapter>-<name>.json |
results/13-system-readiness-report.json |
| Figure | results/figures/<chapter>-<name>.png |
results/figures/13-end-to-end-system-readiness.png |
| Chapter source | <chapter>-<topic>.qmd |
07-scaling-reliability-and-cost.qmd |
Generated artifacts are evidence, not hand-edited inputs. When a script changes, rerun it and review the regenerated output before committing both the program and the intended results.
Environment setup
Create and activate the virtual environment
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtConfirm that the active interpreter belongs to the repository:
which python
python --version
python -m pip --versionThe interpreter path should include models-to-systems/.venv. Activate the environment again whenever a new terminal session begins.
Verify the core toolchain
python -c "import fastapi, matplotlib, numpy, pandas, sklearn; print('Python dependencies OK')"
python -m pytest --version
quarto --version
docker --versionDocker is needed only for container-based exercises. Quarto is installed separately from Python and is required to render the guide.
Reproduce dependencies
Install from the committed dependency file rather than adding packages ad hoc:
python -m pip install -r requirements.txtWhen an intentional dependency change is made, update requirements.txt, recreate or refresh the environment, run the tests and chapter workflows, and render the book before committing.
Command reference
Run automated tests
python -m pytest -qUsing python -m pytest ensures that the test runner uses the same interpreter as the active environment.
Start the decision API locally
python -m uvicorn app.main:app --host 127.0.0.1 --port 8000 --reloadIn a second terminal, inspect the health endpoint:
curl --fail --silent http://127.0.0.1:8000/healthThe interactive API documentation is available at http://127.0.0.1:8000/docs while the service is running.
Build and run the container
docker build --tag decision-api:local .
docker run --detach \
--name decision-api-local \
--publish 8000:8000 \
decision-api:local
curl --fail --silent http://127.0.0.1:8000/healthInspect and stop the local container:
docker logs decision-api-local
docker stop decision-api-local
docker rm decision-api-localUse an immutable identifier such as a Git commit SHA for release candidates. A movable tag such as latest is convenient for experimentation but should not be the only identity recorded in deployment evidence.
Run the end-to-end systems checkpoint
bash scripts/bash/13-run-end-to-end-system.shThe workflow writes:
results/13-system-readiness-summary.csv;results/13-system-readiness-report.json; andresults/figures/13-end-to-end-system-readiness.png.
Its deterministic teaching scenario should report GO with all seven readiness gates passing. If the result differs, inspect the changed program, policy, dependencies, and generated report instead of editing the report directly.
Render the complete Quarto book
quarto renderPreview during writing:
quarto previewThe configured book output is written to docs/. A clean render completes all chapters without missing-file, cross-reference, or citation warnings.
Configuration templates
The following compact templates illustrate the minimum structure of common operational artifacts. Adapt names and targets to the service rather than copying thresholds without analysis.
Service-level objective
service: decision-api
window: 30d
objectives:
availability:
target: 0.990
indicator: successful_requests / valid_requests
latency:
target: 0.950
indicator: requests_completed_within_250_ms / successful_requests
alerting:
fast_burn:
owner: on_call
response: page
slow_burn:
owner: service_team
response: ticketModel and policy manifest
model_version: decision-model-v1
data_version: production-window-2026-08
policy_version: decision-policy-v1
code_revision: <git-sha>
artifact_digest: <sha256-digest>
intended_use: risk-based decision support
review_threshold: 0.35
intervention_threshold: 0.75
approval_record: <approval-id>
rollback_target: <previous-artifact-digest>Monitoring contract
window: 24h
service:
availability_min: 0.990
p95_latency_ms_max: 250
error_rate_max: 0.010
model:
score_psi_max: 0.200
false_negative_rate_max: 0.150
human_review:
review_rate_max: 0.450
required_dimensions:
- model_version
- policy_version
- deployment_environmentIncident record
incident_id: <id>
detected_at: <iso-8601-time>
severity: <severity>
affected_service: decision-api
model_version: <version>
policy_version: <version>
customer_effect: <observed-effect>
immediate_action: <mitigation>
owner: <role-or-team>
status: investigating
rollback_target: <version-or-not-applicable>
follow_up_due: <date>Do not place secrets, access tokens, raw sensitive features, or personally identifying values in these manifests. Store sensitive configuration in an approved secrets system and reference it by name.
Metric definitions
Metrics should have an explicit population, time window, unit, and aggregation rule. The following definitions match the guide’s operating vocabulary.
Availability
For a defined measurement window:
\[ \text{Availability} = \frac{\text{valid requests served successfully}}{\text{valid requests received}} \]
The denominator must be documented. Excluding failed dependencies or selected error classes can make an unreliable service appear healthy.
Error rate
\[ \text{Error rate} = \frac{\text{failed requests}}{\text{requests received}} \]
Separate client validation failures, server failures, timeouts, and dependency failures when those categories require different owners or responses.
Latency percentile
The p95 latency is the value at or below which 95% of measured request latencies fall. Percentiles reveal tail behaviour that an arithmetic mean can hide. State whether the measurement covers model inference alone or the complete request pathway.
End-to-end availability
When independent components must all succeed, a first approximation is:
\[ A_{\text{system}} = \prod_{i=1}^{n} A_i \]
This product is a useful warning that individually reliable components can form a less reliable pathway. Real dependencies may not fail independently, so production analysis should also examine correlated failures and shared infrastructure.
Error budget
For an availability objective \(SLO\) over \(N\) eligible events:
\[ \text{Error budget} = (1 - SLO)N \]
Budget consumption connects reliability evidence to delivery pace. Rapid consumption should constrain risky changes and prioritize stabilization.
Population stability index
For reference proportion \(r_i\) and current proportion \(c_i\) in bin \(i\):
\[ PSI = \sum_i (c_i-r_i)\ln\left(\frac{c_i}{r_i}\right) \]
PSI signals a distribution change; it does not explain the cause or prove that model quality declined. Investigate binning, sample size, seasonality, pipeline changes, and outcome evidence before selecting a response.
False-negative rate
\[ \text{False-negative rate} = \frac{FN}{TP + FN} \]
In a decision system, define what counts as a positive action. In the Chapter 13 policy, human review and urgent intervention both prevent a positive case from being treated as a routine automated outcome. The calculation therefore follows the deployed decision policy, not only a generic classifier threshold.
Unit cost
\[ \text{Unit cost} = \frac{\text{total attributable operating cost}}{\text{completed decision units}} \]
Define the decision unit and included costs. Infrastructure cost alone can omit review labour, external APIs, observability, and incident response.
Readiness checklists
Before merging a change
- The change has a clear owner and purpose.
- Unit, contract, integration, and relevant model tests pass.
- No secret or sensitive record is committed.
- Dependency and schema changes are explicit.
- Generated evidence was recreated from source programs.
- Documentation and runbooks reflect changed behaviour.
- A rollback or safe-disable path is known.
Before promoting a release
- The candidate artifact is immutable and identified by digest or commit.
- The same artifact passed the preceding environment.
- Configuration differences between environments are reviewed.
- Health, readiness, and prediction contracts pass.
- Service, data, model, outcome, and human-workload gates pass.
- Model, data, code, policy, and approval versions are linked.
- Dashboards, alerts, owners, and runbooks are active.
- The rollback target has been verified.
During an incident
- Confirm customer and decision impact before optimizing the diagnosis.
- Name an incident lead and record a timeline.
- Stabilize the system by rollback, traffic shift, policy restriction, or safe shutdown.
- Preserve evidence without placing sensitive values in broad-access logs.
- Identify whether the failure is in delivery, infrastructure, data, model, policy, human workflow, or governance.
- Communicate status at a defined cadence.
- Record follow-up actions with owners and deadlines.
Before initiating retraining
- The triggering signal is reproducible and investigated.
- Data or outcome quality is sufficient for training and evaluation.
- The problem cannot be solved more directly by repairing infrastructure, data contracts, or policy.
- Training data use is permitted and lineage is recorded.
- Subgroup and temporal evaluation plans are defined.
- Acceptance criteria are specified before candidate results are inspected.
- Retraining creates a candidate; it does not overwrite production automatically.
Before retiring a system
- The replacement, manual process, or service shutdown path is approved.
- Upstream producers and downstream consumers are identified.
- Traffic, scheduled jobs, alerts, and retraining workflows are disabled safely.
- Required evidence and audit records are retained according to policy.
- Models, datasets, secrets, and infrastructure are archived or removed appropriately.
- Owners and users are notified, and the retirement decision is documented.
Troubleshooting
| Symptom | Likely cause | Corrective action |
|---|---|---|
ModuleNotFoundError |
environment inactive or dependencies missing | activate .venv; run python -m pip install -r requirements.txt |
pytest uses unexpected packages |
executable comes from another Python installation | use python -m pytest -q; verify which python |
| API health check cannot connect | server not started, wrong port, or startup failure | inspect the Uvicorn terminal; confirm host and port; read the traceback |
| Port 8000 is already allocated | another service or container owns the port | stop that process or publish the container on another local port |
| Docker daemon connection error | Docker Desktop or engine is unavailable | start Docker and rerun docker version |
| Container starts but health fails | incorrect command, dependency, model path, or bind address | inspect docker logs; ensure the app binds to 0.0.0.0 in the container |
| Bash wrapper uses the wrong interpreter | python resolves outside .venv |
activate .venv or set PYTHON_BIN to the intended interpreter |
| Matplotlib cache warning | default cache directory is not writable | set MPLCONFIGDIR to a writable temporary directory |
| Generated figure is missing | chapter program did not complete or ran from an unexpected location | run its Bash wrapper from the repository root and inspect the error |
| Citeproc reports a missing key | citation key is absent or misspelled in library/references.bib |
match the QMD key and BibTeX identifier exactly, then render again |
| Quarto cross-reference is unresolved | label is missing or does not use the expected prefix | give sections sec- labels and figures fig- labels; verify the reference |
| CI passes locally but fails remotely | environment, path, version, or uncommitted-file difference | reproduce the CI version; inspect the workflow log and repository status |
| Readiness decision changes unexpectedly | source, seed, threshold, or dependency changed | compare the program and manifest; regenerate all Chapter 13 outputs |
Avoid fixing generated CSV, JSON, HTML, or PNG files by hand. Correct the source, configuration, or environment and regenerate the evidence.
Evidence and versioning matrix
A production decision should be reconstructable from linked identities.
| Evidence | Minimum identity | Why it matters |
|---|---|---|
| Source code | commit SHA | identifies implementation |
| Service artifact | image digest | identifies deployed bytes |
| Model | registry version or artifact digest | identifies scoring behaviour |
| Training data | immutable dataset or snapshot version | establishes lineage |
| Feature/schema contract | contract version | identifies accepted inputs |
| Decision policy | policy version | explains action routing |
| Deployment | environment and release identifier | locates operational change |
| Monitoring window | start, end, and query/version | makes metrics reproducible |
| Approval | approver role, record ID, and time | establishes authority |
| Incident or override | record ID and owner | preserves exceptions and response |
Version the model and decision policy independently. A threshold or routing change can alter real decisions even when the model artifact is unchanged.
Glossary
Alert
A notification that a defined condition requires attention. An alert should have an owner and response; not every dashboard movement should page a person.
Artifact
An immutable output such as a container image, trained model, package, or report that can be identified and promoted.
Canary deployment
A release strategy that sends a limited portion of live traffic to a candidate before broader promotion.
Concept drift
A change in the relationship between inputs and the outcome the system is intended to predict.
Data drift
A change in the distribution or characteristics of model inputs or scores relative to a reference.
Decision policy
Versioned rules that translate model outputs and contextual constraints into actions such as automation, review, or intervention.
Error budget
The amount of failure permitted by a service-level objective during a defined window.
Human-in-the-loop
A system design in which people perform specified review, judgment, escalation, or override functions rather than serving as an undefined fallback.
Immutable promotion
Moving the same identified artifact through environments instead of rebuilding a potentially different artifact at each stage.
Incident
An event that degrades service, decision quality, safety, security, compliance, or user outcomes and requires coordinated response.
Model registry
A controlled record of model candidates, versions, metadata, evaluation evidence, stages, and approvals.
Observability
The ability to investigate a system’s internal state from telemetry such as metrics, logs, traces, and domain events.
Rollback
Restoring a previously verified model, policy, configuration, or deployment after a change causes unacceptable behaviour.
Runbook
An operational procedure that links a known condition to diagnostic, mitigation, escalation, and verification steps.
Service-level indicator (SLI)
A measured property of service behaviour, such as the proportion of successful requests.
Service-level objective (SLO)
A target for an SLI over a defined period, based on the reliability users need.
Shadow deployment
Running a candidate on copied production inputs without allowing its output to control live decisions.
Telemetry
Operational evidence emitted by a system, including metrics, logs, traces, events, and model-specific signals.
Final reproducibility check
Before declaring the repository complete, run:
source .venv/bin/activate
python -m pip install -r requirements.txt
python -m pytest -q
bash scripts/bash/13-run-end-to-end-system.sh
quarto render
git status --shortReview the test output, the Chapter 13 readiness report and figure, the Quarto render log, and every changed file reported by Git. A successful run should leave the guide with passing tests, reproducible system evidence, a clean render, and only intentional repository changes.
This is the final operating lesson of the guide: production readiness is not a single test or deployment event. It is a repeatable chain of evidence connecting code, infrastructure, data, models, policies, people, and outcomes.