Continuous Integration and Delivery
Turning a Tested Model Service into a Safe, Repeatable Release
Continuous Integration and Delivery
Model deployment established that the case-study model can be packaged, started, and queried through an API. That is necessary, but it does not answer the next operational question:
What happens after a developer changes the service, its dependencies, or the model artifact?
Without an automated release process, every change depends on someone remembering the correct commands, running the right tests, building the correct image, and deploying the intended version. The process may work once and still be unreliable as a system.
Continuous integration and continuous delivery replace that fragile sequence with a versioned, repeatable pathway. Continuous integration (CI) provides rapid evidence that a proposed change is safe to merge. Continuous delivery (CD) produces a release candidate that is always ready for controlled promotion. Continuous deployment goes one step further and promotes every qualifying change to production automatically.
This chapter applies those ideas to the guide’s running case study: a containerized prediction API used to support a human decision. The goal is not merely to make deployment faster. It is to make every release identifiable, testable, recoverable, and auditable.
Learning objectives
By the end of this chapter, you should be able to:
- distinguish continuous integration, continuous delivery, and continuous deployment;
- design a layered validation pipeline for a model-backed API;
- build one immutable container image and promote the same artifact between environments;
- separate application, data, model, and infrastructure checks;
- define release gates, approvals, rollback conditions, and deployment evidence;
- implement a practical GitHub Actions workflow; and
- evaluate the trade-off between fast feedback and release confidence.
From a local command to a release system
A local workflow often looks like this:
pytest
docker build -t decision-api:latest .
docker run -p 8000:8000 decision-api:latestThese commands are useful during development, but they do not establish who ran them, which commit was tested, whether the deployed image is the tested image, or whether production remained healthy after release.
A delivery pipeline turns those commands into an evidence-producing system.
| Stage | Main question | Typical evidence |
|---|---|---|
| Source validation | Is the proposed change reviewable and reproducible? | Commit SHA, pull request, dependency files |
| Fast CI | Does the code meet basic quality and correctness expectations? | Lint, unit tests, schema tests |
| Integration CI | Do the service, model, and container work together? | API tests, model loading test, container smoke test |
| Artifact creation | What exactly could be released? | Image digest, version tag, build metadata |
| Staging | Does the release candidate work in a production-like environment? | Deployment status, smoke tests, compatibility checks |
| Production gate | Is the remaining risk acceptable? | Approval, policy checks, change record |
| Production verification | Did the release remain healthy after promotion? | Health, latency, error rate, prediction telemetry |
The pipeline is therefore more than a list of commands. It is a chain of claims supported by evidence.
CI: validate every proposed change
CI should begin on every pull request and on changes merged into the protected main branch. Its first responsibility is to fail quickly and clearly.
Order checks by feedback cost
Fast, deterministic checks belong near the beginning. Slower checks should run only after basic failures have been eliminated.
- Validate repository structure and configuration.
- Check formatting, linting, and imports.
- Run unit tests.
- Validate request and response schemas.
- Load the serialized model and make a known prediction.
- Start the complete API and run integration tests.
- Build the container image.
- Start the container and run smoke tests.
- Run security and dependency checks.
This ordering shortens the time between introducing a defect and receiving useful feedback. It does not mean that later checks are less important; it means expensive capacity is not spent on a change that already fails a basic test.
Test the boundaries of the system
A model service can fail at several distinct boundaries:
- code boundary: functions produce incorrect results;
- schema boundary: incoming data no longer matches the API contract;
- feature boundary: training and inference use different feature names or transformations;
- artifact boundary: the service cannot load the intended model version;
- service boundary: the API starts but returns incorrect status codes or response fields;
- container boundary: the application works locally but not inside its image; and
- environment boundary: the image works in CI but fails with staging configuration or infrastructure.
The pipeline should test each boundary explicitly. A single successful prediction cannot substitute for all of these tests.
Prefer deterministic CI fixtures
CI tests should use small, versioned fixtures with known expected behaviour. Tests that depend on a live operational database, an unversioned remote file, or the current state of a third-party API may fail for reasons unrelated to the proposed change.
For the case-study service, a minimal fixture set should include:
- one valid request that produces a successful prediction;
- missing and malformed fields that must be rejected;
- values at documented boundaries;
- a row with an expected class or score range; and
- a health request that confirms the expected model version is loaded.
Tests should verify the contract and safety properties, not reproduce the complete training dataset in CI.
Model-aware validation
Ordinary software tests are necessary for model systems, but they are not sufficient. A pipeline may contain perfectly valid Python code while shipping the wrong model, an incompatible preprocessor, or unacceptable predictive behaviour.
Separate four change types
The pipeline should make the source of a release visible.
| Change type | Example | Required checks |
|---|---|---|
| Application | New endpoint or response field | Unit, schema, integration, container tests |
| Dependency | Updated FastAPI or scikit-learn version | Full tests, vulnerability scan, model-load compatibility |
| Model | New fitted pipeline or threshold | Artifact integrity, evaluation gates, fairness or subgroup checks |
| Configuration | New timeout or resource limit | Configuration validation, staging test, runtime verification |
Different changes can share one delivery pathway while activating different policy gates.
Validate the model artifact
At minimum, CI should confirm that:
- the artifact exists and its checksum is recorded;
- the expected serialization format can be loaded;
- the required feature schema matches the service contract;
- a prediction returns the expected type and valid range;
- the model version is exposed by service metadata; and
- evaluation results meet the thresholds required for promotion.
A metric gate should be defined before the candidate is evaluated. For example, a new model may require recall of at least 0.80, no more than a specified reduction in precision, and acceptable performance for predefined subgroups. Selecting a threshold after observing the candidate makes the gate easy to manipulate.
Build once, promote the same artifact
One of the most important delivery rules is:
Build the release artifact once, identify it immutably, and promote that exact artifact.
Rebuilding separately for staging and production can produce different outputs even when both builds appear to use the same source. Dependencies, base images, or build-time inputs may change between builds.
The container image should therefore be tagged with a traceable identifier such as the Git commit SHA and recorded by digest:
decision-api:8f12c4a
sha256:4c8f...e921
Human-friendly tags such as v1.3.0 may point to the same image, but latest should not be the only release identity. A deployment record should connect:
- source commit;
- workflow run;
- container digest;
- model version and checksum;
- configuration version;
- target environment; and
- deployment time and outcome.
This lineage makes incident investigation and rollback much more reliable.
CD: controlled promotion through environments
The case-study pathway uses three logical environments.
| Environment | Purpose | Expected controls |
|---|---|---|
| CI runner | Validate source and build the candidate | Isolated job, synthetic fixtures, short-lived credentials |
| Staging | Test the candidate under production-like conditions | Separate secrets, smoke tests, representative configuration |
| Production | Serve real decision workflows | Protected environment, monitoring, rollback, audit record |
Configuration should vary between environments, but application code and the container image should not. Secrets must be supplied by the environment and never committed to the repository or baked into the image.
Continuous delivery or continuous deployment?
For a prediction system that influences human decisions, continuous delivery is often the better starting point. Automation creates and validates the release candidate, while an authorized person approves production promotion after reviewing the evidence.
Continuous deployment may become appropriate when:
- tests and monitoring are mature;
- changes are small and reversible;
- the service has reliable progressive delivery controls;
- policy permits automated promotion; and
- the team has demonstrated fast detection and recovery.
The choice is a risk decision, not a maturity badge.
Release gates
A gate blocks promotion when required evidence is missing or unacceptable. Useful gates for the case-study system include:
- all required CI jobs pass;
- code review is complete;
- the image is traceable and has passed a vulnerability policy;
- the model meets predefined evaluation thresholds;
- staging health and prediction smoke tests pass;
- database or schema compatibility is confirmed;
- an authorized reviewer approves production; and
- rollback information is available.
Avoid gates that exist only as manual ceremony. Every gate should name a risk, the evidence used to evaluate it, and the person or policy authorized to accept the remaining risk.
A practical GitHub Actions workflow
The downloadable workflow .github/workflows/04-ci-cd.yml demonstrates a vendor-neutral pattern using GitHub Actions syntax. It intentionally stops short of deploying to a specific cloud provider. The repository can therefore teach the workflow without embedding provider credentials or infrastructure assumptions.
The implemented example contains four dependent jobs:
- test installs dependencies and runs the API test suite;
- build creates a commit-tagged image only after the tests pass, exports it, and retains it as a workflow artifact;
- staging-smoke-test downloads and loads that exact image, starts a container, and checks its
/healthendpoint; and - production-gate records the validated commit as approved for a future provider-specific deployment.
The production job can be connected to a GitHub environment with required reviewers when manual approval is desired. In a real system, the build job would also authenticate to a registry, push the image tagged with the commit SHA, and pass its digest to provider-specific deployment jobs.
The repository includes the minimum working system needed to exercise the pipeline rather than merely display its syntax:
| File | Role in the pipeline |
|---|---|
app/main.py |
Defines the case-study API and /health endpoint |
tests/test_api.py |
Verifies the health contract through the API test client |
requirements.txt |
Declares the Python dependencies used by CI |
Dockerfile |
Packages the runnable service |
.github/workflows/04-ci-cd.yml |
Defines the four-job CI/CD pathway |
The complete downloadable workflow is the authoritative version. The excerpt below highlights the test and immutable-build transition:
name: Model service CI/CD
on:
pull_request:
push:
branches: [main]
workflow_dispatch:
permissions:
contents: read
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- name: Install dependencies
run: |
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
- name: Run tests
run: python -m pytest -q
build:
needs: test
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build immutable candidate
run: docker build --tag decision-api:${{ github.sha }} .
- name: Export candidate image
run: docker save decision-api:${{ github.sha }} | gzip > decision-api.tar.gz
- name: Retain candidate
uses: actions/upload-artifact@v4
with:
name: decision-api-${{ github.sha }}
path: decision-api.tar.gz
retention-days: 7Using python -m pytest is deliberate. It runs pytest through the Python interpreter configured by the workflow and keeps the repository root available when the test imports app.main. The job dependency needs: test ensures that a failing test prevents image construction.
Pin third-party actions to reviewed commit SHAs when applying stronger supply-chain controls. Major-version tags are kept here for readability and should be governed deliberately in production repositories.
What the successful run proves
The workflow was exercised through real GitHub Actions runs rather than treated as pseudocode. The first attempts failed for useful reasons: the runner initially lacked pytest, an empty tests/ directory produced no evidence, and direct pytest -q could not resolve the local app package in CI. Adding a meaningful health test, a runnable FastAPI service, a root-level Dockerfile, and the explicit python -m pytest -q invocation corrected the release pathway.
The successful push then completed the practical sequence:
source checkout
→ dependency installation
→ API test
→ commit-tagged Docker build
→ compressed image artifact upload
→ exact artifact download and load
→ container start and /health smoke test
→ production release-candidate record
This is CI/CD behavior in practice. A failed check stops downstream work; a passing change advances while retaining evidence tied to the Git commit. The workflow does not yet publish the image to a container registry or deploy it to a permanent cloud service. Its production gate prints an approval record for a later provider-specific deployment, so the current implementation is continuous delivery—not continuous deployment.
Reproduce the service checks locally
Before pushing a change, run the same essential checks from the repository root:
python -m pytest -q
docker build -t decision-api:local .
docker run --rm -d --name decision-api-local -p 8000:8000 decision-api:local
curl --fail http://localhost:8000/health
docker stop decision-api-localThe expected health response is:
{"status":"healthy","service":"decision-api","version":"0.1.0"}Local success reduces avoidable CI failures, but the remote run remains important because it verifies the repository in a fresh runner rather than the developer’s already-configured environment.
Simulating pipeline feedback and release risk
The chapter script scripts/python/04-simulate-ci-pipelines.py compares three illustrative designs:
- late integration: most checks run after a long build;
- layered CI: fast checks precede integration and container checks; and
- layered CI with parallel checks: independent checks run concurrently after the fast validation stage.
The simulation is not a production risk model. It shows how pipeline order and parallelism affect two operational quantities:
- feedback time: minutes until the pipeline reports the result; and
- escaped-change risk: the simulated proportion of changes passing the pipeline despite containing an undetected defect.
Run the experiment from the repository root:
bash scripts/bash/04-generate-ci-pipeline-results.shThe script writes a reproducible summary to results/04-ci-pipeline-summary.csv and the figure below to results/figures/04-ci-feedback-and-risk.png.
The expected pattern is that early deterministic checks reduce wasted work, while parallel execution reduces elapsed time without removing validation. The escaped-change result depends primarily on test coverage and detection probabilities, not on speed alone. A fast pipeline with weak tests is still weak.
Deployment verification
A successful deployment command does not prove a healthy release. Verification should occur at several levels:
Immediate smoke checks
- the service starts;
- the health endpoint returns success;
- the expected application and model versions are reported;
- one synthetic prediction completes; and
- critical dependencies are reachable.
Short observation window
- error rate remains within its release threshold;
- latency does not regress materially;
- resource saturation remains acceptable;
- request and response schemas remain valid; and
- prediction distributions do not show an obvious discontinuity.
Decision-workflow checks
- downstream consumers receive the response;
- humans can interpret the output as expected;
- fallbacks work when the model is unavailable; and
- audit records connect decisions to the deployed model version.
These checks connect delivery automation to the broader system rather than stopping at container health.
Rollback and roll-forward
Every release should have a recovery strategy before it reaches production.
Rollback restores the previous known-good artifact. It is appropriate when the prior version remains compatible with current data and infrastructure.
Roll-forward deploys a corrective release. It may be safer when a database migration, external contract, or irreversible data change prevents restoration of the old version.
For model systems, rollback must consider both application and model compatibility. Rolling back only the service while retaining an incompatible model artifact can create a second failure. The release record should identify a compatible bundle of application image, model version, preprocessing contract, and configuration.
Common failure patterns
Testing one artifact and deploying another
Cause: staging and production rebuild from source independently.
Control: promote the same image digest through all environments.
Long pipelines that developers bypass
Cause: every check runs sequentially or expensive tests run before fast checks.
Control: order checks by cost, parallelize independent jobs, and reserve scheduled pipelines for exhaustive tests that do not need to block every commit.
Green tests with a broken model contract
Cause: tests mock the model or verify only HTTP status codes.
Control: load the real candidate artifact and assert feature schema, model metadata, and prediction properties.
Secrets in source or images
Cause: credentials are stored in configuration files or passed as Docker build arguments.
Control: use environment-specific secret stores, short-lived credentials, restricted permissions, and secret scanning.
No reliable rollback target
Cause: mutable tags such as latest overwrite release identity.
Control: retain immutable digests and a deployment history that links each environment to an exact release.
Automatic deployment without automatic verification
Cause: the pipeline ends when the deployment command returns success.
Control: define post-deployment checks and automatically stop or reverse promotion when health thresholds fail.
Practical checklist
Before treating the case-study pipeline as delivery-ready, confirm that:
Chapter summary
Continuous integration gives rapid, repeatable evidence about a proposed change. Continuous delivery turns a validated change into an immutable release candidate and promotes it through controlled environments. For model systems, this pathway must test more than application code: it must preserve the contract among data, preprocessing, model artifacts, the API, and the surrounding decision workflow.
The central operating principle is simple: build once, identify the artifact immutably, promote the same artifact, verify the release, and retain a tested recovery path. The next chapter builds on this release foundation by examining deployment strategies and environment design in greater depth.
Exercises
- Classify each check in your current model service as a code, schema, feature, artifact, service, container, or environment check. Which boundary has no automated test?
- Modify the workflow so lint and unit tests run as separate parallel jobs. What information must a later job receive from them?
- Define three measurable promotion gates for a new case-study model. Include at least one operational and one predictive criterion.
- Explain why rebuilding the same commit separately for staging and production weakens release evidence.
- Run the simulation, change the check durations or detection probabilities, and explain which changes improve speed, confidence, or both.
- Write a rollback record containing the exact application image, model artifact, preprocessing contract, and configuration needed to restore a known-good release.