Deployment Strategies and Environments
The CI/CD workflow in the previous chapter produces an immutable, tested container image. That image is a release candidate, not yet a production change. This chapter explains how to promote the same candidate through controlled environments, expose it to traffic, evaluate operational evidence, and recover when the evidence is unacceptable.
The central principle is simple:
Build once, promote the same artifact, and increase exposure only when evidence supports the next step.
By the end of the chapter, you will be able to:
- distinguish environments from deployment strategies;
- design a promotion path from development to production;
- compare recreate, rolling, blue–green, and canary deployments;
- define release gates using service and model-system signals;
- connect rollout decisions to rollback and data compatibility; and
- run a small simulation that shows how guarded traffic shifting limits impact.
From a candidate image to a running release
A model service release changes more than a Python function. It can change the container image, API contract, preprocessing logic, model artifact, configuration, dependencies, resource requirements, or observability fields. A safe release process treats these elements as one versioned system change.
Suppose Chapter 04 created this image:
decision-api:<git-sha>
The Git commit SHA makes the image traceable to source. Promotion should retain that identity:
build: decision-api:a81f4c2
staging: decision-api:a81f4c2
production: decision-api:a81f4c2
Do not rebuild an image separately for production. A rebuild can resolve newer dependencies, copy different files, or use a changed base image. The production artifact would then differ from the one tested in staging.
Runtime configuration may differ by environment, but the application image should not. Examples of environment-specific configuration include log level, service endpoints, replica count, credentials, and feature flags. Configuration belongs outside the image; secrets belong in a secret-management mechanism rather than source code or the container image.
Environments are evidence boundaries
An environment is not merely another server. It is a boundary with a specific purpose, access policy, configuration, and standard of evidence.
| Environment | Primary purpose | Typical data | Required evidence |
|---|---|---|---|
| Local development | Fast implementation and debugging | Synthetic or small safe samples | Unit tests and local behavior |
| CI | Repeatable automated verification | Fixtures and generated test data | Tests, build success, static checks |
| Staging | Production-like integration validation | Synthetic, de-identified, or controlled test data | Health, contract, dependency, and smoke tests |
| Production | Real decisions and real users | Governed operational data | Release gates, live telemetry, auditability, rollback readiness |
The environments should be similar enough that staging is informative, but they do not need identical scale. The most important forms of parity are:
- the same container image and startup command;
- the same API and model-loading pathway;
- equivalent dependency types and interfaces;
- the same configuration keys, with different values where necessary; and
- the same health checks, metrics, logs, and tracing conventions.
Perfect parity is rarely achievable. Production may have more replicas, stricter network controls, real traffic patterns, larger datasets, and downstream consumers that cannot be reproduced exactly. Record these differences explicitly because each difference limits what staging can prove.
A practical promotion path
For the decision API, a reasonable path is:
- Commit: tests validate the code and model-service contract.
- Build: CI creates one immutable image tagged with the commit SHA.
- Staging: the image starts with staging configuration and passes smoke and integration tests.
- Production candidate: the exact image is registered for release approval.
- Limited exposure: a deployment strategy sends a controlled share of traffic to the candidate.
- Promotion or rollback: observed evidence determines the next action.
Promotion is therefore an evidence-based state transition, not a file-copy operation.
Separate deployment from release
These terms are often used interchangeably, but the distinction is useful:
- Deployment places a version in an environment and makes it capable of serving.
- Release exposes that version’s behavior to users or downstream systems.
A candidate can be deployed but receive no production traffic. Blue–green deployments do this by preparing the new version beside the current version. Feature flags can also separate deployment from release by keeping a new pathway disabled until a later decision.
This separation creates a verification window. The team can check startup, model loading, dependencies, health endpoints, and observability before users depend on the new version.
Deployment strategies
No strategy is universally safest. The correct choice depends on traffic volume, infrastructure, compatibility, rollback time, cost, and the consequences of a faulty prediction.
Recreate
A recreate deployment stops the old version and then starts the new version.
Advantages: simple, inexpensive, and unambiguous—only one version is active.
Limitations: causes downtime and makes rollback slower because the old version must be restarted.
Use recreate for development, low-criticality internal services, or scheduled maintenance where downtime is accepted. It is usually unsuitable for a continuously available decision API.
Rolling deployment
A rolling deployment replaces old replicas gradually. During the rollout, both versions serve traffic.
Advantages: avoids planned downtime and uses existing capacity efficiently.
Limitations: mixed versions can complicate debugging, and requests may encounter incompatible API or state assumptions.
Rolling deployment works well when versions are backward compatible and replica-level health checks are trustworthy. Readiness probes must prevent a replica from receiving traffic before the model and all required resources are loaded.
Blue–green deployment
Blue–green deployment maintains two complete production-capable environments. The current version serves traffic in one environment; the candidate is prepared in the other. Traffic switches only after verification.
Advantages: fast traffic switching, strong pre-release verification, and fast rollback while the old environment remains intact.
Limitations: requires duplicate capacity and does not by itself reveal how the candidate behaves under real production traffic before the switch.
Blue–green is attractive when rollback speed is critical and the additional infrastructure cost is acceptable.
Canary deployment
A canary deployment sends a small proportion of live traffic to the candidate, evaluates evidence, and increases exposure in stages.
Advantages: limits the initial blast radius and reveals behavior under real traffic.
Limitations: requires reliable traffic routing, version-aware telemetry, sufficient traffic at each stage, and explicit promotion rules.
A canary is not simply “deploy to 10% and watch.” It is a controlled experiment with:
- defined traffic stages;
- a minimum observation window or sample size;
- candidate and baseline measurements;
- promotion and rollback thresholds; and
- an accountable decision maker or automated policy.
Strategy comparison
| Strategy | Downtime | Extra capacity | Live exposure before full release | Rollback speed | Main risk |
|---|---|---|---|---|---|
| Recreate | Expected | Low | None | Slow–medium | Service interruption |
| Rolling | Usually none | Low–medium | Gradual by replica | Medium | Mixed-version behavior |
| Blue–green | None at switch | High | Optional | Fast | Untested full traffic switch |
| Canary | None | Medium | Explicit and staged | Fast when automated | Weak or noisy gates |
For the decision API, canary deployment is a strong default when production traffic is high enough to provide useful evidence. Blue–green is a good alternative when traffic is sparse, version mixing is dangerous, or instant rollback is the primary concern.
Release gates for a model system
Application health is necessary but incomplete. A model endpoint can return HTTP 200 while serving the wrong model, applying incorrect preprocessing, producing implausible scores, or changing the distribution of decisions.
A release gate should combine several signal classes.
Service signals
- availability and readiness;
- request error rate;
- latency percentiles, especially p95 or p99;
- saturation, CPU, memory, and restart rate; and
- dependency failures and timeouts.
Model-system signals
- loaded model and schema versions;
- prediction score distribution;
- positive-decision or action rate;
- missing, invalid, or out-of-range input rates;
- divergence from the current production version; and
- slice-level behavior for operationally important groups.
Business and safety signals
- abandoned or failed user workflows;
- manual-review volume;
- downstream rejection or override rate;
- policy constraint violations; and
- immediate outcome proxies, when valid.
Ground-truth model performance is often delayed. A canary decision cannot always wait for labels that arrive weeks later. Immediate release gates should therefore detect operational and behavioral regressions without pretending that proxy signals prove long-term model quality. Delayed outcome monitoring continues after the release and may trigger a later rollback or model-lifecycle response.
Define a promotion policy before deployment
Consider a staged canary with candidate traffic shares of 5%, 25%, 50%, and 100%. A simplified policy could require:
canary:
stages: [0.05, 0.25, 0.50, 1.00]
minimum_requests_per_stage: 1000
promote_when:
error_rate_delta_max: 0.002
p95_latency_ratio_max: 1.20
decision_rate_delta_max: 0.03
rollback_when:
error_rate: 0.02
p95_latency_ms: 500
schema_or_model_mismatch: trueThese numbers are illustrative, not universal defaults. Thresholds should reflect the baseline’s normal variation, user impact, traffic volume, and risk tolerance. A low-volume service may need longer observation windows, while a high-risk decision system may require human approval even when automated gates pass.
The policy should also specify what happens when evidence is inconclusive. Pausing a rollout is different from promoting or rolling back. A pause preserves the current exposure while more evidence accumulates or an operator investigates.
Simulating traffic-shift risk
The companion program creates a deterministic release simulation. The baseline has a 0.5% error rate. The candidate begins at the same rate, then degrades to 4% after minute 12. Three strategies are compared:
- immediate: send all traffic to the candidate;
- linear: increase candidate traffic mechanically over 30 minutes; and
- guarded canary: increase traffic in stages, but roll back when the candidate’s rolling error rate crosses a threshold.
Run the Python program directly:
python scripts/python/05-simulate-deployment-strategies.pyOr use the Bash wrapper:
bash scripts/bash/05-simulate-deployment-strategies.shThe program writes:
results/05-deployment-strategy-simulation.csv
results/figures/05-deployment-strategy-comparison.png
The guarded canary in Figure 6.1 does not prevent all failures. It limits exposure while collecting evidence, detects the regression, and stops further impact. That is the operational value of progressive delivery: reducing the blast radius while preserving a path to learn from production conditions.
The simulation is deliberately transparent rather than statistically exhaustive. In a real deployment, the gate should account for sample size, uncertainty, repeated checks, seasonality, traffic composition, and multiple metrics. Very small canaries can be safer in terms of exposure yet too underpowered to distinguish a regression from random variation.
Rollback is a designed capability
“We can redeploy the old image” is not a complete rollback plan. A usable rollback design specifies:
- the last known-good image identifier;
- who or what can initiate rollback;
- the signal and threshold that trigger it;
- the expected time to restore traffic;
- whether configuration and feature flags must also revert;
- how in-flight requests are handled; and
- how the event is recorded and communicated.
Rollback must be rehearsed. If the process has never been tested, its recovery time is an assumption.
Compatibility can make rollback impossible
A new release may write data that the previous version cannot read. Database migrations, event schemas, cached objects, and feature representations can all create this trap. Prefer backward-compatible, staged changes:
- expand the schema or interface so both versions work;
- deploy code that tolerates old and new representations;
- migrate or backfill data;
- switch consumers to the new representation; and
- remove the old representation in a later release.
This expand–migrate–contract pattern protects both rolling coexistence and rollback. Destructive migrations should not be coupled to the first deployment of code that depends on them.
Version-aware observability
During a progressive rollout, aggregate service metrics can hide a candidate failure because most traffic still reaches the stable version. Every request should be attributable to release identity. Useful dimensions include:
service_version
git_sha
model_version
schema_version
environment
deployment_id
Keep label values bounded. Do not place request IDs, user IDs, or other high-cardinality values in metric labels. Those belong in logs or traces with appropriate privacy controls.
Candidate and baseline dashboards should use the same definitions and time windows. A release decision is unreliable when one version’s error rate excludes timeouts or one latency series measures a different boundary.
Deployment runbook for the decision API
Before production deployment:
- identify the immutable image and last known-good image;
- verify staging smoke, contract, and dependency tests;
- confirm configuration and secret references;
- review API, schema, model, and data compatibility;
- select the strategy, stages, owners, and observation windows;
- define promotion, pause, and rollback rules; and
- confirm dashboards, alerts, and version dimensions.
During rollout:
- verify readiness before routing traffic;
- compare candidate signals with the stable version;
- annotate the deployment timeline;
- record each stage transition and its evidence; and
- stop automatically or manually when a rollback condition is met.
After rollout:
- verify full-traffic behavior and downstream workflows;
- continue delayed outcome and drift monitoring;
- retain the previous version for the agreed recovery window;
- record the release result and any exceptions; and
- convert unexpected behavior into tests, policies, or runbook improvements.
Common failure patterns
Rebuilding for each environment
This breaks artifact identity. Build once and promote the same digest or immutable tag.
Treating a health endpoint as release proof
Health checks show that a process can respond, not that its predictions, dependencies, or decisions are correct.
Canary traffic without version-separated telemetry
Aggregate metrics dilute candidate failures. Attach release identity to operational evidence.
Promotion by elapsed time alone
Waiting ten minutes is not evidence unless the window contains enough representative traffic and the defined gates pass.
Automatic rollback without compatibility analysis
Routing traffic back cannot repair data already written in an incompatible format. Design reversible state changes before release.
Environment-specific code branches
Conditional application code such as if production creates pathways that staging never tests. Prefer external configuration with validation at startup.
Chapter summary
A tested container becomes a production release through controlled promotion, not by merely starting it on a production host. Environments establish distinct evidence boundaries. Deployment strategies control how the candidate becomes capable of serving and how quickly it receives traffic. Recreate, rolling, blue–green, and canary strategies make different trade-offs among downtime, cost, version coexistence, evidence, and rollback speed.
For model systems, release gates must extend beyond availability and latency. They should include model identity, input validity, prediction behavior, decision effects, and relevant safety constraints. Progressive delivery limits impact only when telemetry is version-aware, thresholds are defined in advance, and rollback is both technically possible and operationally rehearsed.
The next chapter builds on this release process by making the system observable in normal operation and actionable during incidents.
Exercises
- Classify each check in your current workflow as local, CI, staging, limited-production, or full-production evidence. Identify one claim being made at the wrong boundary.
- Choose a deployment strategy for a low-traffic but high-consequence model service. Explain why a canary may or may not provide enough evidence.
- Add a second rollback rule to the simulation—for example, cumulative excess failures or a latency threshold—and compare the resulting impact.
- Design an expand–migrate–contract sequence for adding a required response field without breaking the stable version or its consumers.
- Write a release record containing the image SHA, model version, configuration version, deployment stages, evidence, decision maker, and final outcome.