A production model is not a single file that is deployed once and forgotten. It is one version in a governed sequence of datasets, features, training runs, evaluation evidence, approvals, deployments, observations, and eventual retirement. Lifecycle management makes that sequence reproducible and reversible.
This chapter extends the monitoring and retraining workflow from the previous chapters. Monitoring can identify a change and retraining can produce a candidate, but neither answers the operational questions that follow:
Which data, code, parameters, and environment produced this candidate?
Which evaluation evidence justified its promotion?
Which version is serving in each environment?
Can the system restore the previous version without reconstructing it?
When may old artifacts and data be archived or deleted?
The goal is not to retain everything forever. The goal is to preserve enough evidence to reproduce important decisions, operate safely, and meet explicit retention obligations.
Learning objectives
By the end of this chapter, you should be able to:
distinguish model versioning from full lineage;
define lifecycle states for data and model artifacts;
design promotion gates that separate candidate creation from deployment approval;
connect a deployed model to immutable evidence;
plan rollback, retirement, retention, and deletion; and
interpret a practical lifecycle simulation and its inventory outputs.
From files to governed assets
Saving a model as model.joblib records bytes, but not meaning. A production asset needs an identity and relationships to the evidence that created it.
Asset
Minimum identity
Important relationships
Raw data snapshot
source, extraction time, content hash
consent, access policy, schema
Curated training set
dataset version and transformation version
raw inputs, quality checks, labels
Feature definition
code version and feature schema
training and online computation
Model candidate
model version and artifact digest
training data, code, parameters, metrics
Evaluation report
evaluation dataset and test version
acceptance thresholds, reviewer decision
Deployment
environment, release ID, time
approved model, configuration, image digest
Prediction record
request and response identifiers
deployed version, event time, feedback key
Names such as final_model_v2_revised.joblib encode history informally and are easy to overwrite. Immutable identifiers and append-only lifecycle events make history queryable.
Lineage is a graph
Model versioning answers, “Which model artifact is this?” Lineage answers, “How did it come to exist, why was it accepted, where did it run, and what depended on it?”
Code
flowchart TD D["Versioned data"] --> T["Training run"] C["Code and configuration"] --> T T --> M["Model candidate"] M --> E["Evaluation evidence"] E --> R["Approved release"] R --> P["Predictions and monitoring"]
flowchart TD
D["Versioned data"] --> T["Training run"]
C["Code and configuration"] --> T
T --> M["Model candidate"]
M --> E["Evaluation evidence"]
E --> R["Approved release"]
R --> P["Predictions and monitoring"]
For the continuous decision-system case study, a release record should link at least the following immutable values:
The manifest uses references rather than mutable paths such as latest. Human-readable aliases may be convenient, but an operational record should resolve them to immutable versions.
Lifecycle states
Lifecycle states make permitted actions explicit. The exact labels may vary, but their semantics must be stable.
A registry state is not merely descriptive. It should control what automation may do. For example, a candidate may be evaluated automatically, but only an approved model may be promoted to production.
Separate creation, approval, and deployment
Retraining automation should be able to create and evaluate candidates without automatically changing production. A safe promotion workflow has distinct gates.
Gate
Example evidence
Failure action
Integrity
hashes, signatures, readable artifacts
quarantine candidate
Data
schema, missingness, leakage, population checks
stop evaluation
Model
discrimination, calibration, subgroup results
reject candidate
System
contract, load, security, dependency tests
block release
Operational
staging smoke test and observability checks
retain current version
Approval
named reviewer or codified low-risk policy
await decision
The incumbent remains the default when a candidate fails. “No deployment” is a valid and often desirable outcome of retraining.
Immutability, aliases, and reproducibility
Three concepts work together:
Immutable objects preserve exact artifacts and datasets.
Metadata records describe identity, lineage, status, and policy.
Aliases such as champion or production point to immutable objects and may move through controlled operations.
Reproducibility requires more than a random seed. Record the source code commit, dependency lock or container digest, parameters, data snapshot, feature definitions, evaluation protocol, and execution context. Some hardware and library operations are nondeterministic, so the practical standard is often traceable and repeatable within declared tolerances, not necessarily bit-for-bit identical results.
Rollback is a lifecycle operation
A rollback should redeploy a previously approved release, not rebuild an old model from remembered settings. The rollback target therefore needs:
a retained model and container artifact;
compatible feature and API contracts;
required configuration and secrets references;
a known deployment procedure;
evidence that the version remains permitted to serve; and
a recorded reason, initiator, and outcome.
Rollback also has data consequences. If a newer release changes event schemas or feature definitions, restoring only the model may not restore a compatible system. Compatibility tests must cover the complete serving path.
Retention and deletion
Keeping every artifact indefinitely increases cost, security exposure, and privacy risk. Deleting too aggressively destroys reproducibility and rollback capability. Retention should therefore be policy-driven.
A defensible retention record specifies:
asset class and owner;
reason for retention;
minimum and maximum retention periods;
legal, contractual, research, or operational holds;
archive location and access controls;
deletion authority and method; and
evidence that deletion completed.
Derived data does not automatically become harmless. Features, embeddings, logs, and prediction records may still contain sensitive or linkable information. Their lifecycle must be defined explicitly rather than inherited casually from raw data.
Practical lifecycle simulation
The companion program creates a deterministic inventory of six model versions. It applies evaluation gates, promotes eligible versions, records the period for which each version served, simulates a rollback, and assigns retention actions.
Run it from the repository root:
bash scripts/bash/10-simulate-lifecycle.sh
The program writes:
results/10-model-lifecycle-inventory.csv, a machine-readable registry inventory;
results/10-model-lifecycle-events.csv, an append-only event log; and
results/figures/10-model-lifecycle-timeline.png, a lifecycle timeline.
Model versions move through evaluation, production, rollback, and retention states.
The simulation is deliberately small. In a real registry, state changes should be transactional, authenticated, authorized, and written to durable audit storage.
Reading the inventory
The output distinguishes three questions that are often collapsed:
Was the model accepted? See gate_result.
Did it serve production traffic? See production_start and production_end.
What should happen to it now? See retention_action.
A rejected candidate can be retained temporarily for debugging without being deployable. A superseded model may remain rollback-eligible for a defined period. A version involved in an incident may be placed on hold even when the normal policy would archive it.
Operational controls
Lifecycle metadata is useful only when connected to controls. A production implementation should enforce these invariants:
every deployment resolves to an immutable model and container digest;
every production model has passed the required gates;
a registry transition records actor, time, reason, and evidence;
production aliases cannot be changed by the training job alone;
rollback targets are tested for compatibility;
deletion respects active holds and produces an auditable record; and
monitoring events include the served model and release identifiers.
These controls connect CI/CD, observability, monitoring, retraining, and governance. Lifecycle management is therefore not a separate storage concern; it is the system of record joining the operating model system together.
Common failure modes
Treating object storage as a registry
A folder of artifacts provides storage but not controlled states, lineage, approval, or audit history.
Moving latest without preserving its target
An alias is useful for discovery, but operational records must capture the immutable version it referenced at the time of an event.
Reusing a version identifier
Replacing the bytes behind an existing version breaks historical evidence. Corrected artifacts receive new identifiers.
Automatically promoting every retrained model
Successful execution does not establish safety or value. Candidate creation and production promotion remain separate decisions.
Deleting inputs while retaining only the model
The artifact alone cannot explain training provenance, evaluation validity, or appropriate use.
Retaining everything forever
Unlimited retention creates cost and risk. Archive and deletion are normal, governed lifecycle states.
Chapter checklist
Before calling a lifecycle design operational, confirm that:
Key takeaway
Lifecycle management turns models and datasets into governed system assets. Immutable versions, connected lineage, controlled promotion, tested rollback, and explicit retention policies allow the organization to answer not only what is running now, but how it arrived, whether it should remain, and how it can be safely replaced or retired.