Feedback Loops and Retraining

Published

Aug 2026

Feedback Loops and Retraining

A deployed model does not improve merely because it receives more traffic. Improvement requires a controlled feedback loop that connects predictions to later outcomes, evaluates whether new evidence justifies a change, and promotes a replacement only when it is safer and more useful than the current model.

This chapter extends monitoring into action. The central case study remains the decision service introduced earlier in the guide. Its predictions are made immediately, but verified outcomes arrive later. Those delayed labels make retraining possible—and also create risks involving leakage, bias, unstable data, and unsafe automation.

Learning objectives

By the end of this chapter, you should be able to:

  • distinguish monitoring signals from retraining evidence;
  • design a traceable prediction-to-outcome feedback dataset;
  • define retraining triggers, eligibility rules, and promotion gates;
  • compare a candidate model with the production champion on the same holdout data;
  • separate automated retraining from automated deployment; and
  • document rollback, approval, and post-release monitoring requirements.

From prediction to verified outcome

At prediction time, a system usually knows the input features, model version, score, recommendation, and request context. The true outcome may not be known until days or weeks later. A durable feedback record must join these two moments without silently rewriting history.

Field group Example fields Purpose
Prediction identity prediction_id, entity_id, event_time Links the decision to later evidence
Model lineage model_version, feature_version, threshold Reconstructs what produced the decision
Model output score, predicted_class Supports performance and calibration analysis
Decision context action, reviewer_override, policy_version Separates model output from the action taken
Outcome outcome, outcome_time, label_source Supplies verified ground truth
Quality controls label_status, exclusion_reason Prevents provisional or invalid labels entering training

The join key must be stable, unique, and privacy-aware. Features available only after the decision must never be added to the historical training row. That would introduce target leakage and make offline performance impossible to reproduce in production.

A feedback loop is a governed system

The loop contains more than a scheduled training job:

  1. Capture predictions, versions, and decision context.
  2. Observe outcomes after an appropriate maturation delay.
  3. Validate label completeness, provenance, and eligibility.
  4. Evaluate performance by time period and important subgroups.
  5. Trigger retraining only when an approved rule is satisfied.
  6. Train a candidate from versioned data and code.
  7. Compare the candidate with the current champion on a fixed holdout set.
  8. Approve and release through the deployment controls established earlier.
  9. Monitor the new version and roll back if guardrails fail.

Monitoring and retraining therefore have different responsibilities. Monitoring asks whether the current system remains healthy. Retraining asks whether sufficient trustworthy evidence exists to build a demonstrably better replacement.

When should retraining start?

A calendar can initiate an evaluation, but it should not automatically justify a new model. Useful trigger families include:

  • Schedule-based: evaluate monthly or quarterly when labels mature slowly.
  • Volume-based: evaluate after enough new eligible outcomes accumulate.
  • Performance-based: investigate when discrimination, calibration, or error rates breach a sustained threshold.
  • Drift-based: evaluate when feature or score distributions move materially.
  • Event-based: evaluate after a policy, product, population, or measurement change.

No single trigger proves retraining is the correct response. A performance decline may come from a broken upstream pipeline, delayed labels, policy changes, or a serving defect. Retraining on corrupted evidence can reinforce the failure.

Trigger policy

A practical trigger policy combines evidence:

\[ T = (N_{eligible} \geq N_{min}) \land (D_{labels} \geq D_{min}) \land (B_{performance} \lor B_{drift} \lor B_{schedule}), \]

where \(N_{eligible}\) is the number of eligible labelled observations, \(D_{labels}\) is label completeness, and each \(B\) term represents an approved trigger condition. The first two conditions prevent weak or immature feedback from starting a retraining run.

Example: simulate a controlled retraining decision

The companion program creates an earlier training period, a recent labelled period, and a future-style holdout period. It trains a production champion, detects deterioration on recent outcomes, trains a candidate, and applies promotion gates on a shared holdout set.

Run it from the repository root:

bash scripts/bash/09-run-feedback-loop.sh

The program writes:

  • results/09-retraining-evaluation.csv — champion and candidate metrics;
  • results/09-retraining-decision.json — trigger evidence, gates, and final decision; and
  • results/figures/09-feedback-loop-evaluation.png — performance and promotion summary.

The simulation is intentionally reproducible. In a real system, the same structure would read versioned snapshots from controlled storage and register the candidate, data lineage, code revision, and evaluation report.

Champion–candidate evaluation

The production model is the champion. A retrained model is only a candidate until it passes every required gate. Both must be evaluated on the same untouched, time-appropriate holdout data.

Useful gates include:

Gate Example rule Why it matters
Data readiness At least 800 eligible labels and 90% completeness Avoids learning from weak feedback
Discrimination Candidate ROC AUC is no worse than champion Protects ranking performance
Probability quality Candidate log loss improves by at least 1% Rewards better probabilities
Operational fit Artifact size and inference latency remain within limits Prevents an accurate but unusable release
Subgroup safety No protected or operationally important group regresses beyond tolerance Prevents aggregate gains hiding local harm
Reproducibility Data, code, configuration, and environment are versioned Makes the result auditable

Promotion should use conjunctive logic for mandatory gates:

\[ Promote = G_{data} \land G_{quality} \land G_{operations} \land G_{subgroups} \land G_{reproducibility}. \]

A large gain in one metric must not compensate for a failed safety or reproducibility gate.

Why the holdout period comes after training

Random splitting can allow near-duplicate events, repeated entities, or future operating conditions to influence training. A time-based evaluation better represents the deployment question: would the candidate trained on information available then perform on observations arriving next?

The correct split depends on the system:

  • use entity-aware splitting when one person, device, or account produces multiple rows;
  • allow outcomes to mature before declaring labels complete;
  • preserve seasonality when choosing evaluation windows; and
  • freeze the promotion holdout before candidate development begins.

Repeatedly tuning against the promotion holdout turns it into training data. When extensive iteration is required, retain a final release set that remains untouched until the candidate is locked.

Feedback can change the world being measured

Predictions often influence actions, and actions influence observed outcomes. For example, a high-risk prediction may trigger an intervention that prevents the adverse outcome. Naively recording the improved outcome as evidence that the original prediction was wrong creates a selective-label problem.

The feedback dataset should therefore preserve:

  • the prediction before intervention;
  • the action actually taken;
  • overrides and their reasons;
  • whether an outcome was observable; and
  • the policy that determined treatment or review.

This does not automatically identify causal effects, but it prevents the model output, organizational decision, and observed outcome from being collapsed into one misleading label.

Automation boundaries

Retraining automation is valuable for repeatability, not for removing accountability. A safe pipeline may automatically validate data, train a candidate, calculate metrics, and assemble evidence. Deployment can remain subject to explicit approval when errors affect people, finances, safety, rights, or regulated decisions.

At minimum, a retraining run should record:

  • trigger reason and run identifier;
  • training and evaluation window boundaries;
  • dataset and feature versions;
  • code revision, dependencies, and configuration;
  • champion and candidate model versions;
  • global and subgroup metrics;
  • every promotion gate and its result;
  • approver, approval time, and release decision; and
  • rollback target and post-release guardrails.

Release and rollback

A passed offline evaluation permits controlled release; it does not guarantee production success. The candidate should use an appropriate strategy such as shadow evaluation, canary release, or a limited cohort. During the observation window, compare technical and model-level indicators with the champion baseline.

Rollback conditions should be defined before release. Examples include elevated error rate, latency beyond the service objective, unexpected score distribution, deteriorating calibration, or a critical subgroup guardrail breach. Rollback restores a known model version, while the associated feature and policy versions must remain compatible.

Common failure modes

Retraining on every new batch

Frequent retraining increases noise, cost, and change risk. Start evaluation only after label maturity, minimum volume, and an approved trigger.

Using provisional labels as truth

Operational states such as “case opened” may later change. Preserve label status and train only from finalized outcomes unless the modelling design explicitly handles censoring or revisions.

Promoting on one aggregate metric

An average improvement can hide worse calibration, subgroup harm, or operational regression. Use a compact set of mandatory gates.

Letting the candidate choose its own test set

Candidate-specific evaluation data makes comparison unreliable. Freeze a shared holdout and evaluate both models under identical conditions.

Automatically deploying because training succeeded

A completed job proves only that the pipeline ran. Promotion requires evaluation evidence; deployment requires the release controls appropriate to the system’s risk.

Operational checklist

Before enabling a feedback-and-retraining workflow, confirm that:

  • prediction records can be joined to verified outcomes;
  • outcome maturity and label provenance are explicit;
  • leakage and selective-label risks have been reviewed;
  • retraining triggers and minimum evidence thresholds are documented;
  • the champion and candidate share a frozen evaluation set;
  • global, calibration, subgroup, and operational gates are enforced;
  • every artifact has data, feature, code, and configuration lineage;
  • deployment approval is separate from pipeline completion; and
  • rollback and post-release observation rules are tested.

Key takeaways

  • Feedback becomes useful only after predictions are linked to mature, trustworthy outcomes.
  • Monitoring can trigger investigation, but drift alone does not prove that retraining is appropriate.
  • A candidate replaces the champion only after passing all mandatory promotion gates on shared holdout data.
  • Human actions and policy changes must remain visible because predictions can alter the outcomes later used as labels.
  • Automated retraining and automated deployment are separate decisions.
  • Reproducibility, approval, controlled release, and rollback turn model updating into a reliable system capability.

Exercises

  1. Change the minimum label volume in 09-run-feedback-loop.py. Explain how the trigger decision changes.
  2. Add a candidate latency metric and a mandatory operational gate.
  3. Introduce a group column, calculate group-specific log loss, and block promotion when any group regresses by more than 5%.
  4. Replace the simulated data with a versioned snapshot while preserving the output contracts.
  5. Draft a release observation window and rollback policy for a high-impact decision service.