LeetLLM
My PlanLearnGlossaryTracksPracticeBlog
LeetLLM

Your go-to resource for mastering AI & LLM systems.

Product

  • Learn
  • Glossary
  • Tracks
  • Practice
  • Blog
  • RSS

Legal

  • Terms of Service
  • Privacy Policy

© 2026 LeetLLM. All rights reserved.

All Topics
Your Progress
0%

0 of 177 articles completed

🛠️Computing Foundations0/9
Git, Shell, Linux for AIDocker for Reproducible AIPython for AI EngineeringNumPy and Tensor ShapesCUDA for ML TrainingMPS & Metal for ML on MacData Structures for AISQL and Data ModelingAlgorithms for ML Engineers
📊Math & Statistics0/8
Gradients and BackpropVectors, Matrices & TensorsLinear Algebra for MLAdam, Momentum, SchedulersProbability for Machine LearningStatistics and UncertaintyDistributions and SamplingHypothesis Tests, Intervals, and pass@k
📚Preparation & Prerequisites0/13
Neural Networks from ScratchCNNs from ScratchTraining & BackpropagationSoftmax, Cross-Entropy & OptimizationRNNs, LSTMs, GRUs, and Sequence ModelingAutoencoders and VAEsThe Transformer Architecture End-to-EndLanguage Modeling & Next TokensFrom GPT to Modern LLMsPrompt Engineering FundamentalsCalling LLM APIs in ProductionFirst AI App End-to-EndThe LLM Lifecycle
🧮ML Algorithms & Evaluation0/11
Linear Regression from ScratchLogistic Regression and MetricsDecision Trees, Forests, and BoostingReinforcement Learning BasicsValidation and LeakageClustering and PCACore Retrieval AlgorithmsDecoding AlgorithmsExperiment Design and A/B TestingPyTorch Training LoopsDataset Pipelines and Data Quality
📦Production ML Systems0/6
Feature Engineering for Production MLBatch and Streaming Feature PipelinesGradient Boosted Trees in ProductionRanking and Recommendation SystemsForecasting and Anomaly DetectionMonitoring Predictive Models
🧪Core LLM Foundations0/8
The Bitter Lesson & ComputeBPE, WordPiece, and SentencePieceStatic to Contextual EmbeddingsPerplexity & Model EvaluationFile Ingestion for AIChunking StrategiesLLM Benchmarks & LimitationsInstruction Tuning & Chat Templates
🧰Applied LLM Engineering0/24
Dimensionality Reduction for EmbeddingsCoT, ToT & Self-Consistency PromptingFunction Calling & Tool UseMCP & Tool Protocol StandardsContext EngineeringPrompt Injection DefenseResponsible AI GovernanceData Labeling and Human FeedbackEvaluating AI AgentsProduction RAG PipelinesHybrid Search: Dense + SparseReranking and Cross-Encoders for RAGRAG Evaluation for Reliable AnswersLLM-as-a-Judge EvaluationBias & Fairness in LLMsHallucination Detection & MitigationLLM Observability & MonitoringExperiment Tracking with MLflow and W&BPrompt Optimization with DSPyModel Versioning & DeploymentSemantic Caching & Cost OptimizationLLM Cost Engineering & Token EconomicsModel Gateways, Routing, and FallbacksDesign an Automated Support Agent
🎓Portfolio Capstones0/9
Capstone: Delivery ETA PredictionCapstone: Product RankingCapstone: Demand ForecastingCapstone: Image Damage ClassifierCapstone: Production ML PipelineCapstone: Document QACapstone: Eval DashboardCapstone: Fine-Tuned ClassifierCapstone: Reproducible ML Study
🧠Transformer Deep Dives0/8
Sentence Embeddings & Contrastive LossEmbedding Similarity & QuantizationScaled Dot-Product AttentionVision Transformers and Image EncodersPositional Encoding: RoPE & ALiBiLayer Normalization: Pre-LN vs Post-LNMechanistic InterpretabilityDecoding Strategies: Greedy to Nucleus
🧬Advanced Training & Adaptation0/16
Scaling Laws & Compute-Optimal TrainingPre-training Data at ScaleBuild GPT from Scratch LabJAX for PyTorch ResearchersContinued Pretraining for Domain ShiftSynthetic Data PipelinesSupervised Fine-Tuning PipelineMixed Precision TrainingDistributed Training: FSDP & ZeROLoRA & Parameter-Efficient TuningReward Modeling from Preference DataRLHF & DPO AlignmentConstitutional AI & Red TeamingRLVR & Verifiable RewardsKnowledge Distillation for LLMsModel Merging and Weight Interpolation
🤖Advanced Agents & Retrieval0/16
Vector DB Internals: HNSW & IVFAdvanced RAG: HyDE & Self-RAGGraphRAG & Knowledge GraphsRAG Security & Access ControlStructured Output GenerationReAct & Plan-and-ExecuteGuardrails & Safety FiltersCode Generation & SandboxingComputer-Use / GUI / Browser AgentsHuman-in-the-Loop Agent ArchitectureAI Coding Workflow with AgentsAgent Memory & PersistenceAgent Failure & RecoveryRecursive Language Models (RLM)Multi-Agent OrchestrationCapstone: Production Agent
⚡Inference & Production Scale0/19
Inference: TTFT, TPS & KV CacheMulti-Query & Grouped-Query AttentionKV Cache & PagedAttentionPrefix Caching and Prompt CachingFlashAttention & Memory EfficiencyContinuous Batching & SchedulingScaling LLM InferenceModel Parallelism for LLM InferenceModel Quantization: GPTQ, AWQ & GGUFLocal LLM DeploymentSLM Specialization & Edge DeploymentSpeculative DecodingLong Context Window ManagementMixture of Experts ArchitectureMamba & State Space ModelsReasoning & Test-Time ComputeAdvanced MLOps & DevOps for AIGPU Serving & AutoscalingA/B Testing for LLMs
🏗️System Design Capstones0/9
Content Moderation SystemCode Completion SystemMulti-Tenant LLM PlatformLLM-Powered Search EngineVision-Language Models & CLIPMultimodal LLM ArchitectureDiffusion Models: Images & TextReal-Time Voice AI AgentReasoning & Test-Time Compute
🎤AI Lab Interviewing0/4
AI Lab Coding Interview: Python SystemsAI Lab System Design InterviewAI Lab Behavioral InterviewAI Lab Technical Presentation
🔬Project Deep Dives0/17
Deep Dive - vLLMDeep Dive - SkyRLDeep Dive - FlashAttentionDeep Dive - FlashInferDeep Dive - DeepGEMMDeep Dive - NCCLDeep Dive - MegatronDeep Dive - DeepSpeedDeep Dive - RayDeep Dive - MLflowDeep Dive - PyTorchDeep Dive - TransformersDeep Dive - SGLangDeep Dive - slimeDeep Dive - DeepEPDeep Dive - TinkerDeep Dive - Light-PEFT
Back to Topics
LearnApplied LLM EngineeringModel Versioning & Deployment
⚙️MediumMLOps & Deployment

Model Versioning & Deployment

Turn an evaluated LLM change into an immutable release bundle, promote it through measured traffic, and roll back without losing lineage.

19 min read
Learning path
Step 75 of 177 in the full curriculum
Prompt Optimization with DSPySemantic Caching & Cost Optimization

Personalize this lesson

Adapt explanations and teaching visuals to your background and preferred voice.

DSPy produced a compiled prompt-program candidate that passed its frozen holdout and safety slices. You still don't have something safe to deploy. A live answer can change because of weights, an evidence gate, compiled prompt state, a tokenizer, a policy corpus, a container image, or a schema.

For the incident-response assistant you've been building, deployment means answering one exact question: which complete, evaluated release produced this response, and how quickly can traffic return to the last known good release?

Release identity changes when any behavior-producing field changes. The stable and candidate bundles keep the answer model, evidence gate, policy corpus, and evaluation contract pinned, but change the compiled prompt program, compile run, and serving image, producing different release IDs.
Pinned fields stay shared and reviewable. Changing compiled prompt state, its source run, or the serving image creates a new release bundle and a new release ID.

A release is more than model weights

The service contains an answer model, a runbook-evidence-classifier-v2 gate, and the compiled prompt program from the previous lesson. Its evidence gate decides which runbook excerpts are trusted enough to admit. Compiled prompt behavior must still constrain the answer to those excerpts and abstain when admitted evidence doesn't support the requested claim. If you roll back only one component while leaving a newer prompt, policy, or corpus active, you haven't restored previous behavior.

This is the release-bundle idea: store every input that can change visible output or operational safety in one immutable manifest. Keep the declared evaluation contract there too, then attach the resulting report artifact to the promotion decision. ML systems accumulate hidden dependencies between data, code, configuration, and serving infrastructure; deployment records need to make those dependencies reviewable.[1]Reference 1Hidden Technical Debt in Machine Learning Systems.https://research.google/pubs/hidden-technical-debt-in-machine-learning-systems/[2]Reference 2Challenges in Deploying Machine Learning: a Survey of Case Studies.https://arxiv.org/abs/2011.09926

Bundle fieldWhy it belongs in the release
Answer-model and evidence-gate identifiersDetermine model behavior
Prompt-program artifact and compile runTie optimized instructions and demonstrations to their source experiment
Tokenizer and prompt versionChange the text the model sees
Policy and corpus versionsDecide which evidence is available and when it's sufficient to serve
Serving image and schemaChange runtime behavior and API compatibility
Evaluation-suite hash and evaluator versionDeclare the comparable scoring contract required before promotion

Start with two full bundles. stable represents what's serving today. candidate changes the compiled prompt program after a DSPy run has passed development selection and an independent holdout gate.

define-release-bundles.py
1from dataclasses import asdict, dataclass, replace 2import hashlib 3import json 4 5@dataclass(frozen=True) 6class ReleaseBundle: 7 service: str 8 answer_model: str 9 evidence_gate: str 10 prompt_compile_run: str 11 tokenizer: str 12 prompt_version: str 13 policy_version: str 14 corpus_version: str 15 serving_image: str 16 input_schema: str 17 eval_suite: str 18 evaluator_version: str 19 20stable = ReleaseBundle( 21 service="incident-evidence-answerer", 22 answer_model="answer-model@sha256:3b07", 23 evidence_gate="runbook-evidence-classifier-v2@sha256:41d8", 24 prompt_compile_run="dspy-compile-run-041", 25 tokenizer="incident-tokenizer@sha256:91aa", 26 prompt_version="[email protected]", 27 policy_version="[email protected]", 28 corpus_version="incident-runbooks@sha256:corpus-5", 29 serving_image="registry.example/answerer@sha256:image-a", 30 input_schema="incident-answer.v2", 31 eval_suite="incident-grounding-suite@sha256:suite-7", 32 evaluator_version="claim-evidence-eval-v2", 33) 34 35candidate = replace( 36 stable, 37 prompt_compile_run="dspy-compile-run-052", 38 prompt_version="[email protected]", 39 serving_image="registry.example/answerer@sha256:image-b", 40) 41 42print(f"stable_prompt={stable.prompt_version}") 43print(f"candidate_prompt={candidate.prompt_version}") 44print(f"candidate_compile_run={candidate.prompt_compile_run}")
Output
1[email protected] 2[email protected] 3candidate_compile_run=dspy-compile-run-052

A version name such as v2 is useful for people, but it doesn't prove contents stayed fixed. The next cell derives identity from the canonical manifest. Change a prompt or image digest and the release ID changes too.

content-address-the-release.py
1def release_id(bundle: ReleaseBundle) -> str: 2 payload = json.dumps(asdict(bundle), sort_keys=True, separators=(",", ":")) 3 # Short prefix keeps the teaching output readable. Retain the full digest in production. 4 digest = hashlib.sha256(payload.encode("utf-8")).hexdigest()[:12] 5 return f"{bundle.service}@sha256:{digest}" 6 7stable_id = release_id(stable) 8candidate_id = release_id(candidate) 9prompt_patch_id = release_id(replace(candidate, prompt_version="[email protected]")) 10 11print(f"stable={stable_id}") 12print(f"candidate={candidate_id}") 13print(f"prompt_patch={prompt_patch_id}") 14print(f"prompt_patch_is_new_release={prompt_patch_id != candidate_id}")
Output
1stable=incident-evidence-answerer@sha256:2e5b61eae514 2candidate=incident-evidence-answerer@sha256:bdd31d60f3e9 3prompt_patch=incident-evidence-answerer@sha256:b01943fd7fa3 4prompt_patch_is_new_release=True

This lab abbreviates each SHA-256 digest to 12 hexadecimal characters so the state transitions stay readable. A production registry should retain the full digest as identity and use short prefixes only for display.

Why should a one-line prompt correction produce a new release ID?

Answer

The prompt can change served behavior even when model weights don't move. Reusing the old release ID would make traces and rollbacks lie about what users saw.

Artifacts stay fixed; aliases move

A registry record stores an immutable bundle. An alias such as production or canary is a mutable pointer used by traffic. This separation is what makes rollback simple: preserve both bundles and move the pointer back.

MLflow's current Model Registry workflow provides version aliases and tags, and its documentation marks fixed Model Stages as deprecated. That distinction matters: a model version is history; an alias expresses the current deployment decision.[3]Reference 3Model Registry Workflows | MLflow AI Platformhttps://mlflow.org/docs/latest/ml/model-registry/workflow/

An MLflow model alias still points to one registered model version, not automatically to the prompt, corpus, policy, schema, and serving image in the release bundle. Treat it as one component pointer, or register a wrapper artifact whose manifest resolves the complete bundle. Don't mistake a movable model alias for complete release identity.

Alias step chart over two fixed release-bundle rails: canary points to candidate release bdd31d60f3e9 before promotion, production starts on stable release 2e5b61eae514, moves to candidate at promotion, and returns to the retained stable release after an incident; both immutable bundle records remain available throughout.
The fixed rails are stored release IDs; the colored lines are aliases. Canary reaches the candidate first, production moves only after promotion, and rollback returns production to the retained stable bundle without rewriting either manifest.

The small registry below enforces that rule. register() keeps a deep copy of the bundle, and move_alias() only points at a registered ID.

registry-and-aliases.py
1from copy import deepcopy 2 3class ReleaseRegistry: 4 def __init__(self) -> None: 5 self._bundles: dict[str, ReleaseBundle] = {} 6 self._aliases: dict[str, str] = {} 7 8 def register(self, bundle: ReleaseBundle) -> str: 9 bundle_id = release_id(bundle) 10 existing = self._bundles.get(bundle_id) 11 if existing is not None and existing != bundle: 12 raise ValueError("release digest collision") 13 self._bundles[bundle_id] = deepcopy(bundle) 14 return bundle_id 15 16 def move_alias(self, alias: str, bundle_id: str) -> None: 17 if bundle_id not in self._bundles: 18 raise KeyError(f"unregistered release: {bundle_id}") 19 self._aliases[alias] = bundle_id 20 21 def resolve(self, alias: str) -> str: 22 return self._aliases[alias] 23 24registry = ReleaseRegistry() 25assert registry.register(stable) == stable_id 26assert registry.register(candidate) == candidate_id 27registry.move_alias("production", stable_id) 28 29print(f"registered={len(registry._bundles)}") 30print(f"production={registry.resolve('production')}") 31print(f"candidate_registered={candidate_id in registry._bundles}")
Output
1registered=2 2production=incident-evidence-answerer@sha256:2e5b61eae514 3candidate_registered=True

The teaching registry keeps move_alias() deliberately small. A production control plane should make that update atomic, record the expected previous target, and reject a stale promotion if another rollout moved the alias first.

Promotion begins with controlled evidence

Registration isn't approval. Continuous delivery for machine learning adds evaluation gates and monitoring to ordinary build-and-deploy practices because a valid artifact can still produce unacceptable behavior.[4]Reference 4Continuous Delivery for Machine Learning.https://martinfowler.com/articles/cd4ml.html

In this running system, don't introduce a generic "quality score" after spending several lessons defining a grounded evidence contract. The release gate should use the same measurements the incident, evaluation, and experiment lessons established:

  • supported_evidence_f1 measures whether supported incident claims are served correctly.
  • unsupported_serve_count must remain zero in the frozen high-risk suite.
  • p95_latency_ms (p95 latency) prevents a behaviorally acceptable gate from breaking the response budget.
  • Schema hash, evaluation-suite hash, and evaluator version prevent incomparable evidence from entering the decision.
  • evaluation_report preserves the report artifact a reviewer or incident responder can inspect later.
Offline promotion gate matches one candidate bundle to a pinned contract, passes release metrics, opens canary, and leaves production pinned.
Offline evidence counts only when release identity and evaluation contract match. Passing that gate can open canary traffic, while production stays pinned to the stable bundle.
offline-promotion-gate.py
1@dataclass(frozen=True) 2class OfflineEvidence: 3 release_id: str 4 eval_suite: str 5 evaluator_version: str 6 input_schema: str 7 evaluation_report: str 8 supported_evidence_f1: float 9 unsupported_serve_count: int 10 p95_latency_ms: int 11 12@dataclass(frozen=True) 13class Decision: 14 allowed: bool 15 reason: str 16 17def offline_gate(bundle: ReleaseBundle, evidence: OfflineEvidence) -> Decision: 18 if evidence.release_id != release_id(bundle): 19 return Decision(False, "evidence belongs to another release") 20 if evidence.eval_suite != bundle.eval_suite: 21 return Decision(False, "evaluation suite changed") 22 if evidence.evaluator_version != bundle.evaluator_version: 23 return Decision(False, "evaluator version changed") 24 if evidence.input_schema != bundle.input_schema: 25 return Decision(False, "schema mismatch") 26 if not evidence.evaluation_report: 27 return Decision(False, "evaluation report missing") 28 if evidence.supported_evidence_f1 < 0.92: 29 return Decision(False, "supported_evidence_f1 below 0.92") 30 if evidence.unsupported_serve_count != 0: 31 return Decision(False, "unsupported answer was served") 32 if evidence.p95_latency_ms > 500: 33 return Decision(False, "p95 latency exceeds 500 ms") 34 return Decision(True, "offline gate passed") 35 36candidate_offline = OfflineEvidence( 37 release_id=candidate_id, 38 eval_suite=candidate.eval_suite, 39 evaluator_version=candidate.evaluator_version, 40 input_schema=candidate.input_schema, 41 evaluation_report="reports/candidate-suite-7-redacted.json", 42 supported_evidence_f1=0.93, 43 unsupported_serve_count=0, 44 p95_latency_ms=472, 45) 46weaker_candidate = replace(candidate_offline, supported_evidence_f1=0.89) 47changed_evaluator = replace(candidate_offline, evaluator_version="claim-evidence-eval-v3") 48 49print(f"candidate={offline_gate(candidate, candidate_offline)}") 50print(f"weak_metric={offline_gate(candidate, weaker_candidate)}") 51print(f"changed_evaluator={offline_gate(candidate, changed_evaluator)}")
Output
1candidate=Decision(allowed=True, reason='offline gate passed') 2weak_metric=Decision(allowed=False, reason='supported_evidence_f1 below 0.92') 3changed_evaluator=Decision(allowed=False, reason='evaluator version changed')

Passing the offline gate permits further evaluation; it doesn't immediately replace production. Open a canary alias while production still points to the known-good bundle.

open-canary-only-after-gate.py
1def open_canary(bundle: ReleaseBundle, evidence: OfflineEvidence) -> Decision: 2 decision = offline_gate(bundle, evidence) 3 if decision.allowed: 4 registry.move_alias("canary", release_id(bundle)) 5 return decision 6 7canary_decision = open_canary(candidate, candidate_offline) 8 9print(f"canary_opened={canary_decision.allowed}") 10print(f"canary={registry.resolve('canary')}") 11print(f"production_still_stable={registry.resolve('production') == stable_id}")
Output
1canary_opened=True 2canary=incident-evidence-answerer@sha256:bdd31d60f3e9 3production_still_stable=True

Managed models need a documented pin

When your team owns weights, a digest can identify them directly. If a service calls a hosted model, the bundle must instead record the strongest fixed identifier that provider documents. For example, OpenAI model pages that expose snapshots describe snapshots as locking a particular model version for consistent behavior.[5]Reference 5Models | OpenAI APIhttps://platform.openai.com/docs/models

A provider can change the model behind a name you thought was stable, shifting your outputs with no deploy on your side. A controlled study comparing the March and June 2023 releases of GPT-4 and GPT-3.5 reported large behavior swings on the same tasks between snapshots, so the "same" service was not always the same service.[6]Reference 6How Is ChatGPT's Behavior Changing over Time?https://arxiv.org/abs/2307.09009 Pinning a documented snapshot reduces that risk but doesn't remove it: snapshots get deprecated, and a golden-set monitor against production still earns its keep.[7]Reference 7Deprecations | OpenAI APIhttps://developers.openai.com/api/docs/deprecations

Don't generalize that guarantee to every provider or every alias. Verify the exact provider documentation, store the chosen model identifier in the bundle, monitor deprecation notices, and rerun release gates before changing it.

Replay a failed production trace against the candidate

Frozen fixtures test the failures you already imagined. A deterministic replay tests a failure you actually shipped: pull the recorded trace for a request that went wrong under the current release, then re-run its exact prompt and tool-call sequence against the candidate release while holding the recorded evidence snapshot fixed. The only thing that varies is the release under test, so a difference in behavior is attributable to the candidate, not to a new request or a moved corpus.

This answers a precise regression question: would the candidate have made the same mistake on this real request? It sits between the offline gate and the shadow because it reuses recorded inputs rather than inventing fixtures or spending live traffic. The replay is read-only: it feeds recorded tool results back in rather than re-executing tools, so no incident state changes.

The stable release served a claim unsupported by admitted evidence. Replay that trace against the candidate prompt program while keeping the evidence gate and recorded inputs fixed:

deterministic-replay.py
1@dataclass(frozen=True) 2class RecordedTrace: 3 request_id: str 4 origin_release: str 5 prompt_version: str 6 evidence_version: str 7 tool_sequence: tuple[str, ...] 8 evidence_supports_claim: bool 9 served_unsupported_claim: bool 10 temperature: float = 0.0 11 seed: int = 0 12 model_snapshot: str = "" 13 14def prompt_would_serve(bundle: ReleaseBundle, evidence_supports_claim: bool) -> bool: 15 # Toy policy: only the compiled candidate prompt withholds unsupported claims. 16 # Replay must branch on bundle state, not a hardcoded "always good" rule. 17 if bundle.prompt_version == candidate.prompt_version: 18 return evidence_supports_claim 19 return True 20 21def replay_against(trace: RecordedTrace, bundle: ReleaseBundle) -> dict[str, object]: 22 would_serve = prompt_would_serve(bundle, trace.evidence_supports_claim) 23 reproduces = would_serve and not trace.evidence_supports_claim 24 return { 25 "candidate_release": release_id(bundle), 26 "prompt_version": bundle.prompt_version, 27 "tool_sequence": trace.tool_sequence, 28 "pins": (trace.temperature, trace.seed, trace.model_snapshot), 29 "candidate_reproduces_failure": reproduces, 30 } 31 32failed_trace = RecordedTrace( 33 request_id="req_88213", 34 origin_release=stable_id, 35 prompt_version=stable.prompt_version, 36 evidence_version="incident-runbooks@sha256:corpus-5", 37 tool_sequence=("lookup_incident", "fetch_runbook", "answer"), 38 evidence_supports_claim=False, 39 served_unsupported_claim=True, 40 temperature=0.0, 41 seed=7, 42 model_snapshot=stable.answer_model, 43) 44stable_replay = replay_against(failed_trace, stable) 45candidate_replay = replay_against(failed_trace, candidate) 46 47print(f"origin_release={failed_trace.origin_release}") 48print(f"stable_reproduces_failure={stable_replay['candidate_reproduces_failure']}") 49print(f"replayed_against={candidate_replay['candidate_release']}") 50print(f"candidate_prompt={candidate_replay['prompt_version']}") 51print(f"inputs_held_fixed={candidate_replay['tool_sequence']}") 52print(f"pins={candidate_replay['pins']}") 53print(f"candidate_reproduces_failure={candidate_replay['candidate_reproduces_failure']}") 54 55assert stable_replay["candidate_reproduces_failure"] is True 56assert candidate_replay["candidate_reproduces_failure"] is False
Output
1origin_release=incident-evidence-answerer@sha256:2e5b61eae514 2stable_reproduces_failure=True 3replayed_against=incident-evidence-answerer@sha256:bdd31d60f3e9 4[email protected] 5inputs_held_fixed=('lookup_incident', 'fetch_runbook', 'answer') 6pins=(0.0, 7, 'answer-model@sha256:3b07') 7candidate_reproduces_failure=False

A clean replay is real evidence that the candidate fixes this specific incident, but it isn't a general guarantee. The toy gate above reads bundle.prompt_version so two release bundles can disagree on the same fixed inputs while the evidence gate stays constant. A production replay harness still needs the recorded prompt, retrieved evidence version, tool inputs and outputs, and the non-deterministic pins on the trace (temperature, seed, and model snapshot), or the "replay" quietly becomes a fresh run whose difference you can't attribute. Keep a growing library of failed traces and add each new incident to it, so a future candidate must clear every past production mistake before promotion.

Shadow evaluation must be read-only

Offline fixtures can't cover every real request shape. A shadow sends a sanitized copy of a production request to the candidate while the stable release alone supplies the user-visible answer. It can reveal latency or evidence-support problems without exposing the candidate's text to customers.

The safety rule is easy to miss: a shadow isn't permitted to execute tools, send messages, change incident state, or write production state. Its output is evaluation data only. The envelope below also redacts an incident identifier before queuing the shadow request.

shadow-envelope.py
1import re 2 3@dataclass(frozen=True) 4class ShadowEnvelope: 5 candidate_release: str 6 stable_release: str 7 sanitized_text: str 8 side_effects_enabled: bool 9 response_visible_to_user: bool 10 evidence_supports_claim: bool 11 12@dataclass(frozen=True) 13class ShadowComparison: 14 route_agreement: bool 15 candidate_claim_safe: bool 16 candidate_would_serve: bool 17 stable_route: str 18 candidate_route: str 19 dropped: bool 20 21def make_shadow(text: str, evidence_supports_claim: bool) -> ShadowEnvelope: 22 sanitized = re.sub(r"INC-\d+", "[INCIDENT_ID]", text) 23 return ShadowEnvelope( 24 candidate_release=registry.resolve("canary"), 25 stable_release=registry.resolve("production"), 26 sanitized_text=sanitized, 27 side_effects_enabled=False, 28 response_visible_to_user=False, 29 evidence_supports_claim=evidence_supports_claim, 30 ) 31 32def compare_shadow(envelope: ShadowEnvelope, queue_full: bool = False) -> ShadowComparison: 33 if queue_full or envelope.side_effects_enabled: 34 return ShadowComparison(False, False, False, "dropped", "dropped", True) 35 # Stable already answered the user; shadow only scores the candidate copy. 36 stable_route = "serve" # production response already left the system 37 candidate_bundle = ( 38 candidate if envelope.candidate_release == candidate_id else stable 39 ) 40 candidate_would_serve = prompt_would_serve( 41 candidate_bundle, envelope.evidence_supports_claim 42 ) 43 candidate_route = "serve" if candidate_would_serve else "abstain" 44 route_agreement = candidate_route == stable_route 45 candidate_claim_safe = candidate_route == "abstain" or envelope.evidence_supports_claim 46 return ShadowComparison( 47 route_agreement=route_agreement, 48 candidate_claim_safe=candidate_claim_safe, 49 candidate_would_serve=candidate_would_serve, 50 stable_route=stable_route, 51 candidate_route=candidate_route, 52 dropped=False, 53 ) 54 55shadow = make_shadow( 56 "Can INC-48192 roll back based on runbook RB-7?", 57 evidence_supports_claim=False, 58) 59comparison = compare_shadow(shadow) 60dropped = compare_shadow(shadow, queue_full=True) 61 62print(f"shadow_text={shadow.sanitized_text}") 63print(f"candidate={shadow.candidate_release}") 64print(f"side_effects_enabled={shadow.side_effects_enabled}") 65print(f"response_visible={shadow.response_visible_to_user}") 66print(f"candidate_route={comparison.candidate_route}") 67print(f"route_agreement={comparison.route_agreement}") 68print(f"candidate_claim_safe={comparison.candidate_claim_safe}") 69print(f"dropped_when_queue_full={dropped.dropped}") 70 71assert comparison.candidate_would_serve is False 72assert comparison.route_agreement is False 73assert comparison.candidate_claim_safe is True 74assert dropped.dropped is True
Output
1shadow_text=Can [INCIDENT_ID] roll back based on runbook RB-7? 2candidate=incident-evidence-answerer@sha256:bdd31d60f3e9 3side_effects_enabled=False 4response_visible=False 5candidate_route=abstain 6route_agreement=False 7candidate_claim_safe=True 8dropped_when_queue_full=True

The one-pattern redaction is only a teaching fixture. A production shadow path needs schema-aware data minimization that covers every customer identifier and secret before the request reaches a queue, log, or candidate service.

In a deployed service, enqueue this envelope to a bounded worker queue and count dropped or failed comparisons. The compare step above is the missing half of shadow: after the candidate answers the sanitized copy, record route agreement, candidate claim safety, and latency against the stable path. Feed shadow_drop_rate from dropped or timed-out comparisons. Don't start an untracked background task in request scope and assume its evaluation record will survive process restarts.

The boolean in this teaching envelope is metadata, not an authorization boundary. Its worker still needs a read-only tool allowlist and credentials that can't write production state. Reject a shadow request if it asks for a side effect.

Why must a shadow evaluation of an agent candidate be read-only?

Answer

Shadow traffic duplicates production inputs. If the candidate can send messages, modify records, or call irreversible tools, evaluation itself creates duplicate side effects instead of passive evidence.

Canary traffic is visible and sticky

Shadow results can justify limited exposure, not automatic promotion. A canary sends a small share of real conversations to the candidate and returns those candidate responses to users. For conversational systems, one thread must remain on one bundle throughout the rollout. Otherwise a user can receive conflicting answers from stable and candidate releases in adjacent turns.

Use deterministic hashing of a conversation ID to choose new conversations reproducibly across workers. Python's built-in hash() is intentionally process-dependent, so it's the wrong bucketing function for that job. Hashing alone isn't enough, though: when canary traffic widens from 1% to 10%, the higher threshold could move an existing conversation from stable to candidate. Persist the first resolved release ID for the conversation lifetime.

sticky-canary-routing.py
1def bucket(conversation_id: str) -> int: 2 digest = hashlib.sha256(conversation_id.encode("utf-8")).hexdigest() 3 return int(digest[:8], 16) % 100 4 5conversation_assignments: dict[str, str] = {} 6aborted_releases: set[str] = set() 7 8def assigned_release(conversation_id: str, canary_percent: int) -> str: 9 existing = conversation_assignments.get(conversation_id) 10 if existing is not None and existing not in aborted_releases: 11 return existing 12 alias = "canary" if bucket(conversation_id) < canary_percent else "production" 13 bundle_id = registry.resolve(alias) 14 if bundle_id in aborted_releases: 15 bundle_id = registry.resolve("production") 16 conversation_assignments[conversation_id] = bundle_id 17 return bundle_id 18 19canary_thread = next( 20 f"thread-{index}" for index in range(1000) if bucket(f"thread-{index}") < 10 21) 22assignments = [assigned_release(canary_thread, canary_percent=10) for _ in range(3)] 23stable_thread = next( 24 f"thread-{index}" for index in range(1000) if bucket(f"thread-{index}") >= 10 25) 26stable_before_widen = assigned_release(stable_thread, canary_percent=10) 27stable_after_widen = assigned_release(stable_thread, canary_percent=100) 28 29print(f"canary_thread={canary_thread}") 30print(f"bucket={bucket(canary_thread)}") 31print(f"same_release_each_turn={len(set(assignments)) == 1}") 32print(f"assigned_to_candidate={assignments[0] == candidate_id}") 33print(f"existing_stable_thread_pinned_after_widen={stable_before_widen == stable_after_widen == stable_id}")
Output
1canary_thread=thread-6 2bucket=6 3same_release_each_turn=True 4assigned_to_candidate=True 5existing_stable_thread_pinned_after_widen=True

The dictionary is a teaching fixture. A real router persists assignments in conversation state or a rollout store, writes the first assignment atomically so concurrent opening turns can't disagree, expires it when the conversation ends, and records the resolved release ID in traces. Stickiness is normal-routing behavior, not permission to keep serving a failed candidate: an abort must override it.

Progressive rollout trace where 1 percent passes, 10 percent records an unsupported serve against a zero-tolerance gate, and canary traffic rolls back to zero.
Unsupported serves stay at 0% in the 1% window, then appear at 10%. The zero-tolerance support gate aborts the canary, clears candidate traffic, and restores every conversation to stable.

A canary window answers a new question

Offline evidence proves behavior on frozen examples. A live window tests traffic mix, serving latency, error rate, and evidence failures under actual request volume. Define those thresholds and a minimum sample size before sending any candidate traffic. A rate computed from a handful of requests isn't enough evidence to widen exposure. For this incident assistant, support is a hard safety invariant: one live unsupported serve fails the window rather than being absorbed into an error budget.

live-canary-window.py
1@dataclass(frozen=True) 2class LiveWindow: 3 name: str 4 request_count: int 5 p95_latency_ms: int 6 error_rate: float 7 unsupported_serve_rate: float 8 shadow_drop_rate: float 9 10def live_gate(window: LiveWindow) -> Decision: 11 if window.request_count < 1_000: 12 return Decision(False, "fewer than 1000 requests observed") 13 if window.p95_latency_ms > 550: 14 return Decision(False, "live latency exceeded 550 ms") 15 if window.error_rate > 0.01: 16 return Decision(False, "error rate exceeded 1%") 17 if window.unsupported_serve_rate > 0.0: 18 return Decision(False, "unsupported serve rate must remain zero") 19 if window.shadow_drop_rate > 0.02: 20 return Decision(False, "shadow telemetry incomplete") 21 return Decision(True, "live window passed") 22 23window_1_percent = LiveWindow("1%", 1_200, 481, 0.002, 0.000, 0.001) 24window_10_percent = LiveWindow("10%", 8_000, 493, 0.003, 0.014, 0.001) 25 26print(f"one_percent={live_gate(window_1_percent)}") 27print(f"ten_percent={live_gate(window_10_percent)}")
Output
1one_percent=Decision(allowed=True, reason='live window passed') 2ten_percent=Decision(allowed=False, reason='unsupported serve rate must remain zero')

Both windows exceed the 1,000-request minimum, so the metric verdicts are meaningful under this teaching contract. A low-volume window would remain blocked even if every observed rate happened to be zero.

Abort a canary; roll back a promotion

These actions sound similar but refer to different alias states:

  • If the candidate fails while only the canary alias receives traffic, abort the rollout. production never moved.
  • If a candidate was already promoted and later fails, roll back by repointing production to the retained stable release.

Keeping the distinction explicit prevents an incident report from claiming production was rolled back when the candidate was never production.

abort-and-rollback.py
1canary_percent = 10 2failed_window = live_gate(window_10_percent) 3if not failed_window.allowed: 4 aborted_releases.add(candidate_id) 5 canary_percent = 0 6 # Zero traffic is not enough: any consumer of registry.resolve("canary") must leave the failed ID. 7 registry.move_alias("canary", stable_id) 8 9print(f"canary_percent_after_abort={canary_percent}") 10print(f"canary_alias_after_abort={registry.resolve('canary')}") 11print(f"production_after_abort={registry.resolve('production')}") 12print(f"pinned_canary_thread_restored_stable={assigned_release(canary_thread, canary_percent) == stable_id}") 13 14assert registry.resolve("canary") == stable_id 15assert registry.resolve("production") == stable_id 16 17# Separate rollback drill: a promoted candidate later regresses at wider traffic. 18rollback_drill = deepcopy(registry) 19previous_production = rollback_drill.resolve("production") 20rollback_drill.move_alias("production", candidate_id) 21post_promotion_incident = replace(window_10_percent, name="100%") 22 23if not live_gate(post_promotion_incident).allowed: 24 rollback_drill.move_alias("production", previous_production) 25 26print(f"drill_production_after_rollback={rollback_drill.resolve('production')}") 27print(f"drill_restored_stable={rollback_drill.resolve('production') == stable_id}") 28print(f"actual_production_unchanged={registry.resolve('production') == stable_id}")
Output
1canary_percent_after_abort=0 2canary_alias_after_abort=incident-evidence-answerer@sha256:2e5b61eae514 3production_after_abort=incident-evidence-answerer@sha256:2e5b61eae514 4pinned_canary_thread_restored_stable=True 5drill_production_after_rollback=incident-evidence-answerer@sha256:2e5b61eae514 6drill_restored_stable=True 7actual_production_unchanged=True

Real progressive-delivery controllers encode the same operational mechanics. Argo Rollouts supports weighted canary steps and pauses; its background analysis can abort an unsuccessful rollout. With traffic routing, keeping the stable replica set available allows traffic to move back immediately on abort, at the cost of additional capacity during rollout.[8]Reference 8Argo Rollouts - Kubernetes Progressive Delivery Controllerhttps://argoproj.github.io/argo-rollouts/

Why might keeping the stable deployment warm during a canary be worth its GPU cost?

Answer

An alias change is only a fast recovery if the stable release can serve immediately. If old weights must reload after an incident, rollback can trade a behavior failure for a latency outage.

Record the decision a later engineer can reconstruct

A pipeline that moves aliases but doesn't persist why it moved them still creates mystery during an incident. Store the release IDs, evidence references, gate verdicts, rollout windows, actor or controller identity, decision timestamp, and final alias state as append-only events.

The candidate below is correctly rejected for production after its 10% canary window serves unsupported incident claims. The release isn't deleted. It remains available for diagnosis, while traffic stays with the known-good bundle.

release-decision-record.py
1@dataclass(frozen=True) 2class ReleaseEvent: 3 stage: str 4 release_id: str 5 decision: str 6 evidence: str 7 8events = ( 9 ReleaseEvent("register", candidate_id, "RECORDED", "manifest_sha"), 10 ReleaseEvent("offline_gate", candidate_id, "PASSED", candidate_offline.evaluation_report), 11 ReleaseEvent("canary_1_percent", candidate_id, "PASSED", "live:window-001"), 12 ReleaseEvent("canary_10_percent", candidate_id, "ABORTED", "live:window-010"), 13 ReleaseEvent("production", stable_id, "UNCHANGED", "rollback:not-needed"), 14) 15 16for event in events: 17 print(f"{event.stage}:{event.decision}:{event.evidence}") 18print("release_decision=REJECT_CANDIDATE_AFTER_CANARY_REGRESSION") 19print(f"active_production={registry.resolve('production')}")
Output
1register:RECORDED:manifest_sha 2offline_gate:PASSED:reports/candidate-suite-7-redacted.json 3canary_1_percent:PASSED:live:window-001 4canary_10_percent:ABORTED:live:window-010 5production:UNCHANGED:rollback:not-needed 6release_decision=REJECT_CANDIDATE_AFTER_CANARY_REGRESSION 7active_production=incident-evidence-answerer@sha256:2e5b61eae514

Production mapping

The lab deliberately uses plain Python so the state transitions are visible. A deployed stack usually splits the same responsibilities:

ResponsibilityProduction form
Store immutable model or component versionArtifact store plus model registry
Store prompt, policy, corpus, tokenizer, schema, and evaluator pinsRelease manifest in source control or deployment registry
Move candidate/production pointersRegistry aliases, deployment config, or feature flags limited to registered releases
Run offline evidence gatesCI job tied to exact manifest digest with a retained report artifact
Shift live traffic and pause on regressionsProgressive-delivery controller and metric analysis
Reconstruct impactRequest trace logs with resolved release ID and rollout event log

Feature flags remain useful, but their values must resolve to registered immutable release IDs. A flag that points at an arbitrary model name makes rapid changes easy and incident reconstruction impossible.

Complete the lesson

Mastery Check

Answer every question, then check your score. Score 75% or higher to mark this lesson complete.

1.Stable and candidate have identical answer-model and evidence-gate weights. The candidate changes prompt_version from [email protected] to [email protected] and serving_image from image-a to image-b. How should the release registry treat it before any traffic moves?

Correct answer: Register it as a separate immutable bundle with its own content-derived release ID, then require evaluation evidence for that exact ID before aliases move.

A release ID is derived from the complete manifest, not model weights alone. Prompt versions and serving images can change behavior, latency, or compatibility, so they require a new immutable bundle identity and evidence tied to that exact release before an alias such as production or canary moves.

2.A release bundle calls a hosted model using a provider's floating model name rather than a documented fixed snapshot. What risk does that introduce?

Correct answer: The manifest no longer identifies one fixed behavior-producing component. Pin the provider's documented fixed identifier when available and re-evaluate before moving to another identifier.

A floating provider alias can change behavior without any deployment on your side, so an otherwise immutable manifest stops identifying the exact component that produced a response. The safest available practice is to store the strongest fixed provider identifier, monitor deprecations, and rerun gates before changing it.

3.A candidate bundle declares eval_suite=suite-7, evaluator_version=claim-evidence-eval-v2, and input_schema=incident-answer.v2. Its evidence has the same release_id and metrics supported_evidence_f1=0.93, unsupported_serve_count=0, and p95_latency_ms=472, but evaluator_version=claim-evidence-eval-v3. What should the offline gate do?

Correct answer: Reject it because the evaluator version changed, so the evidence is not comparable to the bundle's declared evaluation contract.

The offline gate checks more than metric thresholds. It also requires the evidence release ID, evaluation suite, evaluator version, input schema, and retained report to match the bundle. A changed evaluator can alter scoring, so the candidate must be blocked until evidence is regenerated under the declared contract.

4.A shadow worker receives a sanitized copy of a production request. Which configuration satisfies the shadow safety invariant?

Correct answer: Hide the candidate response, allow only read-only tools, and use credentials that cannot write production state.

A shadow result is evaluation data, not a user response or an action request. Metadata saying that side effects are disabled is insufficient by itself, so the worker also needs a read-only tool allowlist and credentials incapable of changing production state.

5.An active conversation was assigned to stable at a 10% canary. The rollout widens to 100% before that conversation ends. How should the router handle its next turn?

Correct answer: Keep it on stable using the persisted assignment; an abort may override that assignment.

A wider threshold should affect new conversations, not reassign an active one. Persisting the first resolved release ID prevents adjacent turns from receiving conflicting behavior. This stickiness does not prevent an abort from removing a failed candidate.

6.A live gate requires p95 latency at most 550 ms, error rate at most 1%, zero unsupported serves, and shadow drop rate at most 2%. A window reports 540 ms, 0.8%, 0.1%, and 1%, respectively. What is the decision?

Correct answer: Reject it because any nonzero unsupported serve rate violates the support invariant.

Latency, error rate, and shadow drop rate remain within their limits, but the declared support contract is zero tolerance. A 0.1% rate still means unsupported claims reached users, so the controller must abort rather than widen traffic.

7.A 1% canary has observed 600 requests with p95 latency 480 ms, zero errors, zero unsupported serves, and zero shadow drops. The release contract requires at least 1,000 requests before a live verdict. What should the controller do?

Correct answer: Keep the rollout paused because the metric rates come from fewer than the required 1,000 requests.

A clean rate from a tiny sample can miss rare failures. The live gate must satisfy both the minimum-volume requirement and every metric threshold before exposure widens.

8.Production still points at the stable release while a candidate receives 10% of conversations through a canary alias. The live window fails because the candidate serves unsupported incident claims. What should be recorded?

Correct answer: Record a canary abort, set candidate traffic to 0, repoint the canary alias to stable, preserve the evidence, and leave production pointing at the stable release.

A rollback only applies after the production alias has been promoted to the candidate and must be moved back. Here production never moved, so the correct recovery is to abort the canary: remove candidate exposure, move the canary alias off the failed release, keep the stable production pointer, and retain the failed-window evidence for diagnosis.

9.A candidate passes offline gating and a 1% canary, production is moved from stable_id to candidate_id, and a later 100% live window fails the latency threshold. What should the release record show?

Correct answer: Record promotion evidence, then a rollback event that moves production from candidate_id back to stable_id and retains both bundles.

Because production had already moved to the candidate, the later failure is a rollback, not a canary abort. The record should preserve the evidence and rollout events that led to promotion, then append the rollback decision and final alias state. The candidate bundle should remain addressable for audit and diagnosis.

10.A controller can repoint production to stable in seconds, but the stable replicas were scaled to zero and their weights take several minutes to load. What has the rollout failed to guarantee?

Correct answer: Fast recovery, because alias rollback is not operationally fast unless stable is healthy and warm.

Moving an alias is only the control-plane part of recovery. If the retained stable release cannot immediately accept traffic, rollback can replace a behavior incident with timeouts or a latency outage. Stable capacity should remain ready until the candidate passes burn-in.

10 questions remaining.

Next Step
Continue to Semantic Caching & Cost Optimization

You now know that every response must resolve to an exact release bundle. Next you'll reuse prior responses safely by making cache scope and invalidation depend on the model, prompt, policy, corpus, and tenant context that produced them.

PreviousPrompt Optimization with DSPy
Share this article
XFacebookLinkedInBlueskyRedditHacker NewsEmail
References

Hidden Technical Debt in Machine Learning Systems.

Sculley et al. · 2015

https://research.google/pubs/hidden-technical-debt-in-machine-learning-systems/

Challenges in Deploying Machine Learning: a Survey of Case Studies.

Paleyes, A., Urma, R. G., & Lawrence, N. D. · 2022 · ACM Computing Surveys

https://arxiv.org/abs/2011.09926

Model Registry Workflows | MLflow AI Platform

MLflow · 2026

https://mlflow.org/docs/latest/ml/model-registry/workflow/

Continuous Delivery for Machine Learning.

Sato, D., Wider, A., & Windheuser, C. · 2019

https://martinfowler.com/articles/cd4ml.html

Models | OpenAI API

OpenAI · 2026

https://platform.openai.com/docs/models

How Is ChatGPT's Behavior Changing over Time?

Chen, L., Zaharia, M., & Zou, J. · 2023

https://arxiv.org/abs/2307.09009

Deprecations | OpenAI API

OpenAI · 2026

https://developers.openai.com/api/docs/deprecations

Argo Rollouts - Kubernetes Progressive Delivery Controller

Argo Project · 2026

https://argoproj.github.io/argo-rollouts/

Discussion

Questions and insights from fellow learners.

Discussion loads when you reach this section.