Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
Reasoning & Test-Time Compute treated thinking budgets, routers, and verifiers as production controls. Those controls are release inputs too. A prompt, router, judge, feature computation, or model alias can change cost and quality while the base weights stay fixed. Advanced MLOps (Machine Learning Operations) and DevOps (development and operations) for AI keep those inputs versioned, tested, and recoverable.
Consider an illustrative release assistant, DeployBuddy. It uses prompt [email protected], Low-Rank Adaptation (LoRA) adapter [email protected], and regional release-history features. A teammate updates the preprocessing for release_pattern_embedding, a vector used alongside recent failure count and average queue time. These inputs help the assistant estimate deployment timing against a service-level agreement (SLA). A feature store supplies stored model inputs by entity key; it doesn't automatically verify every transformation that produced them.
The change looks green. It passes the offline golden-set evaluation, a fixed collection of reviewed requests and expected behavior, plus the safety scanner and latency budget. It moves through staging, a 2% canary, and finally the production alias.
Forty-eight hours later, users in us-west receive SLA estimates off by a full day. The base weights, system prompt, and LoRA adapter are unchanged. Which release input moved?
In this scenario, investigation finds training-serving skew in release_pattern_embedding. The batch job that prepared the golden set normalized vectors, while online materialization did not. The day-long error is part of the teaching scenario, not a measured incident. A fingerprint difference would identify a changed input; reproducing the response with that input is still necessary to establish causality.
The model registry recorded prompt and adapter versions, but not the feature computation graph or timestamped values used during evaluation. Once features enter serving, model versioning alone leaves a hole.
Earlier in Model Versioning & Continuous Deployment, you learned immutable artifacts, aliases, canary traffic, and automated evaluation gates[1][2]. Apply those controls to the complete release, not just the model binary. DeployBuddy's missing feature computation illustrates the hidden dependencies described by Sculley et al.[3] A release tuple is simply the set of pinned inputs that must move together.
We'll follow that missing input through GitOps, historical feature retrieval, evaluation, and rollback. The examples execute local contracts and synthetic measurements. They don't deploy a cluster, call an LLM, or demonstrate live recovery time.

What failed in the DeployBuddy incident if model weights, prompt, and adapter were unchanged?
Answer
The release_pattern_embedding computation path changed. Offline evaluation normalized the vectors while online materialization did not, creating skew that the release lineage did not capture. Chunking is another input worth pinning, but this scenario hasn't established that it changed.
What DeployBuddy adds on top of model versioning
Classic MLOps already deals with data drift, feature stores, reproducibility, and probabilistic models. LLMOps widens the release boundary because application behavior can also change through prompts, retrieval indexes, tool schemas, judges, routing, and generation settings, even when model weights stay fixed.[4][3]
| Dimension | Classic MLOps | LLMOps |
|---|---|---|
| Core artifact | Model binary, feature store | Prompts, RAG configs, embeddings, eval sets, plus the model |
| Output behavior | Often scored against labeled targets | Open-ended or sampled output often needs rubric-based scoring |
| Tests | Unit tests plus task metrics | Those tests plus semantic eval, golden sets, and safety checks |
| CI cost | Data and model evaluation can be expensive | Generation-based evals add token, latency, and judge costs |
| Monitoring | Accuracy, data drift, service health | Groundedness, refusals, policy violations, cost, plus drift |
Three differences change how you release:
- Open-ended and sampled output. Repeated calls can produce different text, and deterministic generation can still have several acceptable responses. Exact equality works for structured subcomponents. End-to-end quality usually needs a score and a rubric.
- Prompts (and RAG config, guardrails, eval sets) are first-class artifacts. A one-line prompt edit can move refusal rate or grounding without changing weights. Version, review, and roll back that prompt like code instead of burying it in application strings.
- Eval gates join the test suite. Generated answers often have no single exact-string target. Promotion needs scored evaluation against curated golden sets alongside ordinary unit tests. Write the golden set before changing the prompt so the target doesn't move with the candidate.
Together they define the release path: GitOps versions every input; feature parity checks expose skew; eval gates decide whether a tuple is publishable; progressive delivery limits exposure; observability supplies the evidence for rollback.
Name three LLM application concerns that extend ordinary MLOps.
Answer
Open-ended or sampled generation that needs scored evaluation, prompts and retrieval/tool config as first-class versioned artifacts, and generation-based evals whose token, judge, and latency costs become release concerns.
GitOps for LLM systems
GitOps treats prompts, feature definitions, model aliases, and deployment configuration as desired state, not strings hidden in application code. OpenGitOps names four properties: the state is declarative, versioned and immutable, pulled by an agent, and continuously reconciled with the live cluster.[5] Those properties make an ordinary promotion a reviewed commit.
Keep artifact bytes in appropriate immutable storage and their verified digests in the manifest; Git needn't contain model weights or private feature values. An incident also needs an audited break-glass path and a durable desired-state repair.
Before looking at the controller, look at what the repository makes visible:
1llm-ops-repo/
2├── prompts/
3│ ├── support/
4│ │ ├── [email protected]
5│ │ ├── [email protected]
6│ │ └── tests/
7│ │ └── support_golden.jsonl
8│ └── policy/
9│ └── [email protected]
10├── features/
11│ └── release_history/
12│ ├── feature_view.py # Feast-style definition
13│ └── entity_keys.yaml # region_id join keys
14├── models/
15│ └── deploy-assistant/
16│ ├── [email protected]
17│ ├── [email protected] # points to registry URI + hash
18│ └── promotion-policy.yaml
19├── deployments/
20│ └── production/
21│ └── aliases.yaml # production -> [email protected]
22└── .github/workflows/
23 └── llm-promote.ymlEvaluate behavior-changing inputs before approving the merge. The release pipeline can then publish the approved candidate and reconcile deployment state:
- Parse changed files and determine what must be re-evaluated.
- Run fast checks (prompt linting, schema validation, cost estimation).
- Run the offline evaluation suite against the exact candidate and feature snapshot pinned in the PR.
- If gates pass, approve those immutable prompt/model versions. Registration may happen earlier so evaluation can load the candidate; don't rebuild it afterward.
- Update the alias file or deployment manifest.
- The GitOps operator (Argo CD or Flux) pulls the new desired state and reconciles the cluster to it.
Step 6 only makes the cluster match Git. For an ordinary rollback, use a revert commit or an alias-manifest change and run the pipeline again.
An emergency controller may shift traffic first. Record the action, coordinate with reconciliation, and update desired state so a later sync can't restore the failed candidate. Controller behavior matters: automatic drift correction is a configurable policy, not an inevitable immediate action in every GitOps installation.
For example, Argo CD's selfHeal controls automatic correction of live-state drift, and its rollback command isn't available with automated sync enabled. Use a reviewed desired-state change for ordinary rollback; design and rehearse any temporary reconciliation suspension used during an emergency.[6]
Model registries such as MLflow separate a mutable alias from a version. Loading models:/deploy-assistant@production resolves its current target; reassigning the alias doesn't rewrite a model object already loaded by a worker.[7] A serving controller must resolve the intended version, prepare healthy capacity, switch routing, and verify the version actually served. Pin the base checkpoint separately from the adapter: [email protected] alone doesn't identify a model.
Putting prompt Markdown files in Git is the easy part. The pull request must evaluate the same feature definitions, golden sets, guardrails, and alias state that production will use. Pinning a feature-view name still allows batch and online code to compute release_pattern_embedding differently. The release needs a check for that computation, not only its name.
Why is GitOps more than "store prompts in git" for LLM systems?
Answer
The repo must also version feature definitions, eval datasets, guardrail policies, aliases, and promotion rules. Otherwise a prompt commit can pass tests against one hidden environment and ship into another.
Ask what the registry should compare. Hash the composed tuple, not the prompt string. releases@5 is a different release from releases@4 even when [email protected] and [email protected] stay put.
1import hashlib
2import json
3
4def release_digest(release: dict[str, str]) -> str:
5 payload = json.dumps(release, sort_keys=True).encode()
6 return hashlib.sha256(payload).hexdigest()
7
8baseline = {
9 "prompt": "[email protected]", "base": "base@42", "adapter": "[email protected]",
10 "serving_image": "image@7", "feature_view": "releases@4",
11}
12candidate = baseline | {"feature_view": "releases@5"}
13baseline_digest = release_digest(baseline)
14candidate_digest = release_digest(candidate)
15print(f"baseline_display={baseline_digest[:10]} candidate_display={candidate_digest[:10]}")
16print(f"stored_digest_length={len(baseline_digest)}")
17print(f"same_release={baseline_digest == candidate_digest}")1baseline_display=89b2794cde candidate_display=7c6edf0ae2
2stored_digest_length=64
3same_release=FalseThe output makes the hidden change visible: prompt and model identifiers match, but the full SHA-256 digest doesn't. Registry equality and trace binding use that full digest; the ten-character values above are display labels only.
Now follow the failure path. If the controller moved runtime traffic to last-known-good while Git still names release-a91, reconciliation is still open. The next snippet checks that mismatch.
1git_desired_alias = "release-a91"
2runtime_alias = "last-known-good"
3break_glass_event = {"reason": "p95 TTFT 910>850", "actor": "rollout-controller"}
4
5needs_reconciliation = git_desired_alias != runtime_alias
6print(f"runtime_alias={runtime_alias} audited={bool(break_glass_event)}")
7print(f"reconciliation_required={needs_reconciliation}")1runtime_alias=last-known-good audited=True
2reconciliation_required=TrueThis script detects a disagreement, not a completed recovery. A nonempty dictionary isn't proof that an audit event was durably stored, and the print statement opens no commit. Confirm runtime health and the served release, then reconcile desired state. The local lifecycle exercise later makes these distinct state transitions explicit.

Feature stores for embeddings and LLM features
Git pinned releases@5, but a name isn't proof that two paths computed the same values. DeployBuddy still failed because batch and online materialization used different logic. Embedding pipelines are a common source of training-serving skew in LLM systems. A RAG pipeline or personalized prompt may depend on dozens of embedding features: user-history summary vectors, document-chunk embeddings, category-affinity vectors, and time-decayed engagement signals.
A feature store (such as Feast, Tecton, or a custom layer around Redis, object storage, and a vector database) helps answer three different questions, provided definitions and materialization paths are versioned and tested:[8]
- Which value existed then? Point-in-time correctness. Join feature event timestamps at or before each example's timestamp, within the feature's allowed age. Separately handle late arrivals and corrections: an old event timestamp doesn't prove that its value was available then.[9]
- Did both paths compute the same value? Consistency checks. A shared feature view and parity tests can detect incompatible batch and online values. A feature store doesn't repair divergent transformation logic by itself.
- What produced this vector? Lineage. Every embedding vector should carry the exact model checkpoint, chunker version, and preprocessing hash that produced it.
Feature platforms such as Feast can provide offline stores, online stores, and point-in-time feature retrieval. Treat that platform as one part of lineage, not as a substitute for versioning. If your RAG stack uses a separate vector database, version the vector index, corpus snapshot, embedding model, chunker, and normalization hash beside the feature view.
In DeployBuddy, the join key is region_id, not a user ID. Online retrieval asks for the current stored features for us-west; historical retrieval supplies an example timestamp. These are different interfaces, not one arbitrary-time query against the online store.[10]
This Feast API sketch declares a file source and a feature view. Constructing the objects doesn't create the Parquet data, materialize an online store, or prove parity. The source file must contain the entity key, event and created timestamps, and declared feature columns.
1from datetime import timedelta
2
3from feast import Entity, FeatureView, Field, FileSource
4from feast.types import Array, Float32, Int64
5
6region = Entity(name="region", join_keys=["region_id"])
7release_history_source = FileSource(
8 path="release_history.parquet",
9 timestamp_field="event_timestamp",
10 created_timestamp_column="available_at",
11)
12
13release_history_view = FeatureView(
14 name="release_history_embeddings",
15 entities=[region],
16 ttl=timedelta(days=90),
17 schema=[
18 Field(name="recent_failure_count", dtype=Int64),
19 Field(name="avg_queue_minutes_30d", dtype=Float32),
20 Field(name="release_pattern_embedding", dtype=Array(Float32)),
21 ],
22 source=release_history_source,
23)At serving time, Feast's get_online_features takes feature references or a feature service plus entity rows. Historical get_historical_features instead joins against timestamped examples. A feature-view TTL constrains the historical lookback; don't treat that setting alone as proof of an online freshness SLA. Check materialization lag and the timestamps of served values.
Feature-view version syntax and backend support also matter. Treat releases@4 and releases@5 here as application release labels, not portable Feast lookup strings. As checked September 22, 2026, Feast documents references such as drivers_activity@v2:trips_today, but version-qualified online reads require enable_online_feature_view_versioning: true and are currently supported only by its SQLite online store. A registry version alone doesn't establish that your chosen backend can serve an older snapshot.[11]
Current Feast documentation also exposes filter_by_created_timestamp=True for historical retrieval. With a suitable non-null created-timestamp column and a supported offline backend, it excludes values created after the example time. Without that option, created timestamps can deduplicate corrections without excluding future-created corrections. Unsupported stores raise an error. Even the filtered join reproduces online availability only if the timestamp records actual availability, not merely when an upstream event occurred.[9]
Without a declared feature layer and parity checks, teams copy embedding code between training notebooks and serving services. The copies drift, and quality degrades as it did for us-west. Feature parity is one gate; prompt text is another. Both belong in the same promotion machinery.
What does point-in-time correctness prevent?
Answer
An event-time join excludes future events. Reproducing what serving could actually know also requires excluding late-arriving or corrected values that became available after the example time, and enforcing a freshness window. Event time and availability time answer different questions.
For a noon us-west request, a 09:00 reading available at 09:05 is eligible. A correction to that same 09:00 event arriving at 13:00 isn't. This local join applies both clocks and a one-day lookback; it returns None when no valid feature exists rather than borrowing a future value.
1from datetime import datetime, timedelta
2
3def utc(value: str) -> datetime:
4 return datetime.fromisoformat(value + "+00:00")
5
6def historical_value(rows, request_time, ttl):
7 eligible = [
8 row for row in rows
9 if request_time - ttl <= row[0] <= request_time and row[1] <= request_time
10 ]
11 return max(eligible, key=lambda row: (row[0], row[1]))[2] if eligible else None
12
13feature_history = [
14 (utc("2026-04-09T09:00:00"), utc("2026-04-09T09:05:00"), 4),
15 (utc("2026-04-10T09:00:00"), utc("2026-04-10T09:05:00"), 5),
16 (utc("2026-04-10T09:00:00"), utc("2026-04-10T13:00:00"), 9),
17 (utc("2026-04-11T09:00:00"), utc("2026-04-11T09:05:00"), 8),
18]
19request_time = utc("2026-04-10T12:00:00")
20print("available by noon:", historical_value(feature_history, request_time, timedelta(days=1)))
21print("within one hour:", historical_value(feature_history, request_time, timedelta(hours=1)))
22assert historical_value(feature_history, request_time, timedelta(days=1)) == 51available by noon: 5
2within one hour: NoneThe join excludes the future correction, but says nothing about vector preprocessing. Next compare the declared computation configurations. Keep full digests for equality and shorten them only in printed labels.
1import hashlib
2import json
3
4def fingerprint(config: dict[str, str]) -> str:
5 return hashlib.sha256(json.dumps(config, sort_keys=True).encode()).hexdigest()
6
7offline = {"chunker": "sentence@3", "embedding": "embed@2", "normalize": "l2"}
8online = {"chunker": "sentence@3", "embedding": "embed@2", "normalize": "none"}
9print(f"offline={fingerprint(offline)[:8]} online={fingerprint(online)[:8]}")
10print(f"parity={fingerprint(offline) == fingerprint(online)}")1offline=ce667443 online=660645c4
2parity=Falseparity=False blocks this proposed release because its declared transforms differ. Equal hashes would establish only equal serialized configuration, not equal code, source data, or computed values. Run both paths on the same fixtures and compare actual vectors with a declared tolerance.
Normalization can matter concretely. Predict which document wins with query [1, 0] before running both metrics:
1import math
2
3documents = {"A": [3.0, 4.0], "B": [2.0, 0.0]}
4query = [1.0, 0.0]
5
6def dot(left, right):
7 return sum(a * b for a, b in zip(left, right, strict=True))
8
9def unit(vector):
10 length = math.hypot(*vector)
11 return [value / length for value in vector]
12
13raw = {key: dot(query, value) for key, value in documents.items()}
14normalized = {key: dot(unit(query), unit(value)) for key, value in documents.items()}
15print(f"raw_dot_winner={max(raw, key=raw.get)} scores={list(raw.values())}")
16print(f"normalized_dot_winner={max(normalized, key=normalized.get)} scores={list(normalized.values())}")
17assert raw["A"] > raw["B"] and normalized["A"] < normalized["B"]1raw_dot_winner=A scores=[3.0, 2.0]
2normalized_dot_winner=B scores=[0.6, 1.0]Raw dot-product ranking reverses. Cosine scoring normalizes by vector length itself, so this particular scale change wouldn't alter its ranking. Zero vectors need an explicit policy, and robust implementations should avoid overflowing or underflowing the squared norm. The downloadable lab rescales finite nonzero vectors before normalization. Test the downstream metric rather than assuming every normalization mismatch causes the same failure.
Eval-gated prompt and model promotion
Once the feature path passes, the template can still change behavior. A one-line system-prompt edit can shift refusal rates, tone, or factual grounding without touching weights.
Version prompts independently from weights, with tests and promotion policies. Semantic versions can express your team's compatibility rules, but registries may instead use numeric versions or commit IDs. Version labels don't establish compatibility without a defined input/output contract and evaluation evidence.
MLflow illustrates another boundary: its prompt template versions are immutable, but attached model_config is mutable. Snapshot generation configuration in your release manifest rather than assuming a prompt version freezes it. Alias-based prompt loading can also cache a resolved value, so test refresh behavior instead of treating a registry update as instantaneous propagation.[12]
Ask what someone would need to reproduce a candidate later. A production prompt registry entry typically records:[13][4]
- The exact template text (with variable placeholders)
- Few-shot examples, rubric snippets, or tool schemas pinned by hash
- Associated guardrail policy version
- Evaluation results on multiple golden sets (helpfulness, groundedness, safety, task-specific accuracy)
- Cost and latency characteristics under different generation settings
Security belongs in the same tuple. A changed tool schema or guardrail policy can change which actions a candidate may request, so safety gates must compare those versions alongside output scores.
Because evals call the model, they take longer and cost more than ordinary unit tests. Tier the gates so checks without model calls fail before paid generation runs[4]:
- Fast gate (no model calls, every PR): schema and template validation, artifact digests, fixture parity, and input-size estimates. Cached responses test the parser, not the new prompt's generated behavior.
- Eval gate (on behavior-changing inputs): use a pinned model version, not a mutable alias, with the pinned feature snapshot, golden set, generation settings, and judge. Evaluate feature, retrieval, tool, and policy changes too. Record sampling uncertainty, critical slices, and actual input/output/tool costs.
You may need to register an immutable candidate before evaluating it. Separate registration from approval: retain failed candidates for audit, but move the approved deployment reference only after the required gates pass. Avoid rebuilding a different artifact after evaluation.
For a hosted model, pin the strongest provider-available version identifier and retain request IDs and observed response metadata. As an operational limit, a manifest can freeze your controlled inputs; it can't by itself freeze an external service. Retained evaluated responses and subsequent regression checks still matter.
Why does a prompt registry need cost and latency metadata instead of template text alone?
Answer
Prompt changes can increase output length, tool use, retrieved context, refusal rate, and latency. A registry entry must capture serving behavior so promotion gates can block expensive or slow "quality wins."
Predict the gate result before running it. Quality rises from 0.84 to 0.87, but tokens jump 27.4% against a 15% cap. The candidate is a quality win and still not promotable.
1baseline = {"quality": 0.84, "tokens": 620, "p95_ms": 820}
2candidate = {"quality": 0.87, "tokens": 790, "p95_ms": 970}
3limits = {"max_token_increase": 0.15, "max_p95_ms": 1000}
4
5token_increase = candidate["tokens"] / baseline["tokens"] - 1
6quality_improved = candidate["quality"] > baseline["quality"]
7passes = (
8 quality_improved
9 and token_increase <= limits["max_token_increase"]
10 and candidate["p95_ms"] <= limits["max_p95_ms"]
11)
12print(f"quality_improved={quality_improved} token_increase={token_increase:.1%}")
13print(f"promote={passes}")1quality_improved=True token_increase=27.4%
2promote=Falsequality_improved=True is not enough. The 27.4% token increase breaks the 15% budget, even though p95 latency stays under its 1000 ms limit. This fixture tests policy arithmetic; it doesn't establish that a 0.03 score difference is statistically reliable. Token count is also only a cost proxy when token types, model prices, cache discounts, and tool use are held comparable.
In shadow traffic, the gateway returns the control response and samples a second call to the candidate. Hold inputs constant when isolating a prompt change; replay the full candidate tuple when testing a broader release. Shadow responses aren't shown to users, but their compute, data access, and side effects are real.
DeployBuddy's shadow path must not deploy software, send messages, or modify tickets. Use recorded or read-only tool results and separately isolated credentials. Bound concurrency, timeouts, and spend so shadow calls don't starve production. Apply the same access and retention policy to copied inputs and logged outputs as to the original request.
After enough requests for the predefined decision rule, compare judge scores, latency, and cost before starting a small sticky canary.

Mirroring a request is therefore only the transport mechanism. Safe shadowing also requires tool isolation, resource limits, and a representative sample.
What does shadow traffic prove, and what does it not prove?
Answer
Shadowing provides evidence about the sampled requests and allows paired comparisons of quality, latency, and cost. It doesn't prove user preference, coverage of rare cases, or safety of unrestricted tool execution. Users still see only control responses.
Shadow can't show user preference because users see only control. A canary does expose users, so the gateway should hash a stable subject.
region:us-west:conversation-1 stays on control (bucket 14). conversation-2 stays on the candidate (bucket 0). Neither conversation flips version mid-thread.
1import hashlib
2
3def canary_bucket(subject_id: str) -> int:
4 return int(hashlib.sha256(subject_id.encode()).hexdigest()[:8], 16) % 100
5
6def route(subject_id: str, percentage: int) -> str:
7 if type(percentage) is not int or not 0 <= percentage <= 100:
8 raise ValueError("percentage must be an integer between 0 and 100")
9 return "candidate" if canary_bucket(subject_id) < percentage else "control"
10
11control = "region:us-west:conversation-1"
12canary = "region:us-west:conversation-2"
13control_routes = [route(control, 10) for _ in range(3)]
14canary_routes = [route(canary, 10) for _ in range(3)]
15print(f"control_bucket={canary_bucket(control)} control_routes={control_routes}")
16print(f"canary_bucket={canary_bucket(canary)} canary_routes={canary_routes}")
17print(f"sticky={len(set(control_routes)) == 1 and len(set(canary_routes)) == 1}")
18for invalid in (True, 2.5, float("nan"), "10"):
19 try:
20 route(control, invalid)
21 except ValueError:
22 pass
23 else:
24 raise AssertionError("invalid canary percentage accepted")1control_bucket=14 control_routes=['control', 'control', 'control']
2canary_bucket=0 canary_routes=['candidate', 'candidate', 'candidate']
3sticky=TrueThe hash is stable while the percentage and candidate mapping are fixed. Raising the percentage can move a previously controlled conversation into the canary; changing the candidate behind the same label can also switch versions mid-thread. If a conversation must stay pinned, persist its assigned immutable release at creation and change assignment only under a deliberate migration or emergency policy. A 10% hash bucket targets a fraction of subjects, not necessarily 10% of requests or tokens.
Stable assignment still says nothing about economics. The next snippet compares mean judge score and mean tokens on the same three us-west requests. A +0.020 judge lift with +17.2% tokens should pause for a cost owner, not auto-promote.
1control = {"judge_scores": [0.82, 0.86, 0.84], "tokens": [600, 640, 620]}
2shadow = {"judge_scores": [0.85, 0.88, 0.85], "tokens": [700, 760, 720]}
3
4quality_delta = sum(shadow["judge_scores"]) / 3 - sum(control["judge_scores"]) / 3
5token_delta = sum(shadow["tokens"]) / 3 / (sum(control["tokens"]) / 3) - 1
6print(f"quality_delta={quality_delta:+.3f} token_delta={token_delta:+.1%}")
7print(f"needs_cost_review={token_delta > 0.10}")1quality_delta=+0.020 token_delta=+17.2%
2needs_cost_review=TrueThe three-request sample illustrates arithmetic, not a statistically justified quality lift. It uses a separate 10% shadow-cost review trigger; the earlier offline gate used 15%. Define these policies before measuring, and don't repeatedly rescore the same tiny sample until it happens to pass.
Automated rollback on evaluation regression
A canary is live user traffic, not a health check. DeployBuddy's us-west SLA miss didn't page a 5xx; the release needed feature fingerprints to expose it.
The canary also has a hard operational boundary. In this lesson's window, p95 TTFT is 910 ms against an 850 ms budget. A controller can revert that objective breach without waiting for a postmortem.
Observability supplies three kinds of evidence, because no single signal describes an LLM release:
- Infrastructure metrics (TTFT p95, 5xx rate, queue depth)
- Business metrics (task completion rate, thumbs down rate)
- Quality signals (LLM-as-judge groundedness, refusal rate, toxicity classifier score on a continuous sample)
The policy must distinguish hard boundaries from noisy evidence. A sustained latency, error, cost, or safety breach can shift traffic to the last known good version. A small sampled-judge or business-metric delta may pause promotion and request review.
Missing, stale, or non-finite required telemetry isn't evidence of health: its declared action must pause promotion or roll back exposure rather than continue. Updating an alias is fast control-plane work, but the runtime may still need to warm or load the restored artifact.

The policy below returns a recommendation, not an infrastructure mutation. Its declared operational and safety rules outrank review-only quality rules. That precedence is a policy choice, not a universal rule that latency matters more than safety. Feed it the fixture window from the figure: TTFT 910 ms and groundedness 0.79. Before comparing thresholds, validate the release identity, aggregation window, sample count, freshness, and numeric domain. A real collector must produce those aggregates and evaluate confidence intervals; labeling a number p95 doesn't compute one.
1from dataclasses import dataclass
2import sys
3
4@dataclass(frozen=True)
5class Threshold:
6 max_value: float | None = None
7 min_value: float | None = None
8 window: str = "5m"
9 action: str = "rollback"
10 max_age_seconds: int = 120
11 unavailable_action: str = "rollback"
12 min_samples: int = 100
13
14@dataclass(frozen=True)
15class MetricSample:
16 value: float
17 age_seconds: int
18 release_id: str
19 window: str
20 sample_count: int
21
22class RollbackPolicy:
23 def __init__(self, release_id: str):
24 self.release_id = release_id
25 self.thresholds = {
26 "p95_ttft_ms": Threshold(max_value=850, window="5m"),
27 "groundedness_score": Threshold(
28 min_value=0.82,
29 window="15m",
30 action="pause_for_review",
31 max_age_seconds=900,
32 unavailable_action="pause_for_review",
33 ),
34 "toxicity_rate": Threshold(max_value=0.004, window="10m"),
35 "task_completion_rate": Threshold(
36 min_value=0.91,
37 window="30m",
38 action="pause_for_review",
39 max_age_seconds=600,
40 unavailable_action="pause_for_review",
41 ),
42 }
43
44 def violations(self, metrics: dict[str, MetricSample]) -> list[tuple[str, str]]:
45 violations: list[tuple[str, str]] = []
46
47 for metric, rule in self.thresholds.items():
48 sample = metrics.get(metric)
49 if sample is None:
50 violations.append(
51 (rule.unavailable_action, f"{metric} required telemetry is missing")
52 )
53 continue
54 if not isinstance(sample, MetricSample):
55 violations.append(
56 (rule.unavailable_action, f"{metric} required telemetry has an invalid record type")
57 )
58 continue
59 if (
60 type(sample.value) not in (int, float)
61 or not -sys.float_info.max <= sample.value <= sys.float_info.max
62 ):
63 violations.append(
64 (rule.unavailable_action, f"{metric} required telemetry is not a finite real number")
65 )
66 continue
67 if (
68 sample.release_id != self.release_id
69 or sample.window != rule.window
70 or type(sample.sample_count) is not int
71 or type(sample.age_seconds) is not int
72 or sample.sample_count < rule.min_samples
73 or sample.age_seconds < 0
74 ):
75 violations.append(
76 (rule.unavailable_action, f"{metric} telemetry identity, window, count, or age is invalid")
77 )
78 continue
79 if sample.value < 0 or (metric != "p95_ttft_ms" and sample.value > 1):
80 violations.append(
81 (rule.unavailable_action, f"{metric} telemetry is outside its numeric domain")
82 )
83 continue
84 if sample.age_seconds > rule.max_age_seconds:
85 violations.append(
86 (
87 rule.unavailable_action,
88 f"{metric} telemetry is stale at {sample.age_seconds}s",
89 )
90 )
91 continue
92
93 value = sample.value
94 if rule.max_value is not None and value > rule.max_value:
95 violations.append(
96 (rule.action, f"{metric}={value} exceeds max {rule.max_value} over {rule.window}")
97 )
98 if rule.min_value is not None and value < rule.min_value:
99 violations.append(
100 (rule.action, f"{metric}={value} below min {rule.min_value} over {rule.window}")
101 )
102
103 return violations
104
105 def decision(self, metrics: dict[str, MetricSample]) -> str:
106 actions = {action for action, _ in self.violations(metrics)}
107 if "rollback" in actions:
108 return "rollback"
109 if "pause_for_review" in actions:
110 return "pause_for_review"
111 return "continue"
112
113policy = RollbackPolicy("release-a91")
114canary_metrics = {
115 "p95_ttft_ms": MetricSample(910, 30, "release-a91", "5m", 2000),
116 "groundedness_score": MetricSample(0.79, 30, "release-a91", "15m", 300),
117 "toxicity_rate": MetricSample(0.001, 30, "release-a91", "10m", 2000),
118 "task_completion_rate": MetricSample(0.93, 30, "release-a91", "30m", 1000),
119}
120
121violations = policy.violations(canary_metrics)
122print(f"decision={policy.decision(canary_metrics)}")
123for action, reason in violations:
124 print(f"{action}: {reason}")
125
126invalid_telemetry = {
127 "groundedness_score": MetricSample(0.85, 1_000, "release-a91", "15m", 300),
128 "toxicity_rate": MetricSample(float("nan"), 30, "release-a91", "10m", 2000),
129 "task_completion_rate": MetricSample(0.93, 30, "release-a91", "30m", 1000),
130}
131print(f"invalid_telemetry_decision={policy.decision(invalid_telemetry)}")
132for action, reason in policy.violations(invalid_telemetry):
133 print(f"{action}: {reason}")1decision=rollback
2rollback: p95_ttft_ms=910 exceeds max 850 over 5m
3pause_for_review: groundedness_score=0.79 below min 0.82 over 15m
4invalid_telemetry_decision=rollback
5rollback: p95_ttft_ms required telemetry is missing
6pause_for_review: groundedness_score telemetry is stale at 1000s
7rollback: toxicity_rate required telemetry is not a finite real numberThe check accepts aggregates bound to one immutable release. It requires real, finite values plus integer sample counts and nonnegative whole-second ages. Python type annotations don't enforce those fields. In particular, comparing a NaN count with a minimum doesn't reject it: both nan < 100 and nan >= 100 are false.
1healthy = canary_metrics | {
2 "p95_ttft_ms": MetricSample(790, 30, "release-a91", "5m", 2000),
3 "groundedness_score": MetricSample(0.85, 30, "release-a91", "15m", 300),
4}
5assert policy.decision(healthy) == "continue"
6bad_records = [
7 ("nan_age", "p95_ttft_ms", MetricSample(790, float("nan"), "release-a91", "5m", 2000)),
8 ("nan_count", "p95_ttft_ms", MetricSample(790, 30, "release-a91", "5m", float("nan"))),
9 ("fractional_count", "p95_ttft_ms", MetricSample(790, 30, "release-a91", "5m", 2000.5)),
10 ("boolean_rate", "toxicity_rate", MetricSample(False, 30, "release-a91", "10m", 2000)),
11 ("huge_value", "p95_ttft_ms", MetricSample(10**1000, 30, "release-a91", "5m", 2000)),
12 ("wrong_record_type", "p95_ttft_ms", {"value": 790}),
13]
14for label, metric, record in bad_records:
15 decision = policy.decision(healthy | {metric: record})
16 assert decision == "rollback"
17 print(f"{label}={decision}")1nan_age=rollback
2nan_count=rollback
3fractional_count=rollback
4boolean_rate=rollback
5huge_value=rollback
6wrong_record_type=rollbackThe unavailable-data policy decides these actions; malformed quality-only telemetry would pause under its rule. The illustrative minimum of 100 samples is a validity floor, not statistical power for a rare-event safety claim. Argo Rollouts can query metrics and drive promotion, pause, or abort through AnalysisTemplates. Its success/failure conditions decide how NaN and empty results behave; failure-closed handling must be configured explicitly.[14]
A custom gateway can use the same pattern with immutable release manifests and audited changes to mutable routing pointers. The scorer above doesn't read current runtime state; the executor must check it again before applying its recommendation. Compare the current routing generation as well as the served release with the deployment whose metrics triggered it. A delayed alert for yesterday's canary must not overwrite a newer deployment, even if both deployments use the same artifact digest. The local lab below tests that repeated-release case.
Error-rate or latency-only rollback misses slow quality regressions. A groundedness drop may not trigger a 5xx spike, yet it can damage user trust. Sampled LLM-judge signals can expose that class of regression, but judge drift and sampling uncertainty mean those signals should be calibrated before driving automatic rollback.
Why can alias rollback be faster and safer than rebuilding or redeploying containers?
Answer
If immutable versions and registry history are already in place, the controller can change the production alias without building a new candidate image. The restored version may still require cache warming or artifact loading, so rollback latency must be tested.
For this latency policy, require three consecutive valid windows above the limit. Each value below is a p95 already computed over a separate five-minute window, not one request's latency. The example tests the decision logic only; a severe security event can justify a different immediate-action policy.
1ttft_limit_ms = 850
2windows = [790, 910, 930, 905]
3required_consecutive_breaches = 3
4consecutive = 0
5
6for value in windows:
7 consecutive = consecutive + 1 if value > ttft_limit_ms else 0
8
9 if consecutive >= required_consecutive_breaches:
10 break
11
12rollback = consecutive >= required_consecutive_breaches
13print(f"last_consecutive_breaches={consecutive}")
14print(f"rollback={rollback}")1last_consecutive_breaches=3
2rollback=TrueThe healthy window resets the counter. Three later windows breach, so the policy recommends rollback and stops processing. This latches the decision instead of allowing a later healthy value to erase a rollback that should already have fired. Overlapping windows are correlated, so three rolling measurements aren't three independent experiments.
Look up the served tuple from a trace
Rollback restored an alias, but it doesn't explain the bad response. Provenance tells you which alias and feature fingerprint req_8a2 actually used. Start with one production trace and answer:
- Which exact prompt version and guardrail policy produced it?
- Which model weights and adapter were active?
- Which feature values (with exact timestamps and computation hashes) were fed into the prompt?
- Which golden set version and judge prompt were used to approve this release?
- What was the commit SHA that last touched any of the above?
Keep this information in a composed release manifest linked from request traces, with model-registry versions, feature-store or index audit records, and the Git commit that triggered promotion. During an incident, resolve the served release tuple from the affected request timestamp before guessing at the cause.

What should be traceable for any single production response?
Answer
The exact prompt version, model weights, adapter, guardrail policy, feature values with timestamps, feature computation hashes, eval gate, judge prompt, serving image, and git commit that promoted the release.
Resolve req_8a2 to release-a91, then diff it against last-known-good. Predict the result: prompt and adapter should match, while the feature fingerprint should point to the restore target releases@4.
1release_manifests = {
2 "release-a91": {
3 "prompt": "[email protected]",
4 "base": "base@42",
5 "adapter": "[email protected]",
6 "feature_view": "releases@5",
7 "normalize": "none",
8 "fingerprint_display": "660645c4",
9 "git_commit": "9f3c2a1",
10 },
11 "last-known-good": {
12 "prompt": "[email protected]",
13 "base": "base@42",
14 "adapter": "[email protected]",
15 "feature_view": "releases@4",
16 "normalize": "l2",
17 "fingerprint_display": "ce667443",
18 "git_commit": "a7d812b",
19 },
20}
21trace = {"request_id": "req_8a2", "release_id": "release-a91", "region": "us-west"}
22served = release_manifests[trace["release_id"]]
23known_good = release_manifests["last-known-good"]
24print(f"request={trace['request_id']} region={trace['region']}")
25print(f"prompt={served['prompt']} adapter={served['adapter']}")
26print(f"served_fp={served['fingerprint_display']} known_good_fp={known_good['fingerprint_display']}")
27print("restore_release=last-known-good")
28print(f"restore_feature_view={known_good['feature_view']}")1request=req_8a2 region=us-west
2[email protected] [email protected]
3served_fp=660645c4 known_good_fp=ce667443
4restore_release=last-known-good
5restore_feature_view=releases@4These abbreviated fingerprints are display-only fixtures. In a real trace, record the resolved immutable release ID and actual feature-value provenance, not just a mutable alias and timestamp. Keep sensitive feature values in access-controlled storage with retention limits. The diff identifies changed preprocessing; restore the complete compatible known-good tuple rather than independently flipping one feature pointer and assuming every other dependency still matches.
Run a release lifecycle locally
Download release_contract_lab.py and run uv run release_contract_lab.py. It uses only Python's standard library and a temporary SQLite database. It rejects mismatched computed vectors, registers a corrected candidate by full digest, changes a routing pointer, resolves a trace, restores the baseline after sustained fixture breaches, and rejects stale controller decisions.
The pointer update uses a compare-and-swap condition on release ID and routing generation. Why both? Candidate A can be rolled back and later redeployed with the same immutable bytes. Its digest identifies the artifact; it doesn't identify which deployment produced an alert. Every local routing change increments the generation, so an old controller token can't match that later deployment. Bind production telemetry to a rollout instance and evaluated window too.
The audit insert and pointer change share one database transaction. This checks a useful local concurrency contract, not distributed atomicity across a registry, Git, and a GPU fleet.
1bad candidate: computed feature parity failed
2approved candidate: routed; trace resolves to its full manifest digest
3rollback: complete baseline tuple restored
4delayed controller: stale routing token; re-read before deciding
5repeated release: same digest, newer routing generation; old decision rejected
6audit: 4 committed transitions; stale attempts changed neither pointer nor audit
7PASS: local contracts only; no infrastructure or model callsThe lab's readiness set is explicitly simulated. A production controller still needs artifact verification, real health probes, rollout acknowledgement, durable audit delivery, Git reconciliation, and a recovery plan when the baseline itself is unhealthy. Rollback also can't undo tool actions already taken by a bad release.
Before treating a green alias as recovery, inspect one new request: does its trace resolve to the restored manifest, its features meet the freshness/parity contract, and its latency and task result meet the recovery gate? Then ask what persists outside that tuple. A corrupted shared index or an already-executed tool action can survive the routing change.