Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
When the incident-evidence assistant answers "How long are revoked API keys retained?", developers will ask it a dozen ways. Reusing one verified answer would be cheaper and faster. Reusing it for "What is incident INC-48192's current status?", after a policy update, or across a tenant boundary would be a confidently wrong cached answer.
The last chapter pinned every answer to one immutable release bundle: model, prompt, policy, corpus, and serving image together. That identity is what makes reuse possible, and also what makes it unsafe when any field moves. The PolicyOps team will cache public-policy answers only, and only after a shadow replay shows safe hits that still save money.
An answer is reusable only inside its contract
A response cache isn't a memory of generally true sentences. It's a store of outputs generated under a particular contract. The last lab's stable release was deploy-answerer@sha256:32b8ed409b8e, pinned to its ReleaseOps evidence corpus. Public API-key policy belongs to a separate policy-answerer release with a different corpus and release identity. A cached sentence from one bundle isn't a cached sentence from the other, even when the question text matches.
| Request | Can an answer be reused? | Why |
|---|---|---|
| "How long are revoked API keys retained?" | Candidate | Public docs-policy answer can remain stable within one policy release. |
| "How long do revoked keys stay in audit logs?" | Candidate | Paraphrase of the same public docs-policy question, after evaluation. |
| "What is incident INC-48192's current status?" | No | Answer depends on live incident state. |
| "Revoke API key KEY-48192." | No | The request asks for a side effect, not a reusable answer. |
Hold two boundaries apart before continuing. If the live incident question lands near the retention question in vector space, it still bypasses. If the policy corpus moves, an exact repeat still misses under the new scope. Similarity can nominate a candidate; the contract decides whether it may serve.
For this system, a reusable answer must match all of these fields:
| Field | Why it matters |
|---|---|
release_id | Pins model, prompt, policy logic, and serving behavior. |
corpus_version | Prevents old policy evidence from surviving a document update. |
tenant_id and access_scope | Prevents one tenant's restricted information from leaking into another response. |
locale and response_schema | Prevents the right content from appearing in the wrong language or output contract. |
Semantic lookup has a second version boundary. The embedding model, vector preprocessing, and index build determine the score distribution. Record them as embedding_model_id and embedding_index_version in the cache policy and hard-gate semantic hits on those fields. If either changes, rebuild the vectors and recalibrate the threshold in shadow mode before serving semantic hits. An old under a new index isn't a safe score, even when the admitted answer text still looks right.

The lab starts from one promoted public-policy release of that same service, plus one admitted answer generated under it.
The first cell only establishes identity. Predict its three labels before running it: the release ID and corpus version should be visible, and the seed answer should have a stable ID that later traces can carry forward.
1from dataclasses import asdict, dataclass, replace
2from hashlib import sha256
3import json
4import math
5
6@dataclass(frozen=True)
7class ReleaseScope:
8 release_id: str
9 corpus_version: str
10 tenant_id: str
11 access_scope: str
12 locale: str
13 response_schema: str
14
15@dataclass(frozen=True)
16class Request:
17 text: str
18 tenant_id: str = "docs-public"
19 access_scope: str = "public-policy"
20 locale: str = "en-US"
21 response_schema: str = "cited-answer-v2"
22 requires_live_data: bool = False
23 writes_state: bool = False
24
25@dataclass(frozen=True)
26class CachedAnswer:
27 answer_id: str
28 source_query: str
29 response: str
30 scope: ReleaseScope
31 response_class: str
32 admission_evidence_id: str
33
34stable_scope = ReleaseScope(
35 release_id="policy-answerer@sha256:df2d4fe7b0c5",
36 corpus_version="api-key-policy-2026-04",
37 tenant_id="docs-public",
38 access_scope="public-policy",
39 locale="en-US",
40 response_schema="cited-answer-v2",
41)
42
43seed_answer = CachedAnswer(
44 answer_id="ans_api_key_retention_30d",
45 source_query="How long are revoked API keys retained?",
46 response="Revoked API keys remain visible in audit logs for 30 days.",
47 scope=stable_scope,
48 response_class="public-policy",
49 admission_evidence_id="eval-public-policy-api-keys-v4",
50)
51
52approved_admission_evidence_ids = {
53 "eval-public-policy-api-keys-v4",
54 "eval-public-policy-rate-limits-v1",
55}
56
57print(f"release_id={stable_scope.release_id}")
58print(f"corpus_version={stable_scope.corpus_version}")
59print(f"seed_answer={seed_answer.answer_id}")1release_id=policy-answerer@sha256:df2d4fe7b0c5
2corpus_version=api-key-policy-2026-04
3seed_answer=ans_api_key_retention_30dFor requests already eligible for answer reuse, an ordinary key-value cache can safely reuse normalized exact repeats as long as its key contains the full scope. It can't see through paraphrasing.
The scope fields must come from application-owned route policy and authenticated context, not from raw user text or a model's guess. The same rule applies to eligibility flags: requires_live_data and writes_state must come from a route or intent-class policy map (or a deterministic classifier behind an admission gate), never from free-form model labels on user text alone. Unknown routes and unknown response classes should bypass answer reuse. The lab also checks that the caller-selected scope agrees with the request before deriving a key.
Predict the next output before running it: the normalized exact repeat should hit, the paraphrase should miss, the updated policy should miss, and the cross-tenant request should raise a scope error. None of those outcomes needs a vector search.
1def normalized_text(text: str) -> str:
2 return " ".join(text.lower().split())
3
4def request_matches_scope(request: Request, scope: ReleaseScope) -> bool:
5 return (
6 request.tenant_id == scope.tenant_id
7 and request.access_scope == scope.access_scope
8 and request.locale == scope.locale
9 and request.response_schema == scope.response_schema
10 )
11
12def exact_key(scope: ReleaseScope, request: Request) -> str:
13 if not request_matches_scope(request, scope):
14 raise ValueError("request does not match cache scope")
15 payload = {
16 "scope": asdict(scope),
17 "text": normalized_text(request.text),
18 }
19 encoded = json.dumps(payload, sort_keys=True).encode("utf-8")
20 return sha256(encoded).hexdigest()
21
22same_words = Request("How long are revoked API keys retained?")
23paraphrase = Request("How long do revoked keys stay in audit logs?")
24cross_tenant = replace(same_words, tenant_id="internal-admin")
25updated_scope = replace(
26 stable_scope,
27 release_id="policy-answerer@sha256:new-policy",
28 corpus_version="api-key-policy-2026-05",
29)
30
31seed_key = exact_key(stable_scope, Request(seed_answer.source_query))
32print(f"exact_repeat_hit={exact_key(stable_scope, same_words) == seed_key}")
33print(f"paraphrase_exact_hit={exact_key(stable_scope, paraphrase) == seed_key}")
34print(f"updated_policy_hit={exact_key(updated_scope, same_words) == seed_key}")
35try:
36 exact_key(stable_scope, cross_tenant)
37except ValueError as error:
38 print(f"cross_tenant_rejected={error}")1exact_repeat_hit=True
2paraphrase_exact_hit=False
3updated_policy_hit=False
4cross_tenant_rejected=request does not match cache scopeThe exact cache does the correct thing: it refuses a paraphrase and refuses an old answer under a new policy release. Semantic caching adds only the first capability. It must not weaken the second.
Similarity retrieves a candidate, not a truth
A semantic cache embeds a new question, searches stored question embeddings, and proposes a nearby saved answer. Systems such as GPTCache put that candidate lookup before the decision to call an LLM at all.[1]
In a larger store, the lookup is commonly an approximate nearest-neighbor (ANN) search. Sentence-BERT showed why this shape works: sentence embeddings can be compared efficiently with cosine similarity for semantic matching tasks.[2]
For two vectors and , cosine similarity is:
The numerator measures their aligned components. Dividing by both lengths makes the result compare direction rather than vector magnitude. A high score says two encoded questions are close under this embedding model. It does not say their answers are interchangeable.
The tiny vectors below are an instructional fixture, not scores from a commercial embedding model. They let us see the failure mode without downloading a model: a restore-key exception can sit near a general retention question while still needing a different answer. Restore-key is still a public-policy read, so a live/write gate won't save you.
Predict the scores before running the cell. The paraphrase should be nearest, the restore-key exception should look dangerously close, and the live incident should be far away. Which of those predictions would be enough to serve an answer? None by itself.
1fixture_vectors = {
2 seed_answer.source_query: (1.00, 0.00, 0.00),
3 "How long do revoked keys stay in audit logs?": (0.99, 0.04, 0.00),
4 "Can I restore a revoked API key?": (0.94, 0.10, 0.00),
5 "What is incident INC-48192's current status?": (0.00, 0.05, 1.00),
6}
7
8def cosine(left: tuple[float, ...], right: tuple[float, ...]) -> float:
9 dot = sum(a * b for a, b in zip(left, right))
10 left_norm = math.sqrt(sum(value * value for value in left))
11 right_norm = math.sqrt(sum(value * value for value in right))
12 return dot / (left_norm * right_norm)
13
14seed_vector = fixture_vectors[seed_answer.source_query]
15for question in [
16 "How long do revoked keys stay in audit logs?",
17 "Can I restore a revoked API key?",
18 "What is incident INC-48192's current status?",
19]:
20 score = cosine(seed_vector, fixture_vectors[question])
21 print(f"{question} | score={score:.3f}")1How long do revoked keys stay in audit logs? | score=0.999
2Can I restore a revoked API key? | score=0.994
3What is incident INC-48192's current status? | score=0.000
💡 Key insight: Cosine similarity ranks nearby questions. It doesn't certify that two answers are interchangeable under the current release, corpus, and access scope.
Live incident state is far from the policy axis here, so a score threshold would reject it. Restore-key wouldn't. Eligibility can catch live state and writes. It can't catch a nearby policy exception.
Eligibility rules run before the score threshold
The assistant shouldn't cache live incident state or write actions at any threshold. Restore-key is different: it's still a public-policy read, so eligibility won't catch it. Even for an eligible question, a cached answer must be from the same release scope, and the serving index must match the embedding identity that calibrated the threshold.
This decision procedure checks the non-negotiable rules first. Route policy owns live/write classification. Only an eligible, same-scope request with a matching embedding policy reaches the similarity threshold. Similarity is last, not first.
Predict the four decisions before running the next cell: the paraphrase should hit; the incident and revoke request should bypass before scoring; the stale index should miss even with a strong score. The order is the safety property.
1@dataclass(frozen=True)
2class CachePolicy:
3 embedding_model_id: str
4 embedding_index_version: str
5 selected_threshold: float
6
7# Pins the embedding identity that produced selected_threshold evidence.
8ACTIVE_CACHE_POLICY = CachePolicy(
9 embedding_model_id="docs-query-embedder-v3",
10 embedding_index_version="api-key-policy-index-2026-04-v2",
11 selected_threshold=0.98,
12)
13
14# Route or intent-class map owns eligibility. Do not trust free-form model labels alone.
15ROUTE_ELIGIBILITY = {
16 "public-policy": {"requires_live_data": False, "writes_state": False},
17 "incident-state": {"requires_live_data": True, "writes_state": False},
18 "key-admin": {"requires_live_data": False, "writes_state": True},
19}
20
21def request_from_route(text: str, access_scope: str) -> Request:
22 route = ROUTE_ELIGIBILITY.get(access_scope)
23 if route is None:
24 # Unknown routes fail closed: treat as live so answer reuse can't apply.
25 return Request(text, access_scope=access_scope, requires_live_data=True)
26 return Request(text, access_scope=access_scope, **route)
27
28def same_scope(request: Request, record: CachedAnswer, active: ReleaseScope) -> bool:
29 return (
30 record.scope == active
31 and request_matches_scope(request, active)
32 )
33
34def record_is_admitted(record: CachedAnswer) -> bool:
35 return (
36 record.response_class == "public-policy"
37 and record.admission_evidence_id in approved_admission_evidence_ids
38 )
39
40def decide_candidate(
41 request: Request,
42 record: CachedAnswer,
43 active: ReleaseScope,
44 score: float,
45 threshold: float,
46 cache_policy: CachePolicy = ACTIVE_CACHE_POLICY,
47) -> str:
48 if request.requires_live_data or request.writes_state:
49 return "BYPASS_DYNAMIC_OR_WRITE"
50 if (
51 cache_policy.embedding_model_id != ACTIVE_CACHE_POLICY.embedding_model_id
52 or cache_policy.embedding_index_version != ACTIVE_CACHE_POLICY.embedding_index_version
53 ):
54 return "MISS_EMBEDDING_VERSION"
55 if not same_scope(request, record, active):
56 return "MISS_SCOPE_CHANGED"
57 if not record_is_admitted(record):
58 return "BYPASS_UNVALIDATED_RECORD"
59 if score < threshold:
60 return "MISS_BELOW_THRESHOLD"
61 return "SEMANTIC_HIT"
62
63policy_paraphrase = request_from_route(
64 "How long do revoked keys stay in audit logs?",
65 "public-policy",
66)
67live_incident = request_from_route(
68 "What is incident INC-48192's current status?",
69 "incident-state",
70)
71revoke_action = request_from_route(
72 "Revoke API key KEY-48192.",
73 "key-admin",
74)
75stale_index_policy = CachePolicy(
76 embedding_model_id="docs-query-embedder-v3",
77 embedding_index_version="api-key-policy-index-2026-03-v1",
78 selected_threshold=0.98,
79)
80
81policy_score = cosine(seed_vector, fixture_vectors[policy_paraphrase.text])
82print(decide_candidate(policy_paraphrase, seed_answer, stable_scope, policy_score, 0.98))
83print(decide_candidate(live_incident, seed_answer, stable_scope, 1.00, 0.98))
84print(decide_candidate(revoke_action, seed_answer, stable_scope, 1.00, 0.98))
85print(decide_candidate(
86 policy_paraphrase, seed_answer, stable_scope, policy_score, 0.98, stale_index_policy,
87))1SEMANTIC_HIT
2BYPASS_DYNAMIC_OR_WRITE
3BYPASS_DYNAMIC_OR_WRITE
4MISS_EMBEDDING_VERSIONThe function returns as soon as a gate fails. Live incidents and key-revocation writes never reach the score. A stale embedding index misses even when the question is a perfect paraphrase of an admitted answer.

A live incident-status request has a similarity score of 1.00 against a saved response. Can it be a semantic answer-cache hit?
Answer
No. Its answer depends on live operational state. Eligibility is checked before similarity, so the request bypasses response reuse regardless of score.
Version changes invalidate answers without guessing
A time-to-live (TTL) can expire old entries after a period. It can't know that an API-key retention policy changed five minutes after an answer was stored. The release bundle provides a stronger invalidation hook: if policy evidence or answer behavior changes, the release or corpus version changes and old entries are no longer in scope.
Before running the invalidation example, predict both decisions: the old scope should hit, while the same question under the new release should miss immediately. Storage eviction can lag behind correctness.
1policy_update = replace(
2 stable_scope,
3 release_id="policy-answerer@sha256:7a12policy",
4 corpus_version="api-key-policy-2026-05",
5)
6
7same_question = Request(seed_answer.source_query)
8old_release_decision = decide_candidate(
9 same_question, seed_answer, stable_scope, 1.00, 0.98
10)
11new_release_decision = decide_candidate(
12 same_question, seed_answer, policy_update, 1.00, 0.98
13)
14
15print(f"old_release={old_release_decision}")
16print(f"new_policy_release={new_release_decision}")
17print(f"new_release_must_generate={new_release_decision != 'SEMANTIC_HIT'}")1old_release=SEMANTIC_HIT
2new_policy_release=MISS_SCOPE_CHANGED
3new_release_must_generate=TrueCache identity should inherit the release identity from deployment. Eviction can clean up storage later; correctness shouldn't depend on eviction finishing first. Keep a bounded TTL as a cleanup policy and backstop, but don't treat it as the authoritative invalidation signal.
Why should changing the model, prompt, evidence snapshot, or authorization policy invalidate cached answers even when query similarity stays high?
Answer
Similarity says the request resembles an old request; it doesn't prove the old answer was produced under the current behavior and access contract. Version those dependencies in cache eligibility.
Choose a threshold in shadow mode
Serving a semantic hit immediately turns a retrieval mistake into a user-visible wrong answer. Shadow mode runs the lookup decision but still serves the normal fresh path. Reviewers then label whether each proposed reuse would have been acceptable.
A good cache metric separates two questions:
- Proposal rate: how often would the cache return something?
- Hit precision: among proposed hits, how often is answer reuse acceptable?
High proposal rate without high precision is a cheaper system that's wrong more often. A raw cache hit rate can't distinguish those outcomes.
The labeled fixture below contains public-policy paraphrases, the restore-key exception, and ineligible requests. These scores are a separate measurement from the 3-d plot. That plot put restore-key at 0.994 so you could see a wrong neighbor inside the 0.98 cone. Here restore-key sits at 0.965, which is how a 0.980 threshold can reject it after you have labels. Don't mix the two score columns, and don't copy either threshold into production. Seven probes are enough to explain the tradeoff, but not enough to authorize a 99% production precision claim.
Pause at the three candidate thresholds. The lowest threshold should admit the most proposals and the wrong exception; the strictest should reject valid paraphrases as well as the exception. The selected value needs to clear precision while preserving useful proposal volume, then earn confidence on a larger replay.
1@dataclass(frozen=True)
2class ShadowProbe:
3 name: str
4 score: float
5 eligible: bool
6 acceptable_reuse: bool
7
8shadow_probes = [
9 ShadowProbe("key retention paraphrase", 0.995, True, True),
10 ShadowProbe("audit log retention wording", 0.989, True, True),
11 ShadowProbe("revoked key wording", 0.982, True, True),
12 ShadowProbe("policy FAQ reworded", 0.981, True, True),
13 ShadowProbe("restore-key exception", 0.965, True, False),
14 ShadowProbe("live incident state", 0.999, False, False),
15 ShadowProbe("revoke key action", 0.997, False, False),
16]
17
18def replay_at(threshold: float) -> dict[str, float | int]:
19 proposed = [
20 probe for probe in shadow_probes
21 if probe.eligible and probe.score >= threshold
22 ]
23 accepted = [probe for probe in proposed if probe.acceptable_reuse]
24 precision = len(accepted) / len(proposed) if proposed else 1.0
25 return {
26 "proposed": len(proposed),
27 "accepted": len(accepted),
28 "precision": precision,
29 "proposal_rate": len(proposed) / len(shadow_probes),
30 }
31
32for threshold in [0.960, 0.980, 0.990]:
33 metrics = replay_at(threshold)
34 print(
35 f"threshold={threshold:.3f} "
36 f"proposed={metrics['proposed']} "
37 f"precision={metrics['precision']:.1%} "
38 f"proposal_rate={metrics['proposal_rate']:.1%}"
39 )
40
41selected_threshold = 0.980
42
43def wilson_lower_bound(successes: int, trials: int, z: float = 1.96) -> float:
44 if trials == 0:
45 return 0.0
46 observed = successes / trials
47 denominator = 1 + z**2 / trials
48 center = observed + z**2 / (2 * trials)
49 margin = z * math.sqrt(
50 observed * (1 - observed) / trials + z**2 / (4 * trials**2)
51 )
52 return (center - margin) / denominator
53
54# Synthetic full-window fixture at the selected threshold. In production,
55# populate these counts from representative labeled shadow traffic.
56full_shadow_requests = 10_000
57full_shadow_proposed = 5_714
58full_shadow_accepted = 5_714
59full_shadow_precision = full_shadow_accepted / full_shadow_proposed
60precision_lower_bound = wilson_lower_bound(
61 full_shadow_accepted,
62 full_shadow_proposed,
63)
64safe_hit_fraction = full_shadow_accepted / full_shadow_requests
65
66print(
67 f"full_shadow proposed={full_shadow_proposed} "
68 f"precision={full_shadow_precision:.1%} "
69 f"precision_lower_bound={precision_lower_bound:.2%}"
70)1threshold=0.960 proposed=5 precision=80.0% proposal_rate=71.4%
2threshold=0.980 proposed=4 precision=100.0% proposal_rate=57.1%
3threshold=0.990 proposed=1 precision=100.0% proposal_rate=14.3%
4full_shadow proposed=5714 precision=100.0% precision_lower_bound=99.93%
The larger fixture keeps the same 0.980 policy but evaluates 5,714 proposed hits from 10,000 requests. All are accepted, so observed precision is 100%. The 95% Wilson lower bound is 99.93%, above the 99% gate. A confidence bound prevents a tiny perfect sample from looking more reliable than the sample supports. The traffic still needs to represent the production mix; statistical confidence can't repair a biased replay.
Why doesn't a high hit rate prove that semantic caching is helping?
Answer
A hit is only beneficial when the reused answer is acceptable under the active release and scope. A loose threshold can raise hit rate by serving nearby but incorrect answers.
A safe cache still has to pay for itself
Every semantic lookup incurs work, even on a miss: embedding the request, searching an index, and recording metrics. Evaluate cost only after the precision gate passes.
Let:
- be requests in a measured period.
- be average fresh-generation cost per request.
- be semantic-lookup cost per request.
- be the observed safe-hit fraction.
If a hit skips fresh generation, expected period savings are:
These quantities must come from the workload and model you plan to operate. The next example uses labeled measurement fixtures, not provider prices.
Compute the break-even point before running the cell. With and , . The fixture's safe-hit fraction is 57.1%, so its savings should be positive, but that conclusion follows only after safe-hit precision has passed.
1requests_per_day = full_shadow_requests
2fresh_generation_usd = 0.0040 # measured fixture: average full answer cost
3semantic_lookup_usd = 0.00008 # measured fixture: embed + index lookup
4
5without_cache = requests_per_day * fresh_generation_usd
6with_cache = requests_per_day * (
7 semantic_lookup_usd + (1 - safe_hit_fraction) * fresh_generation_usd
8)
9savings = without_cache - with_cache
10break_even_hit_fraction = semantic_lookup_usd / fresh_generation_usd
11
12print(f"safe_hit_fraction={safe_hit_fraction:.1%}")
13print(f"break_even_hit_fraction={break_even_hit_fraction:.1%}")
14print(f"daily_savings_fixture_usd={savings:.2f}")
15print(f"savings_positive={savings > 0}")1safe_hit_fraction=57.1%
2break_even_hit_fraction=2.0%
3daily_savings_fixture_usd=22.06
4savings_positive=TrueDon't guess from list price. Measure generation and lookup cost for the real release and traffic mix, then rerun the gate when either changes.

Those savings only count if the reused answers were allowed into the index. A cheap wrong answer isn't a win.
Authorize writes before they become reusable
The read path now rejects records without approved admission evidence. The write path must enforce the same rule before adding a record to the servable index. A freshly generated answer isn't automatically safe to repeat across paraphrases.
Treat admission as write authorization. Keep unreviewed answers in a quarantine store, attach the evaluation artifact that approved a response class, and admit only records inside the tested release scope.
Predict the write decisions: missing evidence should quarantine, approved public-policy evidence should admit, and a live-incident response class should quarantine. A fresh answer earns reuse only after this gate.
1def admission_decision(answer: CachedAnswer) -> str:
2 if answer.response_class != "public-policy":
3 return "QUARANTINE_RESPONSE_CLASS"
4 if answer.scope != stable_scope:
5 return "QUARANTINE_SCOPE_CHANGED"
6 if answer.admission_evidence_id not in approved_admission_evidence_ids:
7 return "QUARANTINE_MISSING_EVIDENCE"
8 return "ADMIT_SERVABLE"
9
10unreviewed_answer = replace(
11 seed_answer,
12 answer_id="ans_restore_key_review",
13 admission_evidence_id="",
14)
15validated_policy_answer = replace(
16 seed_answer,
17 answer_id="ans_rate_limit_public",
18 source_query="What is the default API rate limit?",
19 response="Default API keys allow 600 requests per minute unless the account policy says otherwise.",
20 admission_evidence_id="eval-public-policy-rate-limits-v1",
21)
22dynamic_incident_answer = replace(
23 seed_answer,
24 answer_id="ans_live_incident_status",
25 response_class="live-incident-status",
26)
27
28for label, answer in [
29 ("unreviewed", unreviewed_answer),
30 ("validated_policy", validated_policy_answer),
31 ("dynamic_incident", dynamic_incident_answer),
32]:
33 print(f"{label}={admission_decision(answer)}")1unreviewed=QUARANTINE_MISSING_EVIDENCE
2validated_policy=ADMIT_SERVABLE
3dynamic_incident=QUARANTINE_RESPONSE_CLASSThe admitted record still isn't a universal truth. Reads must match its release and access scope, then pass the calibrated semantic threshold. Admission prevents one bad fresh generation from silently becoming a high-fanout cached answer.
Promote only the narrow policy you tested
Don't turn on semantic caching for every route PolicyOps owns. Promote only the public-policy scope that passed shadow evidence. Account state, live incidents, and write actions still bypass.
Read the promotion gate as three independent questions. Predict the output before running it: quality, economics, and scope should all pass for the tested public-policy class, while untested dynamic and write routes remain outside promotion.
1@dataclass(frozen=True)
2class CachePromotionGate:
3 minimum_precision_lower_bound: float
4 minimum_daily_savings_usd: float
5 required_scope: ReleaseScope
6
7gate = CachePromotionGate(
8 minimum_precision_lower_bound=0.99,
9 minimum_daily_savings_usd=5.00,
10 required_scope=stable_scope,
11)
12
13passes_quality = (
14 precision_lower_bound >= gate.minimum_precision_lower_bound
15)
16passes_economics = savings >= gate.minimum_daily_savings_usd
17passes_scope = seed_answer.scope == gate.required_scope
18decision = (
19 "PROMOTE_PUBLIC_POLICY_SEMANTIC_CACHE"
20 if passes_quality and passes_economics and passes_scope
21 else "KEEP_SHADOW_ONLY"
22)
23
24print(f"quality_gate={passes_quality}")
25print(f"economics_gate={passes_economics}")
26print(f"scope_gate={passes_scope}")
27print(f"cache_decision={decision}")1quality_gate=True
2economics_gate=True
3scope_gate=True
4cache_decision=PROMOTE_PUBLIC_POLICY_SEMANTIC_CACHEThat promotion is the artifact the next chapter's cost ledger consumes: public-policy hits skip generation, and everything else still has to be priced.
Record why each request hit or bypassed
Once the cache is serving, a trace has to reconstruct a hit: which release generated the stored answer, which cache policy reused it, and why a neighboring request bypassed. Without those fields, a wrong-hit incident is a ghost.
1def trace_decision(request: Request, score: float) -> dict[str, str | float]:
2 cache_decision = decide_candidate(
3 request, seed_answer, stable_scope, score, selected_threshold
4 )
5 return {
6 "request": request.text,
7 "release_id": stable_scope.release_id,
8 "corpus_version": stable_scope.corpus_version,
9 "tenant_id": request.tenant_id,
10 "access_scope": request.access_scope,
11 "cache_policy": "public-policy-semantic-v1",
12 "embedding_model_id": ACTIVE_CACHE_POLICY.embedding_model_id,
13 "embedding_index_version": ACTIVE_CACHE_POLICY.embedding_index_version,
14 "answer_id": seed_answer.answer_id if cache_decision == "SEMANTIC_HIT" else "",
15 "admission_evidence_id": seed_answer.admission_evidence_id,
16 "decision": cache_decision,
17 "score": score,
18 }
19
20hit_trace = trace_decision(policy_paraphrase, policy_score)
21bypass_trace = trace_decision(live_incident, 1.00)
22
23print(f"hit_decision={hit_trace['decision']} answer_id={hit_trace['answer_id']}")
24print(f"bypass_decision={bypass_trace['decision']}")
25print(f"traced_release={hit_trace['release_id'] == stable_scope.release_id}")
26print(f"traced_scope={hit_trace['corpus_version'] == stable_scope.corpus_version and hit_trace['access_scope'] == stable_scope.access_scope}")
27print(f"traced_index={hit_trace['embedding_index_version'] == ACTIVE_CACHE_POLICY.embedding_index_version}")
28print(f"traced_admission={hit_trace['admission_evidence_id'] == seed_answer.admission_evidence_id}")
29print(f"traced_model={hit_trace['embedding_model_id'] == ACTIVE_CACHE_POLICY.embedding_model_id}")1hit_decision=SEMANTIC_HIT answer_id=ans_api_key_retention_30d
2bypass_decision=BYPASS_DYNAMIC_OR_WRITE
3traced_release=True
4traced_scope=True
5traced_index=True
6traced_admission=True
7traced_model=TrueWatch production for accepted-hit review failures, user retries after cache hits, scope bypass volume, p95 latency, and realized saved generation. An incident has to disable this cache policy pointer without changing the production release that generates fresh responses.
Those signals also point to different repairs. A wrong answer after a valid hit sends you to admission, scope, or threshold evidence; rising bypass volume sends you to route policy or version churn; savings below expectation sends you to lookup cost, miss rate, or traffic mix. Keep the decision trace that lets an operator tell those cases apart.
Semantic response caching isn't prompt-prefix caching
The cache in this lab can return a stored answer for a paraphrase and skip generation. Provider prompt caching operates at a different layer: it reuses work for an identical rendered prompt prefix, then generates a new answer. OpenAI says prompt caching is enabled by default for supported models. Current docs list a minimum of 1,024 visible input tokens for GPT-5.6 and later and 2,048 for earlier models, with occasional shorter hits. GPT-5.6 and later support explicit or implicit cache breakpoints, so a shared shorter prefix isn't automatically reusable unless it ends at an eligible breakpoint.[3]
| Layer | Matches | Result on hit | Main correctness risk |
|---|---|---|---|
| Exact response cache | Same scoped request key | Return stored answer, skip generation | Stale or incomplete scope key |
| Semantic response cache | Similar eligible question under same contract | Return stored answer, skip generation | False semantic reuse |
| Provider prompt cache | Matching input prefix under provider rules | Compute a new answer with cheaper/faster repeated input work | Missed cost opportunity, not stored-answer substitution |
The distinction determines the evaluation: semantic answer caching needs labeled reuse precision; prompt-prefix caching needs token and latency accounting. The next chapter expands that economics.
Predict the final two routes before running the last cell. A paraphrased public-policy question should skip answer generation but pay for semantic lookup; a live incident with the same long prefix may reuse input work while still generating a fresh response. Same word "cache," different work and different correctness boundary.
1@dataclass(frozen=True)
2class ReuseCase:
3 name: str
4 semantic_answer_hit: bool
5 repeated_prefix_hit: bool
6
7cases = [
8 ReuseCase(
9 name="paraphrased public API-key question",
10 semantic_answer_hit=True,
11 repeated_prefix_hit=False,
12 ),
13 ReuseCase(
14 name="new live incident question after same long instructions",
15 semantic_answer_hit=False,
16 repeated_prefix_hit=True,
17 ),
18]
19
20for case in cases:
21 print(
22 f"{case.name}: "
23 f"skip_generation={case.semantic_answer_hit}, "
24 f"reuse_input_work={case.repeated_prefix_hit}"
25 )
26
27print("next_measure_token_economics=True")1paraphrased public API-key question: skip_generation=True, reuse_input_work=False
2new live incident question after same long instructions: skip_generation=False, reuse_input_work=True
3next_measure_token_economics=True