LeetLLM
My PlanLearnGlossaryTracksPracticeBlog
LeetLLM

Your go-to resource for mastering AI & LLM systems.

Product

  • Learn
  • Glossary
  • Tracks
  • Practice
  • Blog
  • RSS

Legal

  • Terms of Service
  • Privacy Policy

© 2026 LeetLLM. All rights reserved.

All Topics
Your Progress
0%

0 of 178 articles completed

🛠️Computing Foundations0/9
Git, Shell, Linux for AIDocker for Reproducible AIPython for AI EngineeringNumPy and Tensor ShapesCUDA for ML TrainingMPS & Metal for ML on MacData Structures for AISQL and Data ModelingAlgorithms for ML Engineers
📊Math & Statistics0/8
Gradients and BackpropVectors, Matrices & TensorsLinear Algebra for MLAdam, Momentum, SchedulersProbability for Machine LearningStatistics and UncertaintyDistributions and SamplingHypothesis Tests, Intervals, and pass@k
📚Preparation & Prerequisites0/13
Neural Networks from ScratchCNNs from ScratchTraining & BackpropagationSoftmax, Cross-Entropy & OptimizationRNNs, LSTMs, GRUs, and Sequence ModelingAutoencoders and VAEsThe Transformer Architecture End-to-EndLanguage Modeling & Next TokensFrom GPT to Modern LLMsPrompt Engineering FundamentalsCalling LLM APIs in ProductionFirst AI App End-to-EndThe LLM Lifecycle
🧮ML Algorithms & Evaluation0/11
Linear Regression from ScratchLogistic Regression and MetricsDecision Trees, Forests, and BoostingReinforcement Learning BasicsValidation and LeakageClustering and PCACore Retrieval AlgorithmsDecoding AlgorithmsExperiment Design and A/B TestingPyTorch Training LoopsDataset Pipelines and Data Quality
📦Production ML Systems0/6
Feature Engineering for Production MLBatch and Streaming Feature PipelinesGradient Boosted Trees in ProductionRanking and Recommendation SystemsForecasting and Anomaly DetectionMonitoring Predictive Models
🧪Core LLM Foundations0/8
The Bitter Lesson & ComputeBPE, WordPiece, and SentencePieceStatic to Contextual EmbeddingsPerplexity & Model EvaluationFile Ingestion for AIChunking StrategiesLLM Benchmarks & LimitationsInstruction Tuning & Chat Templates
🧰Applied LLM Engineering0/24
Dimensionality Reduction for EmbeddingsCoT, ToT & Self-Consistency PromptingFunction Calling & Tool UseMCP & Tool Protocol StandardsContext EngineeringPrompt Injection DefenseResponsible AI GovernanceData Labeling and Human FeedbackEvaluating AI AgentsProduction RAG PipelinesHybrid Search: Dense + SparseReranking and Cross-Encoders for RAGRAG Evaluation for Reliable AnswersLLM-as-a-Judge EvaluationBias & Fairness in LLMsHallucination Detection & MitigationLLM Observability & MonitoringExperiment Tracking with MLflow and W&BPrompt Optimization with DSPyModel Versioning & DeploymentSemantic Caching & Cost OptimizationLLM Cost Engineering & Token EconomicsModel Gateways, Routing, and FallbacksDesign an Automated Support Agent
🎓Portfolio Capstones0/9
Capstone: Delivery ETA PredictionCapstone: Product RankingCapstone: Demand ForecastingCapstone: Image Damage ClassifierCapstone: Production ML PipelineCapstone: Document QACapstone: Eval DashboardCapstone: Fine-Tuned ClassifierCapstone: Reproducible ML Study
🧠Transformer Deep Dives0/8
Sentence Embeddings & Contrastive LossEmbedding Similarity & QuantizationScaled Dot-Product AttentionVision Transformers and Image EncodersPositional Encoding: RoPE & ALiBiLayer Normalization: Pre-LN vs Post-LNMechanistic InterpretabilityDecoding Strategies: Greedy to Nucleus
🧬Advanced Training & Adaptation0/17
Scaling Laws & Compute-Optimal TrainingPre-training Data at ScaleBuild GPT from Scratch LabJAX for PyTorch ResearchersContinued Pretraining for Domain ShiftSynthetic Data PipelinesSupervised Fine-Tuning PipelineMixed Precision TrainingDistributed Training: FSDP & ZeROLoRA & Parameter-Efficient TuningTraining Run OperationsReward Modeling from Preference DataRLHF & DPO AlignmentConstitutional AI & Red TeamingRLVR & Verifiable RewardsKnowledge Distillation for LLMsModel Merging and Weight Interpolation
🤖Advanced Agents & Retrieval0/16
Vector DB Internals: HNSW & IVFAdvanced RAG: HyDE & Self-RAGGraphRAG & Knowledge GraphsRAG Security & Access ControlStructured Output GenerationReAct & Plan-and-ExecuteGuardrails & Safety FiltersCode Generation & SandboxingComputer-Use / GUI / Browser AgentsHuman-in-the-Loop Agent ArchitectureAI Coding Workflow with AgentsAgent Memory & PersistenceAgent Failure & RecoveryRecursive Language Models (RLM)Multi-Agent OrchestrationCapstone: Production Agent
⚡Inference & Production Scale0/19
Inference: TTFT, TPS & KV CacheMulti-Query & Grouped-Query AttentionKV Cache & PagedAttentionPrefix Caching and Prompt CachingFlashAttention & Memory EfficiencyContinuous Batching & SchedulingScaling LLM InferenceModel Parallelism for LLM InferenceModel Quantization: GPTQ, AWQ & GGUFLocal LLM DeploymentSLM Specialization & Edge DeploymentSpeculative DecodingLong Context Window ManagementMixture of Experts ArchitectureMamba & State Space ModelsReasoning & Test-Time ComputeAdvanced MLOps & DevOps for AIGPU Serving & AutoscalingA/B Testing for LLMs
🏗️System Design Capstones0/9
Content Moderation SystemCode Completion SystemMulti-Tenant LLM PlatformLLM-Powered Search EngineVision-Language Models & CLIPMultimodal LLM ArchitectureDiffusion Models: Images & TextReal-Time Voice AI AgentReasoning Agent System Design
🎤AI Lab Interviewing0/4
AI Lab Coding Interview: Python SystemsAI Lab System Design InterviewAI Lab Behavioral InterviewAI Lab Technical Presentation
🔬Project Deep Dives0/17
Deep Dive - vLLMDeep Dive - SkyRLDeep Dive - FlashAttentionDeep Dive - FlashInferDeep Dive - DeepGEMMDeep Dive - NCCLDeep Dive - MegatronDeep Dive - DeepSpeedDeep Dive - RayDeep Dive - MLflowDeep Dive - PyTorchDeep Dive - TransformersDeep Dive - SGLangDeep Dive - slimeDeep Dive - DeepEPDeep Dive - TinkerDeep Dive - Light-PEFT
Back to Topics
LearnPortfolio CapstonesCapstone: Product Ranking
📊HardEvaluation & Benchmarks

Capstone: Product Ranking

Ship a marketplace ranking candidate with eligible retrieval, separate recall and NDCG gates, replayable exposure rows, and an A/B-ready rollback receipt.

17 min read
Learning path
Step 81 of 178 in the full curriculum
Capstone: Delivery ETA PredictionCapstone: Demand Forecasting

Personalize this lesson

Adapt explanations and teaching visuals to your background and preferred voice.

A shopper types label-printer. The ranker decides which products sit near the top, and those positions then collect the clicks you'll be tempted to treat as labels.

Capstone: Delivery ETA Prediction shipped a delay warning from time-safe features and a receipt that only qualified shadow traffic. A search ranker is a different kind of service. Ranking and Recommendation Systems already separated eligibility, candidate recall, and NDCG on a docs corpus. Package that funnel into a marketplace artifact. Stock, region, and policy have to win before any learned score. Impressions have to freeze before outcomes arrive. An offline NDCG win still isn't a launch. Ship a candidate that may reorder eligible, in-stock listings, and that may not surface blocked listings or call clicks an unbiased truth signal.

Label-printer retrieval scores before ranking. P8 at 0.98 is rejected for policy, P11 at 0.97 is rejected for region, and eligible P5, P6, and P4 enter ranking. Control orders them P4, P5, P6 with NDCG at 3 of 0.736. Treatment orders P5, P6, P4 with NDCG at 3 of 1.000.
For label-printer, the two strongest raw matches never enter ranking: P8 is policy-blocked (and also out of stock) and P11 can't ship to the shopper's region. The ranker only reorders the three survivors. Control puts weak P4 first (NDCG@3 0.736); treatment restores the judged order (NDCG@3 1.000).

Specify the ranking surface

Pin the product before you pick a model:

ContractDecision
surfacemarketplace search results
query setfrozen judged search queries plus online experiment traffic
eligibilityin-stock, deliverable region, policy-approved listing
offline metricrecall@K and NDCG@K on the fixture (this lesson uses K=3K=3K=3; production targets are often recall@100 and NDCG@10)
online primary metricpurchase conversion per search
guardrailsreturns, latency, unsafe listings, seller concentration

Candidate generation and reranking stay separate artifacts. The candidate component has to recover relevant eligible items. The ranker can only change order inside that set.

The earlier ranking lesson used pairwise preferences the way RankNet does: a relevant product should outscore a weaker candidate.[1]Reference 1Learning to Rank using Gradient Descent.https://www.microsoft.com/en-us/research/publication/learning-to-rank-using-gradient-descent/ That modeling choice still doesn't let a learned score override eligibility. Filter first, then score, and pin the eligibility version on every impression log.

Diagram showing Catalog snapshot, Eligible set recall@K, Ranker v1 NDCG@K, and Displayed slate impression log.
Catalog snapshot, Eligible set recall@K, Ranker v1 NDCG@K, and Displayed slate impression log.

Build offline evidence first

A live shopper log can't grade that contract. Start with a judged fixture: each query has eligible candidates and graded relevance from human review.

text
1ranking-product/ 2 data/ 3 catalog_snapshot.jsonl 4 judged_queries.jsonl 5 eligibility_policy.json 6 retrieval/ 7 candidate_generator.py 8 ranking/ 9 ranker.py 10 evaluate_ndcg.py 11 experiments/ 12 impression_schema.json 13 ab_plan.md 14 tests/ 15 test_blocked_listing_never_surfaces.py 16 test_ndcg_regression_gate.py

Offline gates in this lesson score NDCG@3 / recall@3 on the small judged fixture. Production contracts often pin NDCG@10 and recall@100; keep the K explicit so the receipt and the surface table match.

Required offline checks:

CheckWhy it blocks release
eligible-only resultsa relevant prohibited listing still can't display
candidate recall@K (fixture @3; prod often @100)the ranker can't repair missing products
NDCG@K by query category (fixture @3; prod often @10)overall improvement may hide poor critical categories
scoring latencya slower list damages shopping experience
diversity/seller concentrationone seller with many feedback events shouldn't crowd out catalog

Those gates only mean something if blocked listings never enter scoring. Start there.

Filter policy before retrieval

A prohibited or sold-out product can still have an excellent text match and a high learned score. It still can't reach the ranker. Eligibility is the boundary a model can't learn safely from clicks.

The compact fixture below keeps two judged searches in memory. Production might retrieve 100 candidates per query; this local receipt uses 3 so each listing stays visible.

01-marketplace-eligibility-contract.py
1from collections import Counter 2from dataclasses import dataclass 3from datetime import datetime 4from math import ceil, log2 5 6@dataclass(frozen=True) 7class Listing: 8 product_id: str 9 query: str 10 relevance: int 11 retrieval_score: float 12 baseline_rank_score: float 13 candidate_rank_score: float 14 in_stock: bool 15 policy_approved: bool 16 deliverable_region: bool 17 seller_id: str 18 19CATALOG = [ 20 Listing("P1", "insulated-bag", 3, 0.96, 0.82, 0.97, True, True, True, "S1"), 21 Listing("P2", "insulated-bag", 1, 0.88, 0.75, 0.62, True, True, True, "S2"), 22 Listing("P3", "insulated-bag", 2, 0.84, 0.65, 0.90, True, True, True, "S3"), 23 Listing("P9", "insulated-bag", 3, 0.99, 0.99, 0.99, True, False, True, "S9"), 24 Listing("P10", "insulated-bag", 2, 0.95, 0.98, 0.96, False, True, True, "S1"), 25 Listing("P4", "label-printer", 1, 0.74, 0.90, 0.45, True, True, True, "S4"), 26 Listing("P5", "label-printer", 3, 0.93, 0.72, 0.96, True, True, True, "S5"), 27 Listing("P6", "label-printer", 2, 0.86, 0.65, 0.84, True, True, True, "S6"), 28 Listing("P8", "label-printer", 3, 0.98, 0.99, 0.99, False, False, True, "S8"), 29 Listing("P11", "label-printer", 2, 0.97, 0.97, 0.97, True, True, False, "S11"), 30] 31 32QUERIES = ("insulated-bag", "label-printer") 33 34def exclusion_reason(listing: Listing) -> str | None: 35 if not listing.policy_approved: 36 return "policy_block" 37 if not listing.in_stock: 38 return "out_of_stock" 39 if not listing.deliverable_region: 40 return "unavailable_region" 41 return None 42 43print("catalog listings:", len(CATALOG)) 44print("queries:", QUERIES)
Output
1catalog listings: 10 2queries: ('insulated-bag', 'label-printer')
02-filter-and-retrieve-candidates.py
1def retrieve(query: str, budget: int = 3) -> list[Listing]: 2 eligible = [ 3 listing for listing in CATALOG 4 if listing.query == query and exclusion_reason(listing) is None 5 ] 6 return sorted(eligible, key=lambda listing: (-listing.retrieval_score, listing.product_id))[:budget] 7 8def relevant_ids(query: str) -> set[str]: 9 return { 10 listing.product_id for listing in CATALOG 11 if listing.query == query and exclusion_reason(listing) is None and listing.relevance >= 2 12 } 13 14def candidate_recall(query: str, candidates: list[Listing]) -> float: 15 relevant = relevant_ids(query) 16 if not relevant: 17 raise ValueError(f"{query}: no relevant eligible judgments") 18 return len(relevant & {listing.product_id for listing in candidates}) / len(relevant) 19 20retrieved = {query: retrieve(query) for query in QUERIES} 21rejected = { 22 listing.product_id: exclusion_reason(listing) 23 for listing in CATALOG 24 if exclusion_reason(listing) is not None 25} 26 27for query, candidates in retrieved.items(): 28 print(query, [listing.product_id for listing in candidates], f"recall@3={candidate_recall(query, candidates):.3f}") 29print("rejected:", rejected)
Output
1insulated-bag ['P1', 'P2', 'P3'] recall@3=1.000 2label-printer ['P5', 'P6', 'P4'] recall@3=1.000 3rejected: {'P9': 'policy_block', 'P10': 'out_of_stock', 'P8': 'policy_block', 'P11': 'unavailable_region'}

Products P9 and P8 have the strongest scores for their searches, but policy excludes them. P8 is also out of stock; exclusion_reason() logs the first failing check, so the receipt shows policy_block. Product P10 is approved but sold out. Product P11 can't ship to the shopper's region. Filtering first stops the ranker from treating a business invariant as a preference it can trade away.

Measure retrieval and ordering separately

Candidate recall asks whether relevant eligible products reached the ranker. NDCG then asks whether stronger judgments sit near the top of the returned list. A ranker can't repair a relevant product that retrieval dropped.

Use the same exponential-gain form as the ranking chapter. At cutoff KKK, discounted cumulative gain is

DCG@K=∑i=1K2reli−1log⁡2(i+1)\mathrm{DCG}@K = \sum_{i=1}^{K} \frac{2^{rel_i}-1}{\log_2(i+1)}DCG@K=∑i=1K​log2​(i+1)2reli​−1​

where relirel_ireli​ is the judged grade at 1-based position iii. Position 1 has log⁡22=1\log_2 2 = 1log2​2=1, so it keeps full weight. NDCG divides by the DCG of the ideal ordering of that same candidate set. Manning, Raghavan, and Schütze document this exponential-gain form; some libraries use linear gain reli/log⁡2(i+1)rel_i / \log_2(i+1)reli​/log2​(i+1) instead, so the eval setup has to pin one.[2]Reference 2Introduction to Information Retrieval.https://nlp.stanford.edu/IR-book/ In the helper below, enumerate is 0-based, so the denominator is log2(rank + 2), which equals log⁡2(i+1)\log_2(i+1)log2​(i+1).

Work label-printer by hand before looking at code. Control shows P4, P5, P6 with grades 111, 333, and 222:

DCG=11+7log⁡23+32≈1+4.417+1.500=6.917\mathrm{DCG} = \frac{1}{1} + \frac{7}{\log_2 3} + \frac{3}{2} \approx 1 + 4.417 + 1.500 = 6.917DCG=11​+log2​37​+23​≈1+4.417+1.500=6.917

The ideal order is P5, P6, P4 with grades 333, 222, and 111:

IDCG=7+3log⁡23+0.5≈7+1.893+0.500=9.393\mathrm{IDCG} = 7 + \frac{3}{\log_2 3} + 0.5 \approx 7 + 1.893 + 0.500 = 9.393IDCG=7+log2​33​+0.5≈7+1.893+0.500=9.393

So control NDCG@3 is 6.917/9.393≈0.7366.917 / 9.393 \approx 0.7366.917/9.393≈0.736. Treatment is already the ideal order, so its NDCG@3 is 1.0001.0001.000. Recall@3 is still 1.0001.0001.000 on both arms because P5 and P6 (the eligible grades ≥2\ge 2≥2) are in the candidate set either way. Ordering quality moved; retrieval didn't.

The next cells rerank each eligible set, print those NDCG values, and recheck that scoring didn't sneak a blocked listing back in.

03-ndcg-ranking-helpers.py
1def dcg(rows: list[Listing]) -> float: 2 return sum((2**listing.relevance - 1) / log2(rank + 2) for rank, listing in enumerate(rows)) 3 4def ndcg(rows: list[Listing]) -> float: 5 ideal = sorted(rows, key=lambda listing: (-listing.relevance, listing.product_id)) 6 ideal_dcg = dcg(ideal) 7 return dcg(rows) / ideal_dcg if ideal_dcg else 0.0 8 9def rerank(rows: list[Listing], score_field: str) -> list[Listing]: 10 return sorted(rows, key=lambda listing: (-getattr(listing, score_field), listing.product_id)) 11 12baseline_ranked = { 13 query: rerank(candidates, "baseline_rank_score") 14 for query, candidates in retrieved.items() 15} 16candidate_ranked = { 17 query: rerank(candidates, "candidate_rank_score") 18 for query, candidates in retrieved.items() 19} 20 21print("queries ranked:", list(candidate_ranked))
04-compare-retrieval-and-ranking.py
1for query in QUERIES: 2 print( 3 query, 4 f"baseline_ndcg@3={ndcg(baseline_ranked[query]):.3f}", 5 f"candidate_ndcg@3={ndcg(candidate_ranked[query]):.3f}", 6 ) 7 8blocked_hits = [ 9 listing.product_id 10 for rows in candidate_ranked.values() 11 for listing in rows 12 if exclusion_reason(listing) is not None 13] 14print("blocked hits:", blocked_hits)
Output
1insulated-bag baseline_ndcg@3=0.972 candidate_ndcg@3=1.000 2label-printer baseline_ndcg@3=0.736 candidate_ndcg@3=1.000 3blocked hits: []

The candidate improves ordering on both fixture queries. Two edge policies stay explicit: a query with no relevant eligible judgments is a fixture error here, and NDCG returns 0.0 when a ranked set has no positive gain. That still isn't a launch decision. Offline judgments only qualify a controlled experiment, and the displayed slate is the data you'll later be tempted to train on.

Log exposure before learning from outcomes

If that slate is going to become training data, freeze what was shown before any click arrives.

Every displayed slate should log:

FieldWhy
request_id, query, served_atgroup one displayed decision
shopper_id, experiment_id, experiment_armreplay stable experiment assignment
catalog_snapshot, eligibility_versionshow available choices
candidate_version, ranker_versionidentify scoring path
product_id, positionpreserve exposure
position_propensity, propensity_model_versionenable inverse propensity scoring (IPS) on implicit feedback

Clicks alone aren't reliable targets. Top-ranked items get attention because they're top-ranked. That's position bias. Joachims, Swaminathan, and Schnabel show that using raw clicks as learning-to-rank labels is biased, and they derive a propensity-weighted IPS correction for that implicit feedback.[3]Reference 3Unbiased Learning-to-Rank with Biased Feedbackhttps://arxiv.org/abs/1608.04468 Their method still needs a real propensity model, often from a small randomized logging study, not a made-up curve. A formula such as 1 / position isn't a measured propensity. When training_use=implicit, refuse to train unless the impression row carries the actual randomized action probability or a calibrated examination propensity with a pinned model version.

Use the judged fixture to catch regressions. Use a controlled A/B experiment to compare shopper outcomes. Experiment Design and A/B Testing already made the assignment unit, primary metric, and guardrails explicit; Kohavi, Tang, and Xu treat that online comparison as the way to measure a product change.[4]Reference 4Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing.https://experimentguide.com/ Keep each shopper in one arm. A conversion lift that raises returns or policy violations isn't a successful ranking release.

Clicks, purchases, and returns arrive later. Append them as outcome events with event_id, occurred_at, request_id, product_id, and outcome. Don't overwrite the served impression row.

The remaining cells publish replayable impression rows, two joined outcome events, and one candidate receipt. They don't invent live lift. They prove offline evidence, exposure schema, stable assignment, outcome join, and a rollback pointer are ready before traffic starts.

05-experiment-arm-config.py
1CATALOG_SNAPSHOT = "catalog-2026-05-01" 2ELIGIBILITY_VERSION = "market-eligibility-v1" 3CANDIDATE_VERSION = "retrieval-v1" 4RANKER_VERSION = "market-ranker-v1" 5PREVIOUS_RANKER = "market-ranker-v0" 6EXPERIMENT_ID = "ranking-ab-2026-05" 7PROPENSITY_MODEL_VERSION = "position-examination-randomized-v1" 8# Frozen estimates from a separate randomized-position logging study. 9POSITION_PROPENSITY = {1: 0.95, 2: 0.78, 3: 0.62} 10 11ARM_CONFIG = { 12 "control": (PREVIOUS_RANKER, baseline_ranked), 13 "treatment": (RANKER_VERSION, candidate_ranked), 14} 15 16def impression_rows( 17 request_id: str, 18 shopper_id: str, 19 served_at: str, 20 query: str, 21 arm: str, 22) -> list[dict[str, object]]: 23 ranker_version, ranked_by_query = ARM_CONFIG[arm] 24 return [ 25 { 26 "request_id": request_id, 27 "shopper_id": shopper_id, 28 "served_at": served_at, 29 "query": query, 30 "catalog_snapshot": CATALOG_SNAPSHOT, 31 "eligibility_version": ELIGIBILITY_VERSION, 32 "candidate_version": CANDIDATE_VERSION, 33 "ranker_version": ranker_version, 34 "experiment_id": EXPERIMENT_ID, 35 "experiment_arm": arm, 36 "product_id": listing.product_id, 37 "position": position, 38 "position_propensity": POSITION_PROPENSITY[position], 39 "propensity_model_version": PROPENSITY_MODEL_VERSION, 40 "slate_size": len(ranked_by_query[query]), 41 "training_use": "implicit", 42 } 43 for position, listing in enumerate(ranked_by_query[query], start=1) 44 ] 45 46print("experiment:", EXPERIMENT_ID, "arms:", list(ARM_CONFIG))
06-log-impression-rows.py
1control_impressions = impression_rows("REQ-500", "SHOP-100", "2026-05-01T12:00:00Z", "label-printer", "control") 2treatment_impressions = impression_rows("REQ-501", "SHOP-200", "2026-05-01T12:00:05Z", "label-printer", "treatment") 3impressions = control_impressions + treatment_impressions 4outcomes = [ 5 { 6 "event_id": "EVT-900", 7 "occurred_at": "2026-05-01T12:00:08Z", 8 "request_id": "REQ-501", 9 "product_id": "P5", 10 "outcome": "click", 11 }, 12 { 13 "event_id": "EVT-901", 14 "occurred_at": "2026-05-01T12:04:18Z", 15 "request_id": "REQ-501", 16 "product_id": "P5", 17 "outcome": "purchase", 18 }, 19] 20 21print("impressions:", len(impressions), "outcomes:", len(outcomes))
07-join-outcomes-to-impressions.py
1required_impression_fields = { 2 "request_id", "shopper_id", "served_at", "query", "catalog_snapshot", 3 "eligibility_version", "candidate_version", "ranker_version", 4 "experiment_id", "experiment_arm", "product_id", "position", 5 "position_propensity", "propensity_model_version", "slate_size", "training_use", 6} 7required_outcome_fields = {"event_id", "occurred_at", "request_id", "product_id", "outcome"} 8outcome_event_ids = [row["event_id"] for row in outcomes] 9latencies_ms = [11, 13, 12, 15, 14, 16, 13, 12, 17, 18] 10# Toy online conversion table on the two arms (stop-rule practice). 11arm_searches = {"control": 100, "treatment": 100} 12arm_purchases = {"control": 4, "treatment": 6} 13arm_returns = {"control": 0, "treatment": 1} 14conversion = { 15 arm: arm_purchases[arm] / arm_searches[arm] 16 for arm in arm_searches 17} 18return_rate = { 19 arm: (arm_returns[arm] / arm_purchases[arm]) if arm_purchases[arm] else 0.0 20 for arm in arm_searches 21} 22stop_returns = return_rate["treatment"] > return_rate["control"] + 0.01 23 24def nearest_rank(values: list[int], percentile: float) -> int: 25 if not values: 26 raise ValueError("latency sample must not be empty") 27 rank = max(1, ceil(len(values) * percentile)) 28 return sorted(values)[rank - 1] 29 30baseline_ndcg = sum(ndcg(rows) for rows in baseline_ranked.values()) / len(QUERIES) 31candidate_ndcg = sum(ndcg(rows) for rows in candidate_ranked.values()) / len(QUERIES) 32max_seller_count = max( 33 max(Counter(listing.seller_id for listing in rows).values()) 34 for rows in candidate_ranked.values() 35) 36impression_index = { 37 (row["request_id"], row["product_id"]): row 38 for row in impressions 39} 40shopper_arms: dict[str, set[str]] = {} 41for row in impressions: 42 shopper_arms.setdefault(row["shopper_id"], set()).add(row["experiment_arm"]) 43 44def parse_utc(value: str) -> datetime: 45 return datetime.fromisoformat(value.replace("Z", "+00:00")) 46 47def outcome_follows_impression(row: dict[str, object]) -> bool: 48 impression = impression_index.get((row["request_id"], row["product_id"])) 49 return impression is not None and parse_utc(impression["served_at"]) < parse_utc(row["occurred_at"]) 50 51print("baseline_ndcg:", round(baseline_ndcg, 3), "candidate_ndcg:", round(candidate_ndcg, 3))
08-offline-release-gates.py
1release_gates = { 2 "candidate_recall_complete": all(candidate_recall(query, retrieved[query]) == 1.0 for query in QUERIES), 3 "candidate_ndcg_improves": candidate_ndcg > baseline_ndcg, 4 "ndcg_slices_at_least_0_95": all(ndcg(rows) >= 0.95 for rows in candidate_ranked.values()), 5 "ranker_p95_latency_ms_at_most_20": nearest_rank(latencies_ms, 0.95) <= 20, 6 "blocked_items_absent": not blocked_hits, 7 "impression_schema": all(required_impression_fields <= row.keys() for row in impressions), 8 "implicit_feedback_has_propensity": all( 9 row.get("training_use") != "implicit" 10 or ( 11 isinstance(row.get("position_propensity"), (int, float)) 12 and 0 < row["position_propensity"] <= 1 13 and row.get("propensity_model_version") == PROPENSITY_MODEL_VERSION 14 and isinstance(row.get("slate_size"), int) 15 and row["slate_size"] >= 1 16 ) 17 for row in impressions 18 ), 19 "outcome_schema": all(required_outcome_fields <= row.keys() for row in outcomes), 20 "outcome_event_ids_unique": len(outcome_event_ids) == len(set(outcome_event_ids)), 21 "outcomes_join_impressions": all((row["request_id"], row["product_id"]) in impression_index for row in outcomes), 22 "outcomes_follow_impressions": all(outcome_follows_impression(row) for row in outcomes), 23 "one_arm_per_shopper": all(len(arms) == 1 for arms in shopper_arms.values()), 24 "seller_concentration_at_most_2_per_slate": max_seller_count <= 2, 25} 26 27print("release_gates_pass:", all(release_gates.values()))
09-assemble-experiment-receipt.py
1receipt = { 2 "bundle_id": RANKER_VERSION, 3 "previous_ranker": PREVIOUS_RANKER, 4 "catalog_snapshot": CATALOG_SNAPSHOT, 5 "eligibility_version": ELIGIBILITY_VERSION, 6 "candidate_version": CANDIDATE_VERSION, 7 "offline": { 8 "baseline_ndcg_at_3": round(baseline_ndcg, 3), 9 "candidate_ndcg_at_3": round(candidate_ndcg, 3), 10 "ranker_p95_latency_ms": nearest_rank(latencies_ms, 0.95), 11 }, 12 "release_gates": release_gates, 13 "experiment": { 14 "experiment_id": EXPERIMENT_ID, 15 "assignment_unit": "shopper_id", 16 "control": PREVIOUS_RANKER, 17 "treatment": RANKER_VERSION, 18 "allocation": {"control": 0.5, "treatment": 0.5}, 19 "minimum_eligible_searches": 10000, 20 "primary_metric": "purchase_conversion_per_search", 21 "guardrails": ["returns_per_purchase", "search_p95_latency_ms", "blocked_listing_impressions", "seller_concentration_at_3"], 22 "stop_conditions": [ 23 "blocked_listing_impressions > 0", 24 "search_p95_latency_ms > 120", 25 "returns_per_purchase > control + 0.01", 26 ], 27 }, 28 "illustrative_online_review": { 29 "status": "stop_treatment_return_guardrail", 30 "control_conversion": conversion["control"], 31 "treatment_conversion": conversion["treatment"], 32 "control_return_rate": return_rate["control"], 33 "treatment_return_rate": return_rate["treatment"], 34 "stop_returns": stop_returns, 35 "release_evidence": False, 36 }, 37 "candidate_decision": "candidate_for_ab_test" if all(release_gates.values()) else "hold", 38} 39 40print("candidate_decision:", receipt["candidate_decision"])
10-publish-ranking-experiment-receipt.py
1print("control slate:", [row["product_id"] for row in control_impressions]) 2print("treatment slate:", [row["product_id"] for row in treatment_impressions]) 3print("offline:", receipt["offline"]) 4print("release_gates_pass:", all(receipt["release_gates"].values())) 5print("online_review:", receipt["illustrative_online_review"]["status"]) 6print("candidate_decision:", receipt["candidate_decision"])
Output
1control slate: ['P4', 'P5', 'P6'] 2treatment slate: ['P5', 'P6', 'P4'] 3offline: {'baseline_ndcg_at_3': 0.854, 'candidate_ndcg_at_3': 1.0, 'ranker_p95_latency_ms': 18} 4release_gates_pass: True 5online_review: stop_treatment_return_guardrail 6candidate_decision: candidate_for_ab_test

candidate_for_ab_test is narrower than launch approval. The local 20 millisecond gate measures ranker scoring only; the experiment's 120 millisecond stop applies to end-to-end search latency. The tiny online table is stop-rule practice, not live evidence:

ArmSearchesPurchasesConversionReturns / purchase
control10040.040.00
treatment10060.060.17

Treatment conversion rises, and the return-rate guardrail still fires (0.17 is far above control plus 0.01). The receipt says which ranker deserves controlled exposure, which immutable logs make later outcomes interpretable, and which previous alias stays available if a guardrail fails.

Submission checklist

ArtifactEvidence
eligibility policyblocked items can't enter ranking
candidate evaluationrecall measured before reranking
ranker evaluationNDCG slices and latency recorded
impression schemapositions, versions, and assignment context logged before outcomes
outcome streamlater events append separately and join displayed products
experiment planstable assignment, metrics, guardrails, stop rules
rollbackstable ranker alias remains available

Practice: break the ranking contract

Use the runnable examples as a verification testbed. Change one condition at a time, predict the result, then rerun the examples.

  1. Change P3 retrieval score from 0.84 to 0.10, then change retrieve() default budget from 3 to 2. Which metric exposes missing relevant product?
  2. Set P9.policy_approved to True. Why isn't that a harmless relevance change?
  3. Change candidate score for P6 from 0.84 to 0.20. Does ndcg_slices_at_least_0_95 fail?
  4. Remove position from impression_rows(). Which receipt gate blocks experiment readiness?
  5. Change all three eligible insulated-bag listings to seller S1. Which marketplace guardrail fails?
  6. Use SHOP-100 for both control and treatment impressions. Which assignment gate fails?
  7. Change first outcome's product_id from P5 to P9. Which replay gate fails?
  8. Append a retry with the same event_id as EVT-900. Which deduplication gate fails?

Practice answer sketches

Which metric catches dropped P3 after retrieval budget falls to 2?

Answer

candidate_recall_complete fails because the ranker never receives one relevant eligible product. NDCG on surviving candidates can't repair that omission.

Why isn't approving P9 a harmless relevance edit?

Answer

P9 was excluded by marketplace policy. Changing that field changes eligibility contract, not ranking quality, and requires its own reviewed policy release.

What happens when P6 candidate score falls from 0.84 to 0.20?

Answer

label-printer is no longer ideal (P5, P4, P6 instead of P5, P6, P4), but NDCG@3 only falls to about 0.972, which still clears 0.95. A worse-than-ideal list isn't the same as a failed slice gate. Demoting relevance-3 P5 to the bottom is what drops the slice below 0.95.

Which gate fails when impression rows drop position?

Answer

impression_schema fails. Without displayed position, later click and purchase data can't be interpreted against exposure.

Which gate fails when three products in one displayed slate share seller S1?

Answer

seller_concentration_at_most_2_per_slate fails, blocking an experiment that would crowd too much exposure into one seller.

Which gate fails when SHOP-100 appears in both experiment arms?

Answer

one_arm_per_shopper fails. A shopper-level experiment must keep each shopper in one arm so treatment effects stay interpretable.

Which gate fails when outcome product P9 was never displayed for REQ-501?

Answer

outcomes_join_impressions fails. A click or purchase can't become ranking feedback unless it joins a product exposed by the same request.

Which gate fails when a retry reuses EVT-900?

Answer

outcome_event_ids_unique fails. Event IDs distinguish one outcome from a replayed delivery, so a duplicate must be deduplicated before training or experiment analysis.

Complete the lesson

Mastery Check

Answer every question, then check your score. Score 75% or higher to mark this lesson complete.

1.P9 has the highest scores for insulated-bag but is excluded by policy_block. What does flipping P9.policy_approved to True change?

Correct answer: It changes the eligibility contract, so it needs a reviewed policy release rather than being treated as a ranker-quality tweak.

Eligibility is a hard boundary that runs before retrieval and scoring. A policy-blocked item belongs outside the candidate set, and changing its approval status changes what the marketplace is allowed to expose. That change belongs to the policy version and release review, not to ranker tuning.

2.A relevant eligible product, P3, is dropped when the retrieval budget for insulated-bag is reduced before reranking. Which release gate should fail?

Correct answer: candidate_recall_complete, because the candidate set no longer contains every relevant eligible judged product for the query.

Candidate recall is measured before reranking against the eligible relevant judgments. If P3 never reaches the ranker, recall drops even if the remaining products are ordered well. NDCG only evaluates the order of the candidates it receives, so it can't repair or reveal a missing candidate by itself.

3.In the label-printer fixture, if P5's candidate score falls from 0.96 to 0.20, the treatment order becomes P6, P4, P5 instead of ideal P5, P6, P4. Which release gate catches this regression even though the two-query treatment average can still beat the baseline?

Correct answer: ndcg_slices_at_least_0_95, because the per-query NDCG gate catches the bad label-printer ordering that the average can hide.

Lowering a rank score changes ordering, not retrieval eligibility, so recall and blocked-item checks aren't the issue. With the relevance-3 product P5 pushed to the bottom, the label-printer NDCG slice falls below the 0.95 gate even though the average across both fixture queries can still improve over the baseline. NDCG is gated by query category instead of relying only on an overall average.

4.An impression row for a displayed search slate omits position but keeps request, shopper, query, version, experiment, and product fields. Which release gate should fail?

Correct answer: impression_schema, because position is required to preserve the displayed exposure for later click and purchase interpretation.

Displayed position is part of the immutable impression record because clicks and purchases depend on what the shopper saw and where it appeared. Outcome events have their own schema and join by request plus product, while offline NDCG is computed from judged candidate lists. Dropping position breaks exposure logging, not the outcome schema or the offline metric.

5.A team wants to train the next ranker by labeling every clicked product as preferred and every unclicked product as irrelevant. What evidence must be added before this implicit feedback is eligible for training?

Correct answer: Join each outcome to its immutable displayed slate and require a randomized action probability or calibrated examination propensity with a pinned model version.

Clicks are behavior under a displayed ranking, so position and slate exposure must be preserved before outcomes arrive. Training from implicit feedback also needs the actual randomized action probability or a calibrated examination propensity tied to a versioned model. Arm assignment or a guessed curve such as 1 / position doesn't supply that evidence.

6.An outcome event for request REQ-501 is changed from product P5 to product P9, and the impression rows for REQ-501 displayed only P5, P6, and P4. Which gate directly tests whether the appended outcome is usable as ranking feedback?

Correct answer: outcomes_join_impressions requires request and exposed product to match; unexposed P9 cannot supply ranking feedback.

Outcome rows are appended after the immutable impression rows, and a row can have all required outcome fields while still failing as ranking feedback. The decisive check is whether its (request_id, product_id) pair appears in the impression log. Mentioning P9 in an outcome doesn't prove it was displayed, and offline NDCG isn't updated from these events.

7.A treatment slate contains three eligible products, all from seller S1. The release receipt allows at most two products from one seller in a slate. Which gate should fail?

Correct answer: The seller_concentration_at_most_2_per_slate gate fails: three S1 products exceed the two-item seller limit.

Seller concentration is a marketplace guardrail separate from relevance and eligibility. A slate can contain eligible, relevant products and still fail because it over-concentrates exposure in one seller. NDCG doesn't apply a seller-diversity discount in the formula used here.

8.The experiment assignment unit is shopper_id, but SHOP-100 appears in both control and treatment impression rows. Which gate should fail?

Correct answer: one_arm_per_shopper fails because SHOP-100 appears in both control and treatment, violating one-arm assignment.

Stable assignment is part of the experiment design, not a schema or offline-ranking calculation. If the same shopper sees both arms, later behavior can no longer be cleanly attributed to control or treatment exposure. The impression rows may still contain all required fields, but the assignment invariant is broken.

9.An offline-qualified treatment enters its A/B test. Purchase conversion rises from 0.04 to 0.06, but returns per purchase rise from 0.00 to 0.17. The stop rule is returns per purchase greater than control plus 0.01. What should the team do?

Correct answer: Stop treatment and roll back: returns per purchase (0.17) exceeds control (0.00) by more than 0.01.

The observed increase of 0.17 exceeds the allowed increase of 0.01. A primary-metric lift doesn't override a breached guardrail or stop condition. Offline evidence qualified the candidate for controlled exposure, not unconditional launch, and the recorded previous ranker provides the rollback target.

10.A displayed product is clicked and later purchased. An ingestion retry then replays the click event. Which record design keeps the history auditable?

Correct answer: Leave impression unchanged; append timestamped click and purchase events, then join by request and product.

The impression is the immutable record of what was displayed. Clicks and purchases are later events, so appending them preserves their event times and allows multiple outcomes. Unique event IDs let ingestion deduplicate retries, while request and product identify the exposure to which each outcome belongs. Updating or duplicating impressions would corrupt the replayable exposure history.

10 questions remaining.

Next Step
Continue to Capstone: Demand Forecasting

You can now ship a ranking surface that logs exposure before it learns from clicks, and that refuses to treat an offline NDCG win as a launch. Next you'll ship time-ordered forecasts and alerts, where history has to stay in calendar order instead of being reordered like a search slate.

PreviousCapstone: Delivery ETA Prediction
Share this article
XFacebookLinkedInBlueskyRedditHacker NewsEmail
References

Learning to Rank using Gradient Descent.

Burges, C. J. C., et al. · 2005 · ICML 2005

https://www.microsoft.com/en-us/research/publication/learning-to-rank-using-gradient-descent/

Introduction to Information Retrieval.

Manning, C. D., Raghavan, P., Schutze, H. · 2008 · Cambridge University Press

https://nlp.stanford.edu/IR-book/

Unbiased Learning-to-Rank with Biased Feedback

Joachims, T., Swaminathan, A., & Schnabel, T. · 2017 · WSDM 2017

https://arxiv.org/abs/1608.04468

Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing.

Kohavi, R., Tang, D., Xu, Y. · 2020

https://experimentguide.com/

Discussion

Questions and insights from fellow learners.

Discussion loads when you reach this section.