LeetLLM
My PlanLearnGlossaryTracksPracticeBlog
LeetLLM

Your go-to resource for mastering AI & LLM systems.

Product

  • Learn
  • Glossary
  • Tracks
  • Practice
  • Blog
  • RSS

Legal

  • Terms of Service
  • Privacy Policy

© 2026 LeetLLM. All rights reserved.

All Topics
Your Progress
0%

0 of 178 articles completed

🛠️Computing Foundations0/9
Git, Shell, Linux for AIDocker for Reproducible AIPython for AI EngineeringNumPy and Tensor ShapesCUDA for ML TrainingMPS & Metal for ML on MacData Structures for AISQL and Data ModelingAlgorithms for ML Engineers
📊Math & Statistics0/8
Gradients and BackpropVectors, Matrices & TensorsLinear Algebra for MLAdam, Momentum, SchedulersProbability for Machine LearningStatistics and UncertaintyDistributions and SamplingHypothesis Tests, Intervals, and pass@k
📚Preparation & Prerequisites0/13
Neural Networks from ScratchCNNs from ScratchTraining & BackpropagationSoftmax, Cross-Entropy & OptimizationRNNs, LSTMs, GRUs, and Sequence ModelingAutoencoders and VAEsThe Transformer Architecture End-to-EndLanguage Modeling & Next TokensFrom GPT to Modern LLMsPrompt Engineering FundamentalsCalling LLM APIs in ProductionFirst AI App End-to-EndThe LLM Lifecycle
🧮ML Algorithms & Evaluation0/11
Linear Regression from ScratchLogistic Regression and MetricsDecision Trees, Forests, and BoostingReinforcement Learning BasicsValidation and LeakageClustering and PCACore Retrieval AlgorithmsDecoding AlgorithmsExperiment Design and A/B TestingPyTorch Training LoopsDataset Pipelines and Data Quality
📦Production ML Systems0/6
Feature Engineering for Production MLBatch and Streaming Feature PipelinesGradient Boosted Trees in ProductionRanking and Recommendation SystemsForecasting and Anomaly DetectionMonitoring Predictive Models
🧪Core LLM Foundations0/8
The Bitter Lesson & ComputeBPE, WordPiece, and SentencePieceStatic to Contextual EmbeddingsPerplexity & Model EvaluationFile Ingestion for AIChunking StrategiesLLM Benchmarks & LimitationsInstruction Tuning & Chat Templates
🧰Applied LLM Engineering0/24
Dimensionality Reduction for EmbeddingsCoT, ToT & Self-Consistency PromptingFunction Calling & Tool UseMCP & Tool Protocol StandardsContext EngineeringPrompt Injection DefenseResponsible AI GovernanceData Labeling and Human FeedbackEvaluating AI AgentsProduction RAG PipelinesHybrid Search: Dense + SparseReranking and Cross-Encoders for RAGRAG Evaluation for Reliable AnswersLLM-as-a-Judge EvaluationBias & Fairness in LLMsHallucination Detection & MitigationLLM Observability & MonitoringExperiment Tracking with MLflow and W&BPrompt Optimization with DSPyModel Versioning & DeploymentSemantic Caching & Cost OptimizationLLM Cost Engineering & Token EconomicsModel Gateways, Routing, and FallbacksDesign an Automated Support Agent
🎓Portfolio Capstones0/9
Capstone: Delivery ETA PredictionCapstone: Product RankingCapstone: Demand ForecastingCapstone: Image Damage ClassifierCapstone: Production ML PipelineCapstone: Document QACapstone: Eval DashboardCapstone: Fine-Tuned ClassifierCapstone: Reproducible ML Study
🧠Transformer Deep Dives0/8
Sentence Embeddings & Contrastive LossEmbedding Similarity & QuantizationScaled Dot-Product AttentionVision Transformers and Image EncodersPositional Encoding: RoPE & ALiBiLayer Normalization: Pre-LN vs Post-LNMechanistic InterpretabilityDecoding Strategies: Greedy to Nucleus
🧬Advanced Training & Adaptation0/17
Scaling Laws & Compute-Optimal TrainingPre-training Data at ScaleBuild GPT from Scratch LabJAX for PyTorch ResearchersContinued Pretraining for Domain ShiftSynthetic Data PipelinesSupervised Fine-Tuning PipelineMixed Precision TrainingDistributed Training: FSDP & ZeROLoRA & Parameter-Efficient TuningTraining Run OperationsReward Modeling from Preference DataRLHF & DPO AlignmentConstitutional AI & Red TeamingRLVR & Verifiable RewardsKnowledge Distillation for LLMsModel Merging and Weight Interpolation
🤖Advanced Agents & Retrieval0/16
Vector DB Internals: HNSW & IVFAdvanced RAG: HyDE & Self-RAGGraphRAG & Knowledge GraphsRAG Security & Access ControlStructured Output GenerationReAct & Plan-and-ExecuteGuardrails & Safety FiltersCode Generation & SandboxingComputer-Use / GUI / Browser AgentsHuman-in-the-Loop Agent ArchitectureAI Coding Workflow with AgentsAgent Memory & PersistenceAgent Failure & RecoveryRecursive Language Models (RLM)Multi-Agent OrchestrationCapstone: Production Agent
⚡Inference & Production Scale0/19
Inference: TTFT, TPS & KV CacheMulti-Query & Grouped-Query AttentionKV Cache & PagedAttentionPrefix Caching and Prompt CachingFlashAttention & Memory EfficiencyContinuous Batching & SchedulingScaling LLM InferenceModel Parallelism for LLM InferenceModel Quantization: GPTQ, AWQ & GGUFLocal LLM DeploymentSLM Specialization & Edge DeploymentSpeculative DecodingLong Context Window ManagementMixture of Experts ArchitectureMamba & State Space ModelsReasoning & Test-Time ComputeAdvanced MLOps & DevOps for AIGPU Serving & AutoscalingA/B Testing for LLMs
🏗️System Design Capstones0/9
Content Moderation SystemCode Completion SystemMulti-Tenant LLM PlatformLLM-Powered Search EngineVision-Language Models & CLIPMultimodal LLM ArchitectureDiffusion Models: Images & TextReal-Time Voice AI AgentReasoning Agent System Design
🎤AI Lab Interviewing0/4
AI Lab Coding Interview: Python SystemsAI Lab System Design InterviewAI Lab Behavioral InterviewAI Lab Technical Presentation
🔬Project Deep Dives0/17
Deep Dive - vLLMDeep Dive - SkyRLDeep Dive - FlashAttentionDeep Dive - FlashInferDeep Dive - DeepGEMMDeep Dive - NCCLDeep Dive - MegatronDeep Dive - DeepSpeedDeep Dive - RayDeep Dive - MLflowDeep Dive - PyTorchDeep Dive - TransformersDeep Dive - SGLangDeep Dive - slimeDeep Dive - DeepEPDeep Dive - TinkerDeep Dive - Light-PEFT
Back to Topics
LearnPortfolio CapstonesCapstone: Image Damage Classifier
👁️HardMultimodal Models

Capstone: Image Damage Classifier

Ship a damaged-package photo triage service with quality checks, slice evaluation, serving bundles, and review monitoring.

20 min read
Learning path
Step 83 of 178 in the full curriculum
Capstone: Demand ForecastingCapstone: Production ML Pipeline

Personalize this lesson

Adapt explanations and teaching visuals to your background and preferred voice.

A damage classifier receives photos that may be crushed, torn, blurred, dark, or unrelated to the reported case. Unlike a late-delivery score or a seven-day parcel forecast, the input can fail before classification begins because the evidence itself isn't usable.

The last capstone packaged a demand forecast with shadow evidence, delayed observation joins, and a rollback pointer. Reuse that receipt habit for pixels. Ship an image-triage endpoint that flags likely visible damage, rejects unusable photos, preserves evidence for human review, and never turns an uncertain image score into a refund.

You already traced a convolutional neural network (CNN) over a cracked equipment-panel patch in CNNs from Scratch. That spatial score is the starting point, not the product: quality checks and specialist confirmation still sit in front of any costly action.

Define the photo decision first

The service receives return photos from customer uploads and intake cameras. The useful product question isn't "does the model recognize every defect?" It's: which photo should a specialist inspect first, and when is the photo too weak to support any decision?

Use three operational outcomes:

ActionEvidenceProduct behavior
request_new_photoimage is too blurred, dark, or incompleteask for a clearer upload before assessing damage
normal_reviewusable image, low damage scorekeep ordinary return workflow
priority_damage_reviewusable image, high damage scoresurface to specialist with photo and score trace

The classifier isn't a refund policy. Product eligibility still depends on order ownership, item type, return window, and specialist judgment. This separation prevents a shadow or reflection in a photo from issuing a costly action.

A model card should state intended use (and out-of-scope uses), decision thresholds, the slices you actually evaluated, and known limitations. Mitchell et al. proposed model cards as short reports so people can see those operating conditions instead of a lone metric.[1]Reference 1Model Cards for Model Reportinghttps://arxiv.org/abs/1810.03993

Diagram showing Photo manifest case groups + time, Quality check visible + blur + light, Damage scorer versioned threshold, and Immutable route trace policy + thresholds.
Photo manifest case groups + time, Quality check visible + blur + light, Damage scorer versioned threshold, and Immutable route trace policy + thresholds.

Those three actions only stay honest if related photos of one physical package can't leak into both train and test. That's the dataset job next.

Build a dataset that can't leak

For tabular models, leakage may be a future delivery timestamp. For photos, leakage often hides in nearly identical pixels. A customer may upload three bursts of the same crushed box. A warehouse may photograph one parcel from four angles. If related images land in both train and test sets, the model can memorize one package rather than generalize to new damage.

Your manifest should contain:

FieldWhy it matters
case_id and capture_daygroup all photos for one physical package and preserve time ordering
sourceseparate customer phone uploads from warehouse inspection cameras
quality_labeldistinguish unusable evidence from visible damage
damage_labelrecord specialist-confirmed visible damage only on usable photos
splithold out later cases, never random photos from the same case
reviewer_id and guideline_versionaudit disagreement or changed label definitions

Evaluate at least daylight versus dark uploads, customer versus warehouse source, packaging type, and visible-defect size. A global score can hide the exact failure that matters: small tears disappearing in dark phone images.

Use the CNN from CNNs from Scratch as a baseline. Fine-tune a pretrained image encoder only if you record its preprocessing and measure it under the same grouped, time-ordered split. A later deep-dive covers Vision Transformer image encoders; this capstone doesn't require that architecture.[2]Reference 2An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.https://arxiv.org/abs/2010.11929

Freeze a grouped photo manifest

Start with the dataset boundary. Multiple photos from one physical package belong in one split even when filenames differ. Keep time direction intact too: later return cases should test a model trained on earlier cases.

The local manifest below is small enough to inspect line by line. This historical evaluation fixture contains frozen specialist labels, reviewer IDs, and guideline versions for dataset audit. Damage labels appear only when a specialist judged the image usable. Unusable photos keep confirmed_damage=None because poor evidence shouldn't become supervision. Candidate scores arrive in a separate object, so route code can't quietly read labels as model output. Manifest confirmed_damage fields are offline audit only; the serving path must not read them. MODEL_OUTPUTS below are frozen fixtures for the routing receipt. Fitting train_cnn_baseline.py and applying the preprocessing version to pixels is a separate training lab.

01-freeze-grouped-photo-manifest.py
1from collections import Counter, defaultdict 2from dataclasses import dataclass 3import json 4 5@dataclass(frozen=True) 6class PhotoManifestRow: 7 case_id: str 8 photo_id: str 9 capture_day: int 10 split: str 11 source: str 12 packaging: str 13 quality_label: str 14 confirmed_damage: bool | None 15 reviewer_id: str 16 guideline_version: str 17 18PHOTOS = [ 19 PhotoManifestRow("R-401", "R-401-a", 1, "train", "customer_phone", "corrugated", "usable", True, "S-12", "visible-damage-v1"), 20 PhotoManifestRow("R-401", "R-401-b", 2, "train", "customer_phone", "corrugated", "usable", True, "S-12", "visible-damage-v1"), 21 PhotoManifestRow("R-402", "R-402-a", 3, "validation", "customer_phone", "corrugated", "unusable", None, "S-08", "visible-damage-v1"), 22 PhotoManifestRow("R-403", "R-403-a", 4, "validation", "warehouse_camera", "mailer", "usable", False, "S-08", "visible-damage-v1"), 23 PhotoManifestRow("R-404", "R-404-a", 5, "test", "customer_phone", "mailer", "usable", True, "S-12", "visible-damage-v1"), 24 PhotoManifestRow("R-404", "R-404-b", 6, "test", "customer_phone", "mailer", "usable", True, "S-12", "visible-damage-v1"), 25 PhotoManifestRow("R-405", "R-405-a", 7, "test", "warehouse_camera", "corrugated", "usable", False, "S-08", "visible-damage-v1"), 26 PhotoManifestRow("R-406", "R-406-a", 8, "test", "customer_phone", "corrugated", "unusable", None, "S-12", "visible-damage-v1"), 27 PhotoManifestRow("R-407", "R-407-a", 9, "test", "customer_phone", "corrugated", "usable", False, "S-12", "visible-damage-v1"), 28 PhotoManifestRow("R-408", "R-408-a", 10, "test", "warehouse_camera", "corrugated", "usable", True, "S-08", "visible-damage-v1"), 29 PhotoManifestRow("R-409", "R-409-a", 11, "test", "warehouse_camera", "corrugated", "unusable", None, "S-08", "visible-damage-v1"), 30] 31 32print("photos:", len(PHOTOS)) 33print("cases:", len({photo.case_id for photo in PHOTOS}))
Output
1photos: 11 2cases: 9
02-audit-grouped-split-manifest.py
1splits_by_case = defaultdict(set) 2for photo in PHOTOS: 3 splits_by_case[photo.case_id].add(photo.split) 4 5split_days = { 6 split: [photo.capture_day for photo in PHOTOS if photo.split == split] 7 for split in ("train", "validation", "test") 8} 9manifest_checks = { 10 "case_groups_do_not_cross_splits": all(len(splits) == 1 for splits in splits_by_case.values()), 11 "time_ordered_splits": ( 12 all(split_days.values()) 13 and max(split_days["train"]) < min(split_days["validation"]) 14 and max(split_days["validation"]) < min(split_days["test"]) 15 ), 16 "usable_labels_complete": all(photo.confirmed_damage is not None for photo in PHOTOS if photo.quality_label == "usable"), 17 "unusable_labels_abstain": all(photo.confirmed_damage is None for photo in PHOTOS if photo.quality_label == "unusable"), 18 "label_lineage_recorded": all(photo.reviewer_id and photo.guideline_version for photo in PHOTOS), 19} 20 21print("split photo counts:", dict(Counter(photo.split for photo in PHOTOS))) 22print("R-401 splits:", sorted(splits_by_case["R-401"])) 23print(json.dumps(manifest_checks, indent=2)) 24assert all(manifest_checks.values())
Output
1split photo counts: {'train': 2, 'validation': 2, 'test': 7} 2R-401 splits: ['train'] 3{ 4 "case_groups_do_not_cross_splits": true, 5 "time_ordered_splits": true, 6 "usable_labels_complete": true, 7 "unusable_labels_abstain": true, 8 "label_lineage_recorded": true 9}

The fixture uses explicit splits so the invariant stays visible. A larger pipeline can use a group-aware splitter, then freeze and audit the resulting manifest. The important claim isn't that one splitter solves every dataset: no physical case may cross evaluation boundaries, and later cases remain later.

The manifest proves the split and the label contract. It doesn't yet prove the endpoint refuses to read those labels when it chooses a route.

Encode quality-first review traces

The model endpoint should receive a preprocessing result and a damage score, then choose a review route. Quality checks run before the damage threshold. Otherwise a confidently scored blur, dark frame, or unrelated object can create an unsupported escalation.

The cheap checks in this fixture are three numbers you can inspect without opening a CNN:

  • box_visible: is a parcel in the frame?
  • brightness: mean intensity in [0, 1]. Below 0.20 the photo is too dark to judge a tear.
  • blur_score: higher means more blur. Above 0.45 the edges that would show a crush are gone.

A production service may compute those with a blur filter or an exposure histogram. The routing invariant doesn't depend on the exact formula: quality fails first.

A high damage probability still isn't proof. Guo et al. showed that modern networks are often overconfident relative to their accuracy, so you shouldn't treat a 0.94 as a calibrated chance of visible damage, especially on a dark or empty frame.[3]Reference 3On Calibration of Modern Neural Networkshttps://arxiv.org/abs/1706.04599 That's why the router never lets the score override a failed quality check.

The score fixture below is separate from the manifest. Labels stay available for offline evaluation, but the router never reads them. A production system may run its cheap quality checks before an expensive damage model; this compact fixture keeps both outputs so you can test that an unsupported high damage score still abstains.

03-load-model-output-bundle.py
1@dataclass(frozen=True) 2class ModelOutput: 3 damage_probability: float 4 blur_score: float 5 brightness: float 6 box_visible: bool 7 8MODEL_OUTPUTS = { 9 "R-401-a": ModelOutput(0.91, 0.12, 0.66, True), 10 "R-401-b": ModelOutput(0.88, 0.10, 0.70, True), 11 "R-402-a": ModelOutput(0.93, 0.71, 0.51, True), 12 "R-403-a": ModelOutput(0.18, 0.08, 0.75, True), 13 "R-404-a": ModelOutput(0.83, 0.10, 0.64, True), 14 "R-404-b": ModelOutput(0.79, 0.14, 0.61, True), 15 "R-405-a": ModelOutput(0.32, 0.11, 0.68, True), 16 "R-406-a": ModelOutput(0.94, 0.12, 0.12, True), 17 "R-407-a": ModelOutput(0.73, 0.09, 0.65, True), 18 "R-408-a": ModelOutput(0.63, 0.07, 0.72, True), 19 "R-409-a": ModelOutput(0.81, 0.08, 0.74, False), 20} 21 22BUNDLE = { 23 "bundle_id": "damage-cnn-v1", 24 "previous_bundle": "damage-cnn-v0", 25 "preprocessing": "parcel-rgb-224-center-crop-v1", 26 "label_guideline": "visible-damage-v1", 27 "route_policy": "quality-first-review-v1", 28 "damage_threshold": 0.70, 29 "max_blur": 0.45, 30 "min_brightness": 0.20, 31} 32 33print("bundle:", BUNDLE["bundle_id"], "threshold=", BUNDLE["damage_threshold"])
04-quality-first-route.py
1def route(photo: PhotoManifestRow, model_output: ModelOutput) -> dict[str, object]: 2 if not model_output.box_visible: 3 action, reason = "request_new_photo", "package_not_visible" 4 elif model_output.blur_score > BUNDLE["max_blur"] or model_output.brightness < BUNDLE["min_brightness"]: 5 action, reason = "request_new_photo", "image_quality_gate" 6 elif model_output.damage_probability >= BUNDLE["damage_threshold"]: 7 action, reason = "priority_damage_review", "damage_threshold" 8 else: 9 action, reason = "normal_review", "below_threshold" 10 11 return { 12 "route_id": f"{BUNDLE['bundle_id']}:{photo.photo_id}", 13 "photo_id": photo.photo_id, 14 "case_id": photo.case_id, 15 "routed_day": photo.capture_day, 16 "source": photo.source, 17 "packaging": photo.packaging, 18 "bundle_id": BUNDLE["bundle_id"], 19 "previous_bundle": BUNDLE["previous_bundle"], 20 "preprocessing": BUNDLE["preprocessing"], 21 "label_guideline": BUNDLE["label_guideline"], 22 "route_policy": BUNDLE["route_policy"], 23 "damage_threshold": BUNDLE["damage_threshold"], 24 "max_blur": BUNDLE["max_blur"], 25 "min_brightness": BUNDLE["min_brightness"], 26 "damage_probability": model_output.damage_probability, 27 "blur_score": model_output.blur_score, 28 "brightness": model_output.brightness, 29 "box_visible": model_output.box_visible, 30 "action": action, 31 "reason": reason, 32 } 33 34print("route keys:", len(route(PHOTOS[0], MODEL_OUTPUTS["R-401-a"])))
05-route-test-photos.py
1test_photos = [photo for photo in PHOTOS if photo.split == "test"] 2photo_by_id = {photo.photo_id: photo for photo in PHOTOS} 3traces = [route(photo, MODEL_OUTPUTS[photo.photo_id]) for photo in test_photos] 4trace_by_photo = {trace["photo_id"]: trace for trace in traces} 5unsafe_priority_routes = [ 6 trace["photo_id"] 7 for photo, trace in zip(test_photos, traces) 8 if trace["action"] == "priority_damage_review" and photo.quality_label != "usable" 9] 10 11for photo_id in ("R-404-a", "R-406-a", "R-409-a"): 12 trace = trace_by_photo[photo_id] 13 print(photo_id, trace["action"], trace["reason"], f"score={trace['damage_probability']}") 14print("route counts:", dict(Counter(trace["action"] for trace in traces))) 15print("unsafe priority routes:", unsafe_priority_routes) 16assert unsafe_priority_routes == [] 17assert trace_by_photo["R-406-a"]["action"] == "request_new_photo" 18assert trace_by_photo["R-409-a"]["action"] == "request_new_photo"
Output
1R-404-a priority_damage_review damage_threshold score=0.83 2R-406-a request_new_photo image_quality_gate score=0.94 3R-409-a request_new_photo package_not_visible score=0.81 4route counts: {'priority_damage_review': 3, 'normal_review': 2, 'request_new_photo': 2} 5unsafe priority routes: []

Cases R-406-a and R-409-a are the important failure tests. Both have high damage scores. Neither score counts as usable evidence because quality checks fail first. The endpoint asks for another photo instead of escalating an unsupported claim.

Brightness-versus-blur plot of seven test photos. Usable evidence is brightness at least 0.20 and blur at most 0.45. Dot size is damage score. R-406-a is too dark at brightness 0.12 with score 0.94, so it recaptures. R-409-a sits in the usable region with no visible package and score 0.81, so it also recaptures. Usable R-404-a scores 0.83 and takes priority review.
Brightness runs right and blur runs up. Dashed lines mark usable evidence (brightness at least 0.20, blur at most 0.45). R-406-a sits left of the brightness gate with score 0.94; R-409-a sits in the usable region with no package visible. Both recapture. R-404-a is the allowed high-score priority route.

The scatter makes the same contract visible. Usable photos cluster in the lower-right region. R-406-a is the large dot on the dark side of the brightness gate. R-409-a is the hollow dot in that cluster: light and sharp enough, but the package isn't in the frame. Dot size is the damage score, so both rejected photos are large and still not priority routes.

Package the vision service

Submit an inspectable repository, not a notebook screenshot:

text
1damage-vision-service/ 2 data/ 3 label_guidelines.md 4 photo_manifest.parquet 5 split_manifest.json 6 model/ 7 train_cnn_baseline.py 8 evaluate_slices.py 9 model_card.md 10 service/ 11 preprocess.py 12 route_review.py 13 trace_schema.json 14 monitoring/ 15 input_quality_report.py 16 delayed_review_outcomes.py 17 specialist_shadow_receipt.py 18 tests/ 19 test_case_groups_do_not_cross_splits.py 20 test_blurry_photo_never_escalates.py 21 test_later_outcomes_join_routes.py 22 test_empty_review_window_holds.py 23 test_route_trace_is_versioned.py 24 test_previous_bundle_required.py

The serving bundle must pin image resize and crop behavior, color normalization, model weights, label version, route policy, damage threshold, quality-gate thresholds, and previous-bundle pointer. A change from center crop to full-frame resize may change whether a torn corner remains visible; it's a model behavior change even when weights remain constant.

The routing cells emit that trace shape directly. They let a reviewer reconstruct the route:

Response fieldExample
bundle and preprocessingdamage-cnn-v1, parcel-rgb-224-center-crop-v1
rollback and label contractdamage-cnn-v0, visible-damage-v1
quality valuesblur 0.12, brightness 0.66, box visible true
score and action policydamage 0.91, threshold 0.70, blur maximum 0.45, brightness minimum 0.20
routepriority_damage_review
human outcome laterconfirmed_damage or not_supported

A replayable trace records what the bundle decided. It still isn't a confirmed-damage label. Those arrive later, and they join by case.

Join delayed outcomes before promotion

Photo models drift when the image source changes. A new warehouse camera, winter lighting, a mobile upload compressor, or new packaging graphics can alter pixels before a confirmed-damage label exists.

Separate immediate checks from delayed quality:

WindowMonitorTrigger
immediateunreadable image rate, brightness, blur, missing package, latencyinvestigate capture path or fail to manual intake
delayedspecialist-confirmed precision, missed visible damage, route rate by source and packaginghold promotion or create retraining candidate
safety reviewunsupported escalations, policy actions attempted without specialist approvalrollback and audit workflow

Google Cloud's MLOps guidance treats data validation, model validation, serving, monitoring, metadata, and continuous training as connected stages rather than a one-time deploy.[4]Reference 4MLOps: Continuous Delivery and Automation Pipelines in Machine Learning.https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning Use that here: a jump in dark warehouse-camera photos starts an investigation or a candidate run. Promotion still waits for the delayed specialist join, not for an automatic weight swap.

The final local receipt appends specialist outcomes only after routing, then joins each outcome to a physical case. Each delayed outcome also carries its reviewer and label guideline, so precision and recall can't silently mix definitions. It measures review precision and recall once per usable test case, records immediate request rates by source, proves unusable photos abstain, and keeps the previous bundle beside the candidate. When one package has multiple usable photos, the case escalates if any usable photo crosses the priority threshold. These tiny counts teach the contract; production promotion still needs larger slices and shadow traffic.

06-append-specialist-outcomes.py
1@dataclass(frozen=True) 2class SpecialistOutcome: 3 case_id: str 4 reviewed_day: int 5 confirmed_damage: bool 6 reviewer_id: str 7 guideline_version: str 8 9SPECIALIST_OUTCOMES = [ 10 SpecialistOutcome("R-404", 12, True, "S-12", "visible-damage-v1"), 11 SpecialistOutcome("R-405", 12, False, "S-08", "visible-damage-v1"), 12 SpecialistOutcome("R-407", 12, False, "S-12", "visible-damage-v1"), 13 SpecialistOutcome("R-408", 12, True, "S-08", "visible-damage-v1"), 14] 15 16usable_test_case_ids = { 17 photo.case_id 18 for photo in test_photos 19 if photo.quality_label == "usable" 20} 21traces_by_case = defaultdict(list) 22for trace in traces: 23 traces_by_case[trace["case_id"]].append(trace) 24 25print("usable test cases:", sorted(usable_test_case_ids)) 26print("specialist outcomes:", len(SPECIALIST_OUTCOMES))
07-aggregate-case-actions.py
1def aggregate_usable_case_action(case_id: str) -> str: 2 usable_traces = [ 3 trace 4 for trace in traces_by_case[case_id] 5 if photo_by_id[trace["photo_id"]].quality_label == "usable" 6 ] 7 if not usable_traces: 8 raise ValueError(f"no usable traces for {case_id}") 9 if any(trace["action"] == "priority_damage_review" for trace in usable_traces): 10 return "priority_damage_review" 11 return "normal_review" 12 13case_actions = { 14 case_id: aggregate_usable_case_action(case_id) 15 for case_id in sorted(usable_test_case_ids) 16} 17outcome_by_case = {outcome.case_id: outcome for outcome in SPECIALIST_OUTCOMES} 18joined_outcomes = [ 19 (case_id, case_actions[case_id], outcome_by_case[case_id].confirmed_damage) 20 for case_id in sorted(case_actions.keys() & outcome_by_case.keys()) 21] 22 23print("case_actions:", case_actions) 24print("joined usable cases:", len(joined_outcomes)) 25assert case_actions == { 26 "R-404": "priority_damage_review", 27 "R-405": "normal_review", 28 "R-407": "priority_damage_review", 29 "R-408": "normal_review", 30}
08-measure-delayed-quality.py
1def rate_or_none(numerator: int, denominator: int) -> float | None: 2 return round(numerator / denominator, 3) if denominator else None 3 4true_positives = sum( 5 action == "priority_damage_review" and confirmed_damage 6 for _, action, confirmed_damage in joined_outcomes 7) 8false_positives = sum( 9 action == "priority_damage_review" and not confirmed_damage 10 for _, action, confirmed_damage in joined_outcomes 11) 12false_negatives = sum( 13 action != "priority_damage_review" and confirmed_damage 14 for _, action, confirmed_damage in joined_outcomes 15) 16delayed_quality = { 17 "priority_precision": rate_or_none(true_positives, true_positives + false_positives), 18 "priority_recall": rate_or_none(true_positives, true_positives + false_negatives), 19 "reviewed_usable_cases": len(joined_outcomes), 20} 21 22source_slices = { 23 source: { 24 "photos": len(rows), 25 "request_new_photo_rate": round( 26 sum(trace["action"] == "request_new_photo" for trace in rows) / len(rows), 27 3, 28 ), 29 } 30 for source in sorted({trace["source"] for trace in traces}) 31 for rows in [[trace for trace in traces if trace["source"] == source]] 32} 33 34packaging_slices = { 35 packaging: { 36 "photos": len(rows), 37 "priority_rate": round( 38 sum(trace["action"] == "priority_damage_review" for trace in rows) / len(rows), 39 3, 40 ), 41 } 42 for packaging in sorted({trace["packaging"] for trace in traces}) 43 for rows in [[trace for trace in traces if trace["packaging"] == packaging]] 44} 45 46# Didactic n-aware hold: small review windows teach denominators but are not launch proof. 47small_n_hold = delayed_quality["reviewed_usable_cases"] < 30 48delayed_quality["held_with_small_n"] = small_n_hold 49 50print("delayed_quality:", delayed_quality) 51print("source_slices:", source_slices) 52print("packaging_slices:", packaging_slices)
09-release-gate-checklist.py
1required_trace_fields = { 2 "route_id", "photo_id", "case_id", "routed_day", "source", "packaging", 3 "bundle_id", "previous_bundle", "preprocessing", "label_guideline", 4 "route_policy", "damage_threshold", "max_blur", "min_brightness", 5 "damage_probability", "blur_score", "brightness", "box_visible", 6 "action", "reason", 7} 8outcome_case_ids = [outcome.case_id for outcome in SPECIALIST_OUTCOMES] 9release_gates = { 10 **manifest_checks, 11 "route_ids_unique": len({trace["route_id"] for trace in traces}) == len(traces), 12 "all_test_photos_routed": len(traces) == len(test_photos), 13 "unusable_photos_abstain": all( 14 trace_by_photo[photo.photo_id]["action"] == "request_new_photo" 15 for photo in test_photos 16 if photo.quality_label == "unusable" 17 ), 18 "unsafe_priority_routes_absent": not unsafe_priority_routes, 19 "route_traces_replayable": all(required_trace_fields <= trace.keys() for trace in traces), 20 "specialist_outcome_case_ids_unique": len(set(outcome_case_ids)) == len(outcome_case_ids), 21 "specialist_outcomes_join_usable_test_cases": set(outcome_case_ids) <= usable_test_case_ids, 22 "usable_test_cases_have_specialist_outcomes": usable_test_case_ids <= set(outcome_case_ids), 23 "specialist_outcomes_arrive_after_capture": all( 24 outcome.reviewed_day > max( 25 photo.capture_day for photo in test_photos if photo.case_id == outcome.case_id 26 ) 27 for outcome in SPECIALIST_OUTCOMES 28 if outcome.case_id in usable_test_case_ids 29 ), 30 "specialist_outcome_label_contract_matches_bundle": all( 31 outcome.reviewer_id 32 and outcome.guideline_version == BUNDLE["label_guideline"] 33 for outcome in SPECIALIST_OUTCOMES 34 ), 35 "priority_precision_evidence_at_least_0_50": ( 36 delayed_quality["priority_precision"] is not None 37 and delayed_quality["priority_precision"] >= 0.50 38 ), 39 "priority_recall_evidence_at_least_0_50": ( 40 delayed_quality["priority_recall"] is not None 41 and delayed_quality["priority_recall"] >= 0.50 42 ), 43 "small_n_hold_recorded": delayed_quality.get("held_with_small_n") is True, 44 "source_slices_recorded": set(source_slices) == {"customer_phone", "warehouse_camera"}, 45 "packaging_slices_recorded": set(packaging_slices) >= {"corrugated", "mailer"}, 46 "rollback_pointer_recorded": bool(BUNDLE["previous_bundle"]), 47} 48 49print("release_gates_pass:", all(release_gates.values()))
10-publish-specialist-shadow-receipt.py
1receipt = { 2 "candidate_bundle": BUNDLE["bundle_id"], 3 "previous_bundle": BUNDLE["previous_bundle"], 4 "preprocessing": BUNDLE["preprocessing"], 5 "label_guideline": BUNDLE["label_guideline"], 6 "route_policy": BUNDLE["route_policy"], 7 "route_traces": len(traces), 8 "later_specialist_outcomes": len(SPECIALIST_OUTCOMES), 9 "joined_usable_cases": len(joined_outcomes), 10 "case_actions": case_actions, 11 "delayed_quality": delayed_quality, 12 "source_slices": source_slices, 13 "packaging_slices": packaging_slices, 14 "release_gates": release_gates, 15 "candidate_decision": "candidate_for_specialist_shadow_review" if all(release_gates.values()) else "hold", 16} 17 18print("bundle:", receipt["candidate_bundle"], "rollback:", receipt["previous_bundle"]) 19print("routes/outcomes:", receipt["route_traces"], receipt["later_specialist_outcomes"], receipt["joined_usable_cases"]) 20print("delayed_quality:", receipt["delayed_quality"]) 21print("source_slices:", receipt["source_slices"]) 22print("release_gates_pass:", all(receipt["release_gates"].values())) 23print("candidate_decision:", receipt["candidate_decision"]) 24assert all(receipt["release_gates"].values()) 25assert receipt["delayed_quality"]["priority_precision"] == 0.5 26assert receipt["delayed_quality"]["priority_recall"] == 0.5 27assert receipt["candidate_decision"] == "candidate_for_specialist_shadow_review"
Output
1bundle: damage-cnn-v1 rollback: damage-cnn-v0 2routes/outcomes: 7 4 4 3delayed_quality: {'priority_precision': 0.5, 'priority_recall': 0.5, 'reviewed_usable_cases': 4, 'held_with_small_n': True} 4source_slices: {'customer_phone': {'photos': 4, 'request_new_photo_rate': 0.25}, 'warehouse_camera': {'photos': 3, 'request_new_photo_rate': 0.333}} 5release_gates_pass: True 6candidate_decision: candidate_for_specialist_shadow_review

candidate_for_specialist_shadow_review is narrower than launch approval. The receipt says this frozen bundle deserves comparison beside current production routing. It doesn't claim that four usable local cases prove every camera, lighting condition, package type, or damage shape.

Those delayed counts are case-level, not photo-level. A true positive (TP) is a priority route the specialist later confirmed; a false positive (FP) is a priority route they rejected; a false negative (FN) is a damaged case left on the normal path; a true negative (TN) is a correctly quiet case. The four joined usable test cases are:

CaseUsable actionSpecialistCell
R-404priority_damage_reviewdamagedTP
R-405normal_reviewclearTN
R-407priority_damage_reviewclearFP
R-408normal_reviewdamagedFN

Priority precision is 1/(1+1)=0.501 / (1 + 1) = 0.501/(1+1)=0.50 and recall is 1/(1+1)=0.501 / (1 + 1) = 0.501/(1+1)=0.50. R-404-a and R-404-b collapse to one case. Unusable photos never enter this table.

Practice: break the vision contract

Use the runnable examples as a release harness. Change one condition at a time, predict the failure, then rerun the examples.

  1. Change R-401-b split from train to validation. Which dataset gate fails?
  2. Change R-406-a brightness in MODEL_OUTPUTS from 0.12 to 0.30. Why does this expose a quality-detector failure rather than prove the escalation is safe?
  3. Remove damage_threshold from the response trace. Which receipt gate fails?
  4. Change R-408-a damage probability in MODEL_OUTPUTS from 0.63 to 0.75. Which delayed metrics improve?
  5. Set BUNDLE["previous_bundle"] = "". Why is shadow evidence no longer promotion-ready?
  6. Replace SPECIALIST_OUTCOMES with an empty list. Why do evidence gates fail instead of crashing or passing?
  7. Add SpecialistOutcome("R-999", 12, True, "S-12", "visible-damage-v1"). Which join gate fails?
  8. Change R-408 specialist outcome to guideline visible-damage-v2. Which provenance gate fails?

Practice answer sketches

Which gate fails when R-401-b moves to validation?

Answer

case_groups_do_not_cross_splits fails because two photos of physical return R-401 now cross training and validation. Evaluation can reward memorization of one package.

Why isn't raising R-406-a brightness a safe escalation?

Answer

The router now sees an apparently usable high-score image, but specialist label still says the evidence is unusable. Both unusable_photos_abstain and unsafe_priority_routes_absent fail. That mismatch belongs in quality-model evaluation before promotion.

Which gate fails when route traces omit damage_threshold?

Answer

route_traces_replayable fails. A reviewer can't reconstruct whether a policy threshold change altered the route.

What changes when R-408-a score rises from 0.63 to 0.75?

Answer

It becomes a true-positive priority route. Case-level recall rises from 1 / 2 to 2 / 2 = 1.0, and precision rises from 1 / 2 to 2 / 3 = 0.667.

Why keep previous_bundle beside candidate?

Answer

Shadow evidence may justify a later promotion, but promotion still needs a known rollback target. A candidate without previous alias doesn't prove rollback can restore earlier behavior.

What happens when no specialist outcomes have arrived?

Answer

Precision and recall become None, then outcome-completeness and metric-evidence gates fail. No denominator means no evidence. It isn't a crash, measured success, or measured failure.

Which gate fails when outcome R-999 appears?

Answer

specialist_outcomes_join_usable_test_cases fails. A delayed specialist label can't become evaluation evidence unless it joins a usable test case routed by this bundle.

Which gate fails when R-408 uses visible-damage-v2?

Answer

specialist_outcome_label_contract_matches_bundle fails. Delayed metrics only support this bundle when every outcome records a reviewer and uses the bundle's visible-damage-v1 label contract.

Complete the lesson

Mastery Check

Answer every question, then check your score. Score 75% or higher to mark this lesson complete.

1.A frozen photo manifest has two photos from the same physical return, R-401-a and R-401-b. R-401-a stays in train, but R-401-b is edited to validation. Which release gate should fail?

Correct answer: case_groups_do_not_cross_splits fails because one physical case now appears in both train and validation.

The invariant is grouped evaluation by physical case. Photos of the same returned package may have nearly identical pixels, so putting one in train and one in validation can reward memorization. Usable labels and reviewer lineage are still present, and the chronological split can still pass.

2.R-406-a has damage_probability=0.94, blur_score=0.12, box_visible=True, and a manifest quality_label of unusable. If its model brightness is edited from 0.12 to 0.30 with min_brightness=0.20 and damage threshold 0.70, why should promotion still hold?

Correct answer: The edit makes the router treat an unusable high-score photo as usable, causing an unsupported priority route.

With brightness raised above the minimum, the router passes the quality check and the 0.94 score crosses the damage threshold. But the frozen manifest still marks the evidence unusable, so the abstention contract is broken. This exposes a quality-detector failure, not proof that the damage escalation is safe.

3.A route trace still stores photo ID, bundle ID, preprocessing, quality scores, damage probability, and action, but omits the damage_threshold used to choose priority_damage_review. Which gate should fail?

Correct answer: route_traces_replayable fails because a reviewer cannot reconstruct whether the stored policy threshold produced the route.

Replayability requires the policy inputs needed to reproduce the decision, including the damage threshold. Missing that field doesn't by itself change route IDs, make every priority route unsafe, or prevent counting delayed outcomes against already recorded actions.

4.Use the case policy that a return case is priority_damage_review if any usable photo for that case crosses the threshold. Original delayed quality has TP=1, FP=1, and FN=1 across four usable test cases. R-408-a is confirmed damaged but scores 0.63, below the 0.70 threshold. If its score changes to 0.75, what are the new priority precision and recall?

Correct answer: Precision is 0.667 and recall is 1.0 because the previous false negative becomes a true positive.

Raising R-408-a above the threshold makes the R-408 case a priority route. It was a confirmed damaged case previously counted as a false negative. The counts become TP=2, FP=1, FN=0, so precision is 2 / (2 + 1) = 0.667 and recall is 2 / (2 + 0) = 1.0.

5.Same CNN weights are repackaged, but preprocessing changes from a 224 center crop to full-frame resize, and previous_bundle is left empty. What must be fixed before promotion?

Correct answer: Publish a versioned bundle, pin preprocessing, re-evaluate it, and record a rollback target.

Preprocessing is part of the deployed model behavior. A crop or resize change can alter whether visible damage remains in the image, so it needs a versioned bundle and reevaluation. An empty previous_bundle also removes the recorded rollback target required for safe promotion.

6.A warehouse installs a new camera. During the first day, the endpoint logs a much higher request_new_photo rate from warehouse_camera because many images are dark or missing the package, but no specialist-confirmed outcomes have arrived yet. What should operators do under the monitoring contract?

Correct answer: Treat it as an input-quality issue, investigate capture, or route to manual intake until delayed outcomes arrive.

Dark images and missing packages are immediate input-quality signals, not delayed confirmation metrics. The trigger is to investigate the capture path or fail to manual intake. Source drift can motivate investigation or a candidate run, but it should not automatically replace production, and precision requires later specialist outcomes.

7.A usable customer photo passes quality checks and has damage_probability=0.91 with a damage threshold of 0.70. The customer also asks for an instant refund. What should the vision service's decision support do?

Correct answer: Route priority review with score and trace; refund eligibility stays with order rules.

For a usable photo above threshold, the route is priority_damage_review. That is still only triage for human review, not a refund decision. Eligibility depends on order ownership, item type, return window, and specialist judgment, and later outcomes should be appended separately rather than written into the original route trace.

8.After routing seven test photos, specialists later review the cases. A developer proposes overwriting each original route trace with confirmed_damage and the final case state, then computing precision from those edited traces. What should the pipeline do instead?

Correct answer: Keep traces immutable; append dated specialist outcomes and join them by case.

The original route trace is evidence of what the candidate bundle decided before specialist review. Appending outcomes separately preserves arrival time, exposes missing or orphaned joins, and prevents hindsight from rewriting the decision that was actually served.

9.All photos from each physical return already stay in one split. Cases span capture days 1 through 100. Which split plan preserves the required time direction?

Correct answer: Train on the earliest cases, validate on later cases, and test on the latest cases.

Grouping prevents near-duplicate photos of one package from crossing boundaries, but it doesn't by itself preserve temporal evaluation. Training must use earlier cases, with validation and test containing progressively later cases, so evaluation represents performance on future returns.

10.Aggregate precision is unchanged, but small tears in dark customer-phone photos are being missed after a capture-path change. Which analysis can expose the operational failure?

Correct answer: Compare performance by source, lighting, packaging, and defect size under the frozen split.

A global metric can remain stable while an important subgroup fails. Evaluating the frozen holdout by image source, lighting, packaging, and visible-defect size reveals whether dark customer uploads with small tears have specifically degraded without introducing case leakage or changing policy first.

10 questions remaining.

Next Step
Continue to Capstone: Production ML Pipeline

You now have four shippable artifacts (ETA, ranking, forecast, and this vision bundle) each with its own action gate and rollback pointer. Next you'll put them under one validated promotion, monitoring, and rollback workflow so a reviewer can trace any live decision back to a receipt.

PreviousCapstone: Demand Forecasting
Share this article
XFacebookLinkedInBlueskyRedditHacker NewsEmail
References

Model Cards for Model Reporting

Mitchell, M., Wu, S., Zaldivar, A., et al. · 2019 · FAT* 2019

https://arxiv.org/abs/1810.03993

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Dosovitskiy, A., et al. · 2020 · ICLR 2021

https://arxiv.org/abs/2010.11929

On Calibration of Modern Neural Networks

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. · 2017

https://arxiv.org/abs/1706.04599

MLOps: Continuous Delivery and Automation Pipelines in Machine Learning.

Google Cloud. · 2026 · Official documentation

https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning

Discussion

Questions and insights from fellow learners.

Discussion loads when you reach this section.