LeetLLM
My PlanLearnGlossaryTracksPracticeBlog
LeetLLM

Your go-to resource for mastering AI & LLM systems.

Product

  • Learn
  • Glossary
  • Tracks
  • Practice
  • Blog
  • RSS

Legal

  • Terms of Service
  • Privacy Policy

© 2026 LeetLLM. All rights reserved.

All Topics
Your Progress
0%

0 of 177 articles completed

🛠️Computing Foundations0/9
Git, Shell, Linux for AIDocker for Reproducible AIPython for AI EngineeringNumPy and Tensor ShapesCUDA for ML TrainingMPS & Metal for ML on MacData Structures for AISQL and Data ModelingAlgorithms for ML Engineers
📊Math & Statistics0/8
Gradients and BackpropVectors, Matrices & TensorsLinear Algebra for MLAdam, Momentum, SchedulersProbability for Machine LearningStatistics and UncertaintyDistributions and SamplingHypothesis Tests, Intervals, and pass@k
📚Preparation & Prerequisites0/13
Neural Networks from ScratchCNNs from ScratchTraining & BackpropagationSoftmax, Cross-Entropy & OptimizationRNNs, LSTMs, GRUs, and Sequence ModelingAutoencoders and VAEsThe Transformer Architecture End-to-EndLanguage Modeling & Next TokensFrom GPT to Modern LLMsPrompt Engineering FundamentalsCalling LLM APIs in ProductionFirst AI App End-to-EndThe LLM Lifecycle
🧮ML Algorithms & Evaluation0/11
Linear Regression from ScratchLogistic Regression and MetricsDecision Trees, Forests, and BoostingReinforcement Learning BasicsValidation and LeakageClustering and PCACore Retrieval AlgorithmsDecoding AlgorithmsExperiment Design and A/B TestingPyTorch Training LoopsDataset Pipelines and Data Quality
📦Production ML Systems0/6
Feature Engineering for Production MLBatch and Streaming Feature PipelinesGradient Boosted Trees in ProductionRanking and Recommendation SystemsForecasting and Anomaly DetectionMonitoring Predictive Models
🧪Core LLM Foundations0/8
The Bitter Lesson & ComputeBPE, WordPiece, and SentencePieceStatic to Contextual EmbeddingsPerplexity & Model EvaluationFile Ingestion for AIChunking StrategiesLLM Benchmarks & LimitationsInstruction Tuning & Chat Templates
🧰Applied LLM Engineering0/24
Dimensionality Reduction for EmbeddingsCoT, ToT & Self-Consistency PromptingFunction Calling & Tool UseMCP & Tool Protocol StandardsContext EngineeringPrompt Injection DefenseResponsible AI GovernanceData Labeling and Human FeedbackEvaluating AI AgentsProduction RAG PipelinesHybrid Search: Dense + SparseReranking and Cross-Encoders for RAGRAG Evaluation for Reliable AnswersLLM-as-a-Judge EvaluationBias & Fairness in LLMsHallucination Detection & MitigationLLM Observability & MonitoringExperiment Tracking with MLflow and W&BPrompt Optimization with DSPyModel Versioning & DeploymentSemantic Caching & Cost OptimizationLLM Cost Engineering & Token EconomicsModel Gateways, Routing, and FallbacksDesign an Automated Support Agent
🎓Portfolio Capstones0/9
Capstone: Delivery ETA PredictionCapstone: Product RankingCapstone: Demand ForecastingCapstone: Image Damage ClassifierCapstone: Production ML PipelineCapstone: Document QACapstone: Eval DashboardCapstone: Fine-Tuned ClassifierCapstone: Reproducible ML Study
🧠Transformer Deep Dives0/8
Sentence Embeddings & Contrastive LossEmbedding Similarity & QuantizationScaled Dot-Product AttentionVision Transformers and Image EncodersPositional Encoding: RoPE & ALiBiLayer Normalization: Pre-LN vs Post-LNMechanistic InterpretabilityDecoding Strategies: Greedy to Nucleus
🧬Advanced Training & Adaptation0/16
Scaling Laws & Compute-Optimal TrainingPre-training Data at ScaleBuild GPT from Scratch LabJAX for PyTorch ResearchersContinued Pretraining for Domain ShiftSynthetic Data PipelinesSupervised Fine-Tuning PipelineMixed Precision TrainingDistributed Training: FSDP & ZeROLoRA & Parameter-Efficient TuningReward Modeling from Preference DataRLHF & DPO AlignmentConstitutional AI & Red TeamingRLVR & Verifiable RewardsKnowledge Distillation for LLMsModel Merging and Weight Interpolation
🤖Advanced Agents & Retrieval0/16
Vector DB Internals: HNSW & IVFAdvanced RAG: HyDE & Self-RAGGraphRAG & Knowledge GraphsRAG Security & Access ControlStructured Output GenerationReAct & Plan-and-ExecuteGuardrails & Safety FiltersCode Generation & SandboxingComputer-Use / GUI / Browser AgentsHuman-in-the-Loop Agent ArchitectureAI Coding Workflow with AgentsAgent Memory & PersistenceAgent Failure & RecoveryRecursive Language Models (RLM)Multi-Agent OrchestrationCapstone: Production Agent
⚡Inference & Production Scale0/19
Inference: TTFT, TPS & KV CacheMulti-Query & Grouped-Query AttentionKV Cache & PagedAttentionPrefix Caching and Prompt CachingFlashAttention & Memory EfficiencyContinuous Batching & SchedulingScaling LLM InferenceModel Parallelism for LLM InferenceModel Quantization: GPTQ, AWQ & GGUFLocal LLM DeploymentSLM Specialization & Edge DeploymentSpeculative DecodingLong Context Window ManagementMixture of Experts ArchitectureMamba & State Space ModelsReasoning & Test-Time ComputeAdvanced MLOps & DevOps for AIGPU Serving & AutoscalingA/B Testing for LLMs
🏗️System Design Capstones0/9
Content Moderation SystemCode Completion SystemMulti-Tenant LLM PlatformLLM-Powered Search EngineVision-Language Models & CLIPMultimodal LLM ArchitectureDiffusion Models: Images & TextReal-Time Voice AI AgentReasoning Agent System Design
🎤AI Lab Interviewing0/4
AI Lab Coding Interview: Python SystemsAI Lab System Design InterviewAI Lab Behavioral InterviewAI Lab Technical Presentation
🔬Project Deep Dives0/17
Deep Dive - vLLMDeep Dive - SkyRLDeep Dive - FlashAttentionDeep Dive - FlashInferDeep Dive - DeepGEMMDeep Dive - NCCLDeep Dive - MegatronDeep Dive - DeepSpeedDeep Dive - RayDeep Dive - MLflowDeep Dive - PyTorchDeep Dive - TransformersDeep Dive - SGLangDeep Dive - slimeDeep Dive - DeepEPDeep Dive - TinkerDeep Dive - Light-PEFT
Back to Topics
LearnPortfolio CapstonesCapstone: Demand Forecasting
⚙️HardMLOps & Deployment

Capstone: Demand Forecasting

Ship a demand forecast and capacity-alert artifact with rolling backtests, alert review, and retraining policy.

14 min read
Learning path
Step 82 of 177 in the full curriculum
Capstone: Product RankingCapstone: Image Damage Classifier

Personalize this lesson

Adapt explanations and teaching visuals to your background and preferred voice.

Warehouse teams need to forecast daily parcel volume before staffing and packing decisions are locked. A point forecast alone isn't enough: planners also need to know when actual demand falls outside an expected range and which evidence supports an alert.

This capstone ships a forecast and alert artifact. It doesn't automatically hire labor, move inventory, or page an operator on every miss. It creates a versioned expectation, detects unusually large forecast errors, and records evidence for a planner's human-in-the-loop decision.

Earlier, you built a same-weekday seasonal baseline and separated forecast errors from anomaly review. This capstone packages those mechanics into a release candidate with shadow evidence and a rollback pointer.

Two rolling-origin forecast windows for fulfillment center FC-A show a dashed seasonal baseline, a candidate line, a shaded plus-or-minus-six range, and later actuals. The February 6 campaign point is the only range breach: candidate 140, range 134 to 146, actual 160. A compact evidence strip records 14 issued and joined rows, 13 of 14 observations in range, policy review scores of 0.667, a warehouse-demand-v1 shadow decision, and warehouse-demand-v0 as rollback.
Read the dashed baseline, candidate line, and shaded plus-or-minus-six range across two frozen origins. The February 6 breach routes to review, while the evidence strip keeps joins, coverage, review scores, shadow decision, and rollback separate.

Choose the Series and Decision

Predict daily shipped parcels for each warehouse seven days ahead. Use an explicit planning contract:

FieldContract
entityfulfillment center and shipping service tier
targetparcels shipped per calendar day
horizonnext seven days
decisionplanner reviews capacity when forecast or alert requires it
baselinesame weekday from prior week
evaluationMAE plus underforecast cost by high-volume slice

Demand can change around promotions, holidays, seller campaigns, inventory shortages, and data outages. Those known drivers should appear as features only if they are scheduled and available before the forecast cutoff.

Hyndman and Athanasopoulos explain why forecast evaluation must use later observations and rolling forecasting origins rather than random splits.[1]Reference 1Forecasting: Principles and Practice, Third Edition.https://otexts.com/fpp3/ For this project, each backtest run records its training cutoff, horizon, model version, and the actual values that arrived afterward.

Diagram showing History snapshot through cutoff, Immutable forecast rows baseline + candidate, Join later observations rolling-origin metrics, and Alert queue range + owner + resolution.
History snapshot through cutoff, Immutable forecast rows baseline + candidate, Join later observations rolling-origin metrics, and Alert queue range + owner + resolution.

Build a Reviewable Artifact

Your repository surface should look like:

text
1demand-forecast/ 2 data/ 3 warehouse_daily_counts.parquet 4 planned_events.json 5 split_manifest.json 6 forecasting/ 7 seasonal_baseline.py 8 train_candidate.py 9 rolling_backtest.py 10 alerts/ 11 forecast_error_policy.json 12 evaluate_alerts.py 13 reports/ 14 backtest_metrics.json 15 alert_review.csv 16 tests/ 17 test_future_rows_excluded.py 18 test_observation_join.py 19 test_alert_contract.py

The candidate can be a tree model over lag features, rolling means, service tier, weekday, and known promotions. It must beat the seasonal baseline on later windows, especially where underforecasting is expensive. A candidate that marginally improves MAE but misses peak-volume days should remain blocked.

Freeze issued forecasts before outcomes arrive

The runnable receipt below starts after training. Seasonal baseline and candidate point forecasts are pre-filled fixtures so this lesson grades freeze, join, alert, and receipt contracts. Fitting lag features under a rolling origin is a separate training lab; don't claim the harness trained a tree.

It freezes two seven-day forecast windows issued one week apart. Each issued row stores the cutoff, target date, horizon, baseline, candidate, expected range, and known event context before the target day arrives.

Actual parcel counts belong in a separate append-only observation stream. The backtest joins each later observation to the immutable forecast it evaluates. A real pipeline should materialize the same boundary after every rolling-origin fold.

StreamStored before replayArrival time
issued forecastcenter, tier, issue date, target date, horizon, model outputs, range, known-event contextbefore target day
observationforecast ID, observed date, actual parcel counton or after target day
alertjoined forecast ID, range breach, policies, owner, resolutionafter observation join
01-freeze-forecast-contract.py
1from dataclasses import dataclass, replace 2from datetime import date 3import json 4 5@dataclass(frozen=True) 6class IssuedForecast: 7 forecast_id: str 8 fold: str 9 issued_at: date 10 target_date: date 11 horizon_day: int 12 center: str 13 service_tier: str 14 baseline: int 15 candidate: int 16 lower: int 17 upper: int 18 latest_actual_at: date 19 scheduled_event: str | None = None 20 event_known_at: date | None = None 21 22@dataclass(frozen=True) 23class Observation: 24 forecast_id: str 25 observed_at: date 26 actual: int 27 28@dataclass(frozen=True) 29class BacktestRow: 30 issued: IssuedForecast 31 observation: Observation 32 33INTERVAL_HALF_WIDTH = 6 34INTERVAL_POLICY = "candidate-plus-minus-6-v1" 35LATEST_ACTUAL_BY_FOLD = { 36 "fold-1": date(2026, 1, 31), 37 "fold-2": date(2026, 2, 7), 38} 39 40def issue_forecast( 41 fold: str, 42 issued_at: date, 43 target_date: date, 44 center: str, 45 service_tier: str, 46 baseline: int, 47 candidate: int, 48 scheduled_event: str | None = None, 49 event_known_at: date | None = None, 50) -> IssuedForecast: 51 horizon_day = (target_date - issued_at).days 52 return IssuedForecast( 53 forecast_id=f"{center}:{service_tier}:{issued_at}:{target_date}", 54 fold=fold, 55 issued_at=issued_at, 56 target_date=target_date, 57 horizon_day=horizon_day, 58 center=center, 59 service_tier=service_tier, 60 baseline=baseline, 61 candidate=candidate, 62 lower=candidate - INTERVAL_HALF_WIDTH, 63 upper=candidate + INTERVAL_HALF_WIDTH, 64 latest_actual_at=LATEST_ACTUAL_BY_FOLD[fold], 65 scheduled_event=scheduled_event, 66 event_known_at=event_known_at, 67 ) 68 69print("interval_policy:", INTERVAL_POLICY, "half_width:", INTERVAL_HALF_WIDTH)
Output
1interval_policy: candidate-plus-minus-6-v1 half_width: 6
02-issue-forecasts-and-observations.py
1ISSUED_FORECASTS = [ 2 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 2), "FC-A", "standard", 100, 103), 3 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 3), "FC-A", "standard", 112, 111), 4 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 4), "FC-A", "standard", 115, 117), 5 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 5), "FC-A", "standard", 118, 117), 6 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 6), "FC-A", "standard", 132, 140, "seller-campaign", date(2026, 1, 29)), 7 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 7), "FC-A", "standard", 82, 83), 8 issue_forecast("fold-1", date(2026, 2, 1), date(2026, 2, 8), "FC-A", "standard", 76, 77), 9 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 9), "FC-A", "standard", 104, 105), 10 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 10), "FC-A", "standard", 110, 112), 11 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 11), "FC-A", "standard", 119, 120), 12 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 12), "FC-A", "standard", 116, 119), 13 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 13), "FC-A", "standard", 160, 164, "seller-campaign", date(2026, 2, 5)), 14 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 14), "FC-A", "standard", 84, 85), 15 issue_forecast("fold-2", date(2026, 2, 8), date(2026, 2, 15), "FC-A", "standard", 78, 79), 16] 17 18ACTUALS = [104, 110, 119, 116, 160, 84, 78, 106, 113, 121, 118, 168, 86, 80] 19OBSERVATIONS = [ 20 Observation(forecast.forecast_id, forecast.target_date, actual) 21 for forecast, actual in zip(ISSUED_FORECASTS, ACTUALS, strict=True) 22] 23issued_by_id = {row.forecast_id: row for row in ISSUED_FORECASTS} 24observations_by_forecast_id = {row.forecast_id: row for row in OBSERVATIONS} 25 26def lag_sources_precede_cutoff(forecasts: list[IssuedForecast]) -> bool: 27 return all(row.latest_actual_at < row.issued_at for row in forecasts) 28 29leaky_forecast = replace( 30 ISSUED_FORECASTS[0], 31 latest_actual_at=ISSUED_FORECASTS[0].target_date, 32) 33assert lag_sources_precede_cutoff(ISSUED_FORECASTS) 34assert not lag_sources_precede_cutoff([leaky_forecast]) 35 36print("issued forecasts:", len(ISSUED_FORECASTS)) 37print("later observations:", len(OBSERVATIONS))
03-join-backtest-rows.py
1ROWS = [ 2 BacktestRow(forecast, observations_by_forecast_id[forecast.forecast_id]) 3 for forecast in ISSUED_FORECASTS 4 if forecast.forecast_id in observations_by_forecast_id 5] 6 7folds = sorted({row.fold for row in ISSUED_FORECASTS}) 8for fold in folds: 9 rows = [row for row in ISSUED_FORECASTS if row.fold == fold] 10 print( 11 f"{fold}: cutoff={rows[0].issued_at}", 12 f"window={rows[0].target_date}..{rows[-1].target_date}", 13 f"rows={len(rows)}", 14 ) 15print("first forecast id:", ISSUED_FORECASTS[0].forecast_id) 16print("joined backtest rows:", len(ROWS))
Output
1fold-1: cutoff=2026-02-01 window=2026-02-02..2026-02-08 rows=7 2fold-2: cutoff=2026-02-08 window=2026-02-09..2026-02-15 rows=7 3first forecast id: FC-A:standard:2026-02-01:2026-02-02 4joined backtest rows: 14

The fixture is intentionally compact: one fulfillment center and one service tier make every row inspectable. Its ID includes center, tier, issue date, and target date so additional series can't collide during replay. A production report should repeat the same contract by center, tier, horizon, and event slice.

Evaluate Forecasts and Alerts Separately

Mean absolute error (MAE) answers how far point forecasts miss on average. Capacity planning needs a second view because a large underforecast on a peak day can cost more than a small overforecast on a routine day. The next cell assigns a local cost of 3 units to each underforecast parcel on a day with at least 150 observed parcels.

An expected range answers a different question. It should contain a stated share of later observations when measured across enough held-out windows. Hyndman and Athanasopoulos describe prediction intervals as forecast ranges with a specified coverage probability and explain why distributional forecasts need their own accuracy measures.[1]Reference 1Forecasting: Principles and Practice, Third Edition.https://otexts.com/fpp3/ The local candidate ± 6 policy only demonstrates release plumbing; it isn't a calibrated production interval.

04-evaluate-point-and-range-metrics.py
1def mae(field: str) -> float: 2 return sum( 3 abs(row.observation.actual - getattr(row.issued, field)) 4 for row in ROWS 5 ) / len(ROWS) 6 7def peak_underforecast_cost(field: str) -> int: 8 return sum( 9 3 * max(row.observation.actual - getattr(row.issued, field), 0) 10 for row in ROWS 11 if row.observation.actual >= 150 12 ) 13 14def coverage() -> float: 15 return sum( 16 row.issued.lower <= row.observation.actual <= row.issued.upper 17 for row in ROWS 18 ) / len(ROWS) 19 20ALERT_POLICY = "outside-range-review-v1" 21ALERT_OWNER = "capacity-ops" 22CANDIDATE_FORECAST = "warehouse-demand-v1" 23PREVIOUS_FORECAST = "warehouse-demand-v0" 24 25forecast_metrics = { 26 "baseline_mae": round(mae("baseline"), 3), 27 "candidate_mae": round(mae("candidate"), 3), 28 "baseline_peak_underforecast_cost": peak_underforecast_cost("baseline"), 29 "candidate_peak_underforecast_cost": peak_underforecast_cost("candidate"), 30 "range_coverage": round(coverage(), 3), 31 "range_rows": len(ROWS), 32} 33 34print(json.dumps(forecast_metrics, indent=2))
05-build-range-alerts.py
1new_alerts = [ 2 { 3 "forecast_id": row.issued.forecast_id, 4 "forecast_version": CANDIDATE_FORECAST, 5 "interval_policy": INTERVAL_POLICY, 6 "policy_version": ALERT_POLICY, 7 "owner": ALERT_OWNER, 8 "issued_at": str(row.issued.issued_at), 9 "target_date": str(row.issued.target_date), 10 "horizon_day": row.issued.horizon_day, 11 "center": row.issued.center, 12 "service_tier": row.issued.service_tier, 13 "expected_range": [row.issued.lower, row.issued.upper], 14 "observed_at": str(row.observation.observed_at), 15 "observed": row.observation.actual, 16 "error": row.observation.actual - row.issued.candidate, 17 "scheduled_event": row.issued.scheduled_event, 18 "event_known_at": str(row.issued.event_known_at) if row.issued.event_known_at else None, 19 "resolution": None, 20 } 21 for row in ROWS 22 if not row.issued.lower <= row.observation.actual <= row.issued.upper 23] 24 25print("new_alert_count:", len(new_alerts)) 26print( 27 "first_alert:", 28 {key: new_alerts[0][key] for key in ("forecast_id", "expected_range", "observed", "error", "scheduled_event", "owner")}, 29)
Output
1new_alert_count: 1 2first_alert: {'forecast_id': 'FC-A:standard:2026-02-01:2026-02-06', 'expected_range': [134, 146], 'observed': 160, 'error': 20, 'scheduled_event': 'seller-campaign', 'owner': 'capacity-ops'}

The candidate improves average error and peak-day cost, but one seller-campaign day still escapes its expected range. That row belongs in a planner queue. The queue preserves which immutable forecast was issued, what happened later, who owns review, and which interval and alert policies created the alert.

Publish a Shadow-Review Receipt

Forecast quality and alert usefulness require separate evidence. A narrow range can create an exhausting queue. A broad range can hide events a planner needed to see. Measure alert precision and recall on historical alerts after reviewers label whether each event required action. Keep the forecast and alert-policy versions on those review rows; otherwise a candidate could pass using evidence produced by a different model or threshold.

The final cell publishes one candidate receipt. It advances to planner shadow review, not production replacement. The previous forecast alias remains explicit so a later promotion process can roll back cleanly.

06-review-historical-alerts.py
1# Bind at least one reviewed row to the same backtest breach the harness produced. 2backtest_breach_ids = {alert["forecast_id"] for alert in new_alerts} 3REVIEWED_ALERTS = [ 4 { 5 "alert_id": "A-101", 6 "forecast_id": "FC-A:standard:2026-02-01:2026-02-06", 7 "forecast_version": CANDIDATE_FORECAST, 8 "policy_version": ALERT_POLICY, 9 "triggered": True, 10 "actionable": True, 11 }, 12 { 13 "alert_id": "A-102", 14 "forecast_id": "FC-A:standard:2026-01-25:2026-01-30", 15 "forecast_version": CANDIDATE_FORECAST, 16 "policy_version": ALERT_POLICY, 17 "triggered": True, 18 "actionable": True, 19 }, 20 { 21 "alert_id": "A-103", 22 "forecast_id": "FC-A:standard:2026-01-25:2026-01-28", 23 "forecast_version": CANDIDATE_FORECAST, 24 "policy_version": ALERT_POLICY, 25 "triggered": True, 26 "actionable": False, 27 }, 28 { 29 "alert_id": "A-104", 30 "forecast_id": "FC-A:standard:2026-02-01:2026-02-03", 31 "forecast_version": CANDIDATE_FORECAST, 32 "policy_version": ALERT_POLICY, 33 "triggered": False, 34 "actionable": True, 35 }, 36 { 37 "alert_id": "A-105", 38 "forecast_id": "FC-A:standard:2026-02-01:2026-02-04", 39 "forecast_version": CANDIDATE_FORECAST, 40 "policy_version": ALERT_POLICY, 41 "triggered": False, 42 "actionable": False, 43 }, 44] 45 46true_positives = sum(row["triggered"] and row["actionable"] for row in REVIEWED_ALERTS) 47false_positives = sum(row["triggered"] and not row["actionable"] for row in REVIEWED_ALERTS) 48false_negatives = sum(not row["triggered"] and row["actionable"] for row in REVIEWED_ALERTS) 49 50def rate_or_none(numerator: int, denominator: int) -> float | None: 51 return round(numerator / denominator, 3) if denominator else None 52 53alert_review = { 54 "forecast_version": CANDIDATE_FORECAST, 55 "policy_version": ALERT_POLICY, 56 "precision": rate_or_none(true_positives, true_positives + false_positives), 57 "recall": rate_or_none(true_positives, true_positives + false_negatives), 58 "reviewed_rows": len(REVIEWED_ALERTS), 59} 60 61print("alert_review:", alert_review)
07-backtest-release-gates.py
1required_alert_fields = { 2 "forecast_id", "forecast_version", "interval_policy", "policy_version", "owner", 3 "issued_at", "target_date", "horizon_day", "center", "service_tier", 4 "expected_range", "observed_at", "observed", "error", "scheduled_event", 5 "event_known_at", "resolution", 6} 7release_gates = { 8 "issued_forecast_ids_unique": len(issued_by_id) == len(ISSUED_FORECASTS), 9 "observation_forecast_ids_unique": len(observations_by_forecast_id) == len(OBSERVATIONS), 10 "observations_join_issued_forecasts": all( 11 row.forecast_id in issued_by_id 12 for row in OBSERVATIONS 13 ), 14 "issued_forecasts_have_observations": len(ROWS) == len(ISSUED_FORECASTS), 15 "targets_after_cutoff": all(row.issued.issued_at < row.issued.target_date for row in ROWS), 16 "horizons_match_dates": all( 17 row.issued.horizon_day == (row.issued.target_date - row.issued.issued_at).days 18 for row in ROWS 19 ), 20 "horizons_within_seven_day_contract": all( 21 1 <= row.issued.horizon_day <= 7 22 for row in ROWS 23 ), 24 "observations_arrive_on_or_after_target": all( 25 row.observation.observed_at >= row.issued.target_date 26 for row in ROWS 27 ), 28 "scheduled_events_known_by_cutoff": all( 29 row.event_known_at is None or row.event_known_at <= row.issued_at 30 for row in ISSUED_FORECASTS 31 ), 32 "multiple_rolling_origins": len({row.issued_at for row in ISSUED_FORECASTS}) >= 2, 33 "candidate_beats_baseline_mae": forecast_metrics["candidate_mae"] < forecast_metrics["baseline_mae"], 34 "candidate_reduces_peak_underforecast_cost": ( 35 forecast_metrics["candidate_peak_underforecast_cost"] 36 < forecast_metrics["baseline_peak_underforecast_cost"] 37 ), 38 # ±6 width is local plumbing, not a calibrated prediction interval. 39 "local_range_plumbing_coverage_at_least_0_85": forecast_metrics["range_coverage"] >= 0.85, 40 "shadow_alert_count_at_most_3": len(new_alerts) <= 3, 41 "alert_review_matches_candidate_policy": all( 42 row["forecast_version"] == CANDIDATE_FORECAST 43 and row["policy_version"] == ALERT_POLICY 44 for row in REVIEWED_ALERTS 45 ), 46 "alert_review_bound_to_backtest_breach": bool( 47 backtest_breach_ids 48 and any(row.get("forecast_id") in backtest_breach_ids for row in REVIEWED_ALERTS) 49 ), 50 "alert_precision_evidence_at_least_0_60": ( 51 alert_review["precision"] is not None 52 and alert_review["precision"] >= 0.60 53 ), 54 "alert_recall_evidence_at_least_0_60": ( 55 alert_review["recall"] is not None 56 and alert_review["recall"] >= 0.60 57 ), 58 "alert_rows_replayable": all(required_alert_fields <= row.keys() for row in new_alerts), 59 # Validate stored lag-source timestamps, not a date manufactured by the gate. 60 "lag_k_uses_only_pre_cutoff_actuals": lag_sources_precede_cutoff(ISSUED_FORECASTS), 61 "rollback_pointer_recorded": bool(PREVIOUS_FORECAST), 62} 63 64print("release_gates_pass:", all(release_gates.values()))
08-assemble-shadow-receipt.py
1receipt = { 2 "candidate_forecast": CANDIDATE_FORECAST, 3 "previous_forecast": PREVIOUS_FORECAST, 4 "latest_rolling_origin": str(max(row.issued_at for row in ISSUED_FORECASTS)), 5 "interval_policy": INTERVAL_POLICY, 6 "alert_policy": ALERT_POLICY, 7 "owner": ALERT_OWNER, 8 "replay": { 9 "issued_forecast_rows": len(ISSUED_FORECASTS), 10 "joined_observation_rows": len(ROWS), 11 }, 12 "backtest": forecast_metrics, 13 "alert_review": alert_review, 14 "release_gates": release_gates, 15 "candidate_decision": "candidate_for_planner_shadow_review" if all(release_gates.values()) else "hold", 16} 17 18print("candidate_decision:", receipt["candidate_decision"])
09-verify-receipt-fields.py
1assert receipt["candidate_forecast"] == CANDIDATE_FORECAST 2assert receipt["previous_forecast"] == PREVIOUS_FORECAST 3assert receipt["backtest"]["candidate_mae"] < receipt["backtest"]["baseline_mae"] 4print("receipt keys:", sorted(receipt))
10-publish-shadow-review-receipt.py
1print("replay:", receipt["replay"]) 2print("backtest:", receipt["backtest"]) 3print("alert_review:", receipt["alert_review"]) 4print("release_gates_pass:", all(receipt["release_gates"].values())) 5print("rollback:", receipt["previous_forecast"]) 6print("candidate_decision:", receipt["candidate_decision"])
Output
1replay: {'issued_forecast_rows': 14, 'joined_observation_rows': 14} 2backtest: {'baseline_mae': 4.643, 'candidate_mae': 2.643, 'baseline_peak_underforecast_cost': 108, 'candidate_peak_underforecast_cost': 72, 'range_coverage': 0.929, 'range_rows': 14} 3alert_review: {'forecast_version': 'warehouse-demand-v1', 'policy_version': 'outside-range-review-v1', 'precision': 0.667, 'recall': 0.667, 'reviewed_rows': 5} 4release_gates_pass: True 5rollback: warehouse-demand-v0 6candidate_decision: candidate_for_planner_shadow_review

candidate_for_planner_shadow_review is intentionally narrower than launch approval. Frozen forecasts, later observation joins, and reviewed alerts say this bundle deserves planner observation beside current production forecasts. They don't prove that every center, tier, event slice, or future week will behave well. A reviewed-alert window with no triggered or actionable rows reports None, not an invented precision or recall score.

Plan Refresh and Monitoring

New outcomes arrive daily, but model replacement should happen on a scheduled or triggered review cycle. Store:

Operational itemRequired decision
daily observation joinappend actual count, then join it to immutable issued forecast ID
weekly accuracy reportcompare baseline, production, and shadow candidate by slice
range coverage reportmeasure later-window coverage by center, tier, and horizon
alert resolution reviewclassify actionable, expected, or data issue
retraining triggerinvestigate sustained cost regression before fitting replacement
promotion gatererun rolling backtest, protected slices, shadow review, and rollback check

Practice: break the forecast contract

Use the runnable examples as a release harness. Change one condition at a time, predict the failure, then rerun the examples.

  1. Change every fold-2 issue date from 2026-02-08 to 2026-02-15. Which temporal gates fail?
  2. Change fold-1 Friday candidate from 140 to 132. Which cost gets worse even though only one row changed?
  3. Set INTERVAL_HALF_WIDTH = 0. Why can MAE stay unchanged while range and queue gates fail?
  4. Mark A-104 as not actionable. Which alert metric improves, and which stays unchanged?
  5. Set PREVIOUS_FORECAST = "". Which executable gate fails?
  6. Give first observation forecast ID missing:standard:2026-02-01:2026-02-02. Which replay gate fails?
  7. Set every reviewed alert's triggered and actionable values to False. Why do alert-evidence gates fail instead of crashing or passing?
  8. Change every reviewed alert's policy_version to outside-range-review-v0. Which provenance gate fails?

Practice answer sketches

Which gates fail when fold-2 is issued on 2026-02-15?

Answer

targets_after_cutoff fails because targets from February 9 through February 15 no longer occur after their issue cutoff. horizons_within_seven_day_contract fails too because those derived horizons are no longer between 1 and 7. That is hindsight, not a backtest.

What changes when the first Friday candidate falls from 140 to 132?

Answer

Candidate MAE rises by 8 / 14, but peak underforecast cost rises by 3 * 8 = 24. The slice metric makes capacity risk visible.

Why do zero-width expected ranges fail local range and queue gates without changing MAE?

Answer

MAE reads point forecasts only. Local coverage and queue volume read whether observations stay inside expected ranges. Different artifacts answer different questions.

What changes when A-104 becomes not actionable?

Answer

Recall rises from 2 / 3 to 2 / 2 = 1.0 because the former missed alert is no longer actionable. Precision stays 2 / 3 because triggered rows don't change.

Which gate fails when PREVIOUS_FORECAST becomes an empty string?

Answer

rollback_pointer_recorded fails. Shadow evidence may justify a later promotion, but promotion still needs a known rollback target.

Which gate fails when an observation points at forecast ID missing:standard:2026-02-01:2026-02-02?

Answer

observations_join_issued_forecasts fails. A later count can't become evaluation evidence unless it joins one immutable forecast issued before the target day.

What happens when reviewed history contains no triggered or actionable alerts?

Answer

Precision and recall become None, then both alert-evidence gates fail. No denominator means no evidence. It isn't a crash, measured success, or measured failure.

Which gate fails when reviewed alerts come from outside-range-review-v0?

Answer

alert_review_matches_candidate_policy fails. Precision and recall only support this release when every reviewed row was produced by the candidate forecast and outside-range-review-v1 policy under review.

Complete the lesson

Mastery Check

Answer every question, then check your score. Score 75% or higher to mark this lesson complete.

1.Fold 2 was originally issued on 2026-02-08 for target dates 2026-02-09 through 2026-02-15. If those same target rows are instead issued on 2026-02-15, with IDs and observations regenerated from the modified issued rows, which release gates fail?

Correct answer: targets_after_cutoff and horizons_within_seven_day_contract, because the targets are no longer after the cutoff and the recomputed horizons are zero or negative.

A rolling-origin row is valid only when the forecast is issued before the target day and the horizon stays inside the seven-day contract. Issuing Feb 9 through Feb 15 targets on Feb 15 creates hindsight rows with horizons from -6 through 0. Regenerating IDs and observations avoids join problems, so the temporal gates are the decisive failures.

2.The campaign-day row has actual volume 160 and candidate forecast 140. Peak underforecast cost is 3 * max(actual - forecast, 0) for days with actual at least 150. If that one candidate forecast is changed from 140 to 132 across 14 backtest rows, what changes?

Correct answer: Candidate MAE increases by 8 / 14, and peak underforecast cost increases by 24 because the high-volume underforecast grew by 8 parcels.

The absolute error on that row grows from 20 to 28, adding 8 total error over 14 rows, so MAE rises by 8 / 14. Because the actual volume is at least 150, the extra 8 underforecasted parcels are priced at 3 each, adding 24 to the peak underforecast cost.

3.The candidate point forecasts are left unchanged, but the expected range policy is changed from candidate +/- 6 to zero width, so every range is [candidate, candidate]. In the fixture, no actual count exactly equals its candidate forecast. Why can point-forecast MAE stay the same while range and queue gates fail?

Correct answer: MAE uses only the candidate point value, while coverage and alert count use whether each later observation falls inside the expected range.

Changing interval width does not change the candidate point forecasts, so the MAE formula sees the same inputs. Coverage and alert creation are separate checks against lower <= actual <= upper; with zero-width ranges and no exact hits, coverage collapses and every row becomes a range breach.

4.A fold is issued on 2026-02-01 for target dates 2026-02-02 through 2026-02-08. The candidate may use event features only when they are scheduled and available before the forecast cutoff. Which feature value is valid to freeze on the issued forecast row for the 2026-02-06 target?

Correct answer: A seller-campaign flag for the February 6 target with event_known_at = 2026-01-29.

Issued forecast rows must contain only information available before the cutoff. A campaign known on January 29 can be frozen into a February 1 forecast for February 6. A later-discovered campaign, the actual parcel count, and the post-observation breach flag all depend on information unavailable at issue time.

5.A candidate has the same passing MAE, peak-cost, range-coverage, queue, and alert-review evidence as the receipt, but the receipt sets previous_forecast to an empty string. What should the release decision be?

Correct answer: hold, because rollback_pointer_recorded fails even though the forecast and alert metrics pass.

The release decision requires all gates to pass. Clearing previous_forecast makes rollback_pointer_recorded false, so the executable receipt must hold the candidate. Passing backtest and alert evidence does not replace the need for an explicit rollback target.

6.An actual parcel count arrives for an issued forecast. Which write preserves replayable evidence and makes missing or orphaned joins detectable?

Correct answer: Append an observation with the immutable forecast ID and observed time, then join it to the unchanged issued row.

The issued row is evidence of what was known and predicted before the target day, so later outcomes must not rewrite it. An append-only observation keyed by the immutable forecast ID preserves arrival time and allows replay checks to expose missing observations, orphaned observations, and incorrect joins.

7.The reviewed-alert function returns None when a rate denominator is zero. If every reviewed row has triggered = False and actionable = False, what happens to precision, recall, and the alert-evidence gates that require a non-None rate of at least 0.60?

Correct answer: Precision and recall are None, so both alert-evidence gates fail because there is no denominator-backed evidence.

With no triggered rows and no actionable rows, true positives, false positives, and false negatives are all zero. Both precision and recall would have zero denominators, so the helper returns None. The gates explicitly require a real rate before comparing with 0.60, so absence of evidence does not pass and does not crash.

8.A candidate receipt passes every listed gate: rolling-origin replay, MAE and peak-cost comparisons, local range coverage, queue size, reviewed alert precision and recall, and a nonempty previous forecast alias. What release action follows?

Correct answer: candidate_for_planner_shadow_review, because the receipt supports planner observation beside the current forecast, not production replacement across all slices and future weeks.

Passing the local receipt means the bundle is ready for planner shadow review beside the current forecast. It is not a production replacement yet: promotion still needs broader slice evidence, observed shadow behavior, monitored rollout, and a tested rollback path. New outcomes should be appended and reviewed on a scheduled or triggered cycle, not used to retrain and promote automatically on arrival.

9.A reviewed alert set has 2 true positives, 1 false positive, and 1 false negative. The false-negative event is relabeled not actionable while its triggered value remains false. What changes?

Correct answer: Recall rises from 2/3 to 1.0, while precision remains 2/3.

Relabeling the missed event as not actionable removes the only false negative, so recall becomes 2 / (2 + 0) = 1.0. The triggered rows do not change, leaving 2 true positives and 1 false positive, so precision remains 2 / (2 + 1) = 2/3.

10.Weekly slice reports show a sustained increase in high-volume underforecast cost. What does the operating policy require before a replacement is promoted?

Correct answer: Investigate the regression before fitting a replacement, then rerun rolling backtests, protected slices, shadow review, and the rollback check.

Daily outcomes should be appended and monitored, but sustained cost regression is a trigger for investigation rather than automatic replacement. Any proposed replacement must repeat the time-aware backtest, protected-slice evaluation, shadow review, and rollback checks before promotion.

10 questions remaining.

Next Step
Continue to Capstone: Image Damage Classifier

You can now package time-ordered forecasts, shadow evidence, and reviewed capacity alerts. Next you'll ship a model over pixels, where image quality and human confirmation guard every damage route.

PreviousCapstone: Product Ranking
Share this article
XFacebookLinkedInBlueskyRedditHacker NewsEmail
References

Forecasting: Principles and Practice, Third Edition.

Hyndman, R. J. & Athanasopoulos, G. · 2021

https://otexts.com/fpp3/

Discussion

Questions and insights from fellow learners.

Discussion loads when you reach this section.