LeetLLM
My PlanLearnGlossaryTracksPracticeBlog
LeetLLM

Your go-to resource for mastering AI & LLM systems.

Product

  • Learn
  • Glossary
  • Tracks
  • Practice
  • Blog
  • RSS

Legal

  • Terms of Service
  • Privacy Policy

© 2026 LeetLLM. All rights reserved.

All Topics
Your Progress
0%

0 of 177 articles completed

🛠️Computing Foundations0/9
Git, Shell, Linux for AIDocker for Reproducible AIPython for AI EngineeringNumPy and Tensor ShapesCUDA for ML TrainingMPS & Metal for ML on MacData Structures for AISQL and Data ModelingAlgorithms for ML Engineers
📊Math & Statistics0/8
Gradients and BackpropVectors, Matrices & TensorsLinear Algebra for MLAdam, Momentum, SchedulersProbability for Machine LearningStatistics and UncertaintyDistributions and SamplingHypothesis Tests, Intervals, and pass@k
📚Preparation & Prerequisites0/13
Neural Networks from ScratchCNNs from ScratchTraining & BackpropagationSoftmax, Cross-Entropy & OptimizationRNNs, LSTMs, GRUs, and Sequence ModelingAutoencoders and VAEsThe Transformer Architecture End-to-EndLanguage Modeling & Next TokensFrom GPT to Modern LLMsPrompt Engineering FundamentalsCalling LLM APIs in ProductionFirst AI App End-to-EndThe LLM Lifecycle
🧮ML Algorithms & Evaluation0/11
Linear Regression from ScratchLogistic Regression and MetricsDecision Trees, Forests, and BoostingReinforcement Learning BasicsValidation and LeakageClustering and PCACore Retrieval AlgorithmsDecoding AlgorithmsExperiment Design and A/B TestingPyTorch Training LoopsDataset Pipelines and Data Quality
📦Production ML Systems0/6
Feature Engineering for Production MLBatch and Streaming Feature PipelinesGradient Boosted Trees in ProductionRanking and Recommendation SystemsForecasting and Anomaly DetectionMonitoring Predictive Models
🧪Core LLM Foundations0/8
The Bitter Lesson & ComputeBPE, WordPiece, and SentencePieceStatic to Contextual EmbeddingsPerplexity & Model EvaluationFile Ingestion for AIChunking StrategiesLLM Benchmarks & LimitationsInstruction Tuning & Chat Templates
🧰Applied LLM Engineering0/24
Dimensionality Reduction for EmbeddingsCoT, ToT & Self-Consistency PromptingFunction Calling & Tool UseMCP & Tool Protocol StandardsContext EngineeringPrompt Injection DefenseResponsible AI GovernanceData Labeling and Human FeedbackEvaluating AI AgentsProduction RAG PipelinesHybrid Search: Dense + SparseReranking and Cross-Encoders for RAGRAG Evaluation for Reliable AnswersLLM-as-a-Judge EvaluationBias & Fairness in LLMsHallucination Detection & MitigationLLM Observability & MonitoringExperiment Tracking with MLflow and W&BPrompt Optimization with DSPyModel Versioning & DeploymentSemantic Caching & Cost OptimizationLLM Cost Engineering & Token EconomicsModel Gateways, Routing, and FallbacksDesign an Automated Support Agent
🎓Portfolio Capstones0/9
Capstone: Delivery ETA PredictionCapstone: Product RankingCapstone: Demand ForecastingCapstone: Image Damage ClassifierCapstone: Production ML PipelineCapstone: Document QACapstone: Eval DashboardCapstone: Fine-Tuned ClassifierCapstone: Reproducible ML Study
🧠Transformer Deep Dives0/8
Sentence Embeddings & Contrastive LossEmbedding Similarity & QuantizationScaled Dot-Product AttentionVision Transformers and Image EncodersPositional Encoding: RoPE & ALiBiLayer Normalization: Pre-LN vs Post-LNMechanistic InterpretabilityDecoding Strategies: Greedy to Nucleus
🧬Advanced Training & Adaptation0/16
Scaling Laws & Compute-Optimal TrainingPre-training Data at ScaleBuild GPT from Scratch LabJAX for PyTorch ResearchersContinued Pretraining for Domain ShiftSynthetic Data PipelinesSupervised Fine-Tuning PipelineMixed Precision TrainingDistributed Training: FSDP & ZeROLoRA & Parameter-Efficient TuningReward Modeling from Preference DataRLHF & DPO AlignmentConstitutional AI & Red TeamingRLVR & Verifiable RewardsKnowledge Distillation for LLMsModel Merging and Weight Interpolation
🤖Advanced Agents & Retrieval0/16
Vector DB Internals: HNSW & IVFAdvanced RAG: HyDE & Self-RAGGraphRAG & Knowledge GraphsRAG Security & Access ControlStructured Output GenerationReAct & Plan-and-ExecuteGuardrails & Safety FiltersCode Generation & SandboxingComputer-Use / GUI / Browser AgentsHuman-in-the-Loop Agent ArchitectureAI Coding Workflow with AgentsAgent Memory & PersistenceAgent Failure & RecoveryRecursive Language Models (RLM)Multi-Agent OrchestrationCapstone: Production Agent
⚡Inference & Production Scale0/19
Inference: TTFT, TPS & KV CacheMulti-Query & Grouped-Query AttentionKV Cache & PagedAttentionPrefix Caching and Prompt CachingFlashAttention & Memory EfficiencyContinuous Batching & SchedulingScaling LLM InferenceModel Parallelism for LLM InferenceModel Quantization: GPTQ, AWQ & GGUFLocal LLM DeploymentSLM Specialization & Edge DeploymentSpeculative DecodingLong Context Window ManagementMixture of Experts ArchitectureMamba & State Space ModelsReasoning & Test-Time ComputeAdvanced MLOps & DevOps for AIGPU Serving & AutoscalingA/B Testing for LLMs
🏗️System Design Capstones0/9
Content Moderation SystemCode Completion SystemMulti-Tenant LLM PlatformLLM-Powered Search EngineVision-Language Models & CLIPMultimodal LLM ArchitectureDiffusion Models: Images & TextReal-Time Voice AI AgentReasoning & Test-Time Compute
🎤AI Lab Interviewing0/4
AI Lab Coding Interview: Python SystemsAI Lab System Design InterviewAI Lab Behavioral InterviewAI Lab Technical Presentation
🔬Project Deep Dives0/17
Deep Dive - vLLMDeep Dive - SkyRLDeep Dive - FlashAttentionDeep Dive - FlashInferDeep Dive - DeepGEMMDeep Dive - NCCLDeep Dive - MegatronDeep Dive - DeepSpeedDeep Dive - RayDeep Dive - MLflowDeep Dive - PyTorchDeep Dive - TransformersDeep Dive - SGLangDeep Dive - slimeDeep Dive - DeepEPDeep Dive - TinkerDeep Dive - Light-PEFT
Back to Topics
LearnMath & StatisticsProbability for Machine Learning
📊EasyEvaluation & Benchmarks

Probability for Machine Learning

Use one API abuse-risk detector to learn events, priors, conditional probability, independence, Bayes rule, and base-rate mistakes.

15 min read
Learning path
Step 14 of 177 in the full curriculum
Adam, Momentum, SchedulersStatistics and Uncertainty

Personalize this lesson

Adapt explanations and teaching visuals to your background and preferred voice.

An abuse detector flags an API signup for manual review. It catches 95 percent of abusive signups, but that doesn't mean a flagged signup is 95 percent likely to be abusive. You also need the abuse rate in the traffic population and the detector's false-alarm rate.

Probability makes those quantities precise: name the event, population, and evidence, then count inside the right denominator.[1]Reference 1Machine Learning: A Probabilistic Perspective.https://probml.github.io/pml-book/book0.html[2]Reference 2Pattern Recognition and Machine Learning.https://www.microsoft.com/en-us/research/publication/pattern-recognition-machine-learning/[3]Reference 3Deep Learning.https://www.deeplearningbook.org/ The same discipline later applies to score calibration and token probabilities.

Bayes posterior calculation showing all signups narrowed to flagged signups, then 95 abuse cases divided by 590 flagged signups for 16 percent. Bayes posterior calculation showing all signups narrowed to flagged signups, then 95 abuse cases divided by 590 flagged signups for 16 percent.
Posterior uses abuse among flagged signups, not all signups.

Before the detector runs, you need a world to count in.

NameIn this exampleWhy it matters
sample space10,000 signupsthe full world you're counting over
eventsignup is abusivethe thing you care about
evidencedetector flagged the signupwhat you observed
priorabuse rate before the flagbelief before new evidence
posteriorabuse rate after the flagbelief after new evidence

If any of those pieces are missing, the number is floating. Floating numbers are how teams turn model scores into bad product decisions.

Start with the population

Use 10,000 recent signups from the same API signup surface.

Signup typeCountProbability
Abusive1000.01
Clean9,9000.99
Total10,0001.00

The probability of an abusive signup is a fraction:

P(abusive)=10010000=0.01P(\text{abusive}) = \frac{100}{10000} = 0.01P(abusive)=10000100​=0.01

Read that as: before the detector says anything, 1 percent of signups are abusive.

That starting probability is the prior. A prior isn't a guess pulled from nowhere. In an engineering system, it should usually come from a measured population: production logs, labeled eval data, reviewed tickets, or another concrete sample.

In probability language, a random variable is a number that depends on which case you happened to pick. If we let F = 1 when a signup is abusive and F = 0 when it's clean, the expectation E[F] is the long-run average value, each outcome weighted by its probability:

E[F]=1×0.01+0×0.99=0.01E[F] = 1 \times 0.01 + 0 \times 0.99 = 0.01E[F]=1×0.01+0×0.99=0.01

That 0.01 is the base rate written as an average instead of a fraction. A 0/1 variable like this is called a Bernoulli variable, and its expectation is always just the probability of the 1 outcome.

Expectation alone hides how much outcomes move around. Variance measures the average squared distance from the expectation, Var[F] = E[(F - E[F])^2]. For a Bernoulli variable with success probability p, it simplifies to p(1 - p):

Var[F]=0.01×(1−0.01)=0.0099\text{Var}[F] = 0.01 \times (1 - 0.01) = 0.0099Var[F]=0.01×(1−0.01)=0.0099

Variance names outcome-level spread around the average. Later statistics chapters separate that spread from uncertainty in an estimated rate and from variation across product slices or repeated training runs.

The count table can be written as a Bernoulli indicator. 1 means an abusive signup and 0 means a clean signup. The mean of that indicator is the prior.

indicator-mean-is-probability.py
1abuse_indicator = [1] * 100 + [0] * 9_900 2 3prior = sum(abuse_indicator) / len(abuse_indicator) 4variance = sum((value - prior) ** 2 for value in abuse_indicator) / len(abuse_indicator) 5 6print(f"signups: {len(abuse_indicator):,}") 7print(f"prior = E[F]: {prior:.4f}") 8print(f"Var[F]: {variance:.4f}") 9 10assert prior == 0.01 11assert abs(variance - prior * (1 - prior)) < 1e-12
Indicator summary
1signups: 10,000 2prior = E[F]: 0.0100 3Var[F]: 0.0099

Evidence changes the question

Now add the detector.

If the signup is...Detector behaviorProbability
Abusiveflags it0.95
Cleanfalsely flags it0.05

That top row is strong. If a signup is abusive, the detector catches it 95 percent of the time.

But product teams need a different decision after seeing a flag:

Given that this signup was flagged, how likely is it abusive?

That's a different question. Probability notation makes the difference visible:

NotationPlain English
P(flagged∣abusive)P(\text{flagged} \mid \text{abusive})P(flagged∣abusive)If a signup is abusive, how often does the detector flag it?
P(abusive∣flagged)P(\text{abusive} \mid \text{flagged})P(abusive∣flagged)If a signup is flagged, how often is it abusive?

Those two lines aren't interchangeable. Reversing them is one of the most common probability mistakes in ML systems.

Count the flagged pile

Work from the 10,000 signups. Don't start with Bayes rule yet.

Abusive signups:

100×0.95=95100 \times 0.95 = 95100×0.95=95

Clean signups:

9900×0.05=4959900 \times 0.05 = 4959900×0.05=495

Now put the flagged signups into one pile.

Source of flagged signupCount
abusive and flagged95
clean and flagged495
all flagged signups590

The false-positive rate is small, but it acts on the huge clean pile. Five percent of 9,900 is larger than 95 percent of 100.

That's why the posterior is surprising:

P(abusive∣flagged)=95590≈0.161P(\text{abusive} \mid \text{flagged}) = \frac{95}{590} \approx 0.161P(abusive∣flagged)=59095​≈0.161

A flagged signup is about 16 percent likely to be abusive, not 95 percent likely.

The detector didn't become bad. The question changed. We stopped asking how often abusive signups are caught and started asking what lives inside the flagged pile.

Turn the arithmetic into a confusion-table calculation. Notice that the program prints counts before it prints the posterior.

count-the-flagged-pile.py
1total_signups = 10_000 2abuse_signups = 100 3clean_signups = total_signups - abuse_signups 4true_positive_rate = 0.95 5false_positive_rate = 0.05 6 7true_flags = round(abuse_signups * true_positive_rate) 8false_flags = round(clean_signups * false_positive_rate) 9flagged_signups = true_flags + false_flags 10posterior = true_flags / flagged_signups 11 12print(f"true flags: {true_flags}") 13print(f"false flags: {false_flags}") 14print(f"all flags: {flagged_signups}") 15print(f"P(abuse | flagged): {posterior:.3f}") 16 17assert (true_flags, false_flags, flagged_signups) == (95, 495, 590)
Flagged-pile counts
1true flags: 95 2false flags: 495 3all flags: 590 4P(abuse | flagged): 0.161

Conditioning means narrowing the world

Conditional probability always narrows the world first.

For P(abusive∣flagged)P(\text{abusive} \mid \text{flagged})P(abusive∣flagged), the denominator isn't all 10,000 signups. The denominator is only the 590 flagged signups.

ProbabilityWorld you count insideNumerator
P(abusive)P(\text{abusive})P(abusive)all 10,000 signups100 abusive signups
P(abusive∣flagged)P(\text{abusive} \mid \text{flagged})P(abusive∣flagged)590 flagged signups95 abusive flagged signups

That's the whole mental move. Ask "among which cases?" before you divide.

This is the same flow as the diagram:

Conditional probability diagram comparing the 1% base-rate abuse prior over 10,000 signups with the 16% posterior inside the 590 flagged signups, computed as 95 divided by 590. Conditional probability diagram comparing the 1% base-rate abuse prior over 10,000 signups with the 16% posterior inside the 590 flagged signups, computed as 95 divided by 590.
Conditioning changes the denominator. Once you condition on a flag, the relevant world is the 590 flagged signups.

When a probability problem feels abstract, draw that flow. The formula should summarize the drawing, not replace it.

Joint and marginal probability

Two more words tie the counts together.

A joint probability is the chance that two things happen together. In the population, 95 signups are both abusive and flagged, so:

P(abusive and flagged)=9510000=0.0095P(\text{abusive and flagged}) = \frac{95}{10000} = 0.0095P(abusive and flagged)=1000095​=0.0095

A marginal probability is the chance of one event by itself, ignoring the other. The marginal probability of a flag adds up every way a flag can happen:

P(flagged)=95+49510000=59010000=0.059P(\text{flagged}) = \frac{95 + 495}{10000} = \frac{590}{10000} = 0.059P(flagged)=1000095+495​=10000590​=0.059

Joint, conditional, and marginal probabilities are linked by one identity. When P(B)>0P(B) > 0P(B)>0, conditional probability is the joint divided by the world you conditioned on:

P(A∣B)=P(A and B)P(B)P(A \mid B) = \frac{P(A \text{ and } B)}{P(B)}P(A∣B)=P(B)P(A and B)​

Rearranged, that gives the multiplication rule: a joint probability is a conditional times a marginal.

P(A and B)=P(A∣B) P(B)=P(B∣A) P(A)P(A \text{ and } B) = P(A \mid B)\,P(B) = P(B \mid A)\,P(A)P(A and B)=P(A∣B)P(B)=P(B∣A)P(A)

Check it against the counts: P(abusive∣flagged)=0.0095/0.059≈0.161P(\text{abusive} \mid \text{flagged}) = 0.0095 / 0.059 \approx 0.161P(abusive∣flagged)=0.0095/0.059≈0.161, the same posterior as before. This identity is also where Bayes rule comes from. The next section just reads it from the other direction.

Bayes rule after the counts

Bayes rule is the formula version of the flagged-pile count.

For events AAA and BBB with P(B)>0P(B) > 0P(B)>0:

P(A∣B)=P(B∣A)P(A)P(B)P(A \mid B) = \frac{P(B \mid A)P(A)}{P(B)}P(A∣B)=P(B)P(B∣A)P(A)​

For this example:

SymbolMeaningValue
AAAsignup is abusive
BBBdetector flagged the signup
P(A)P(A)P(A)abuse base rate0.01
P(B∣A)P(B \mid A)P(B∣A)true-positive rate0.95
P(B∣not A)P(B \mid \text{not }A)P(B∣not A)false-positive rate0.05

When BBB is the observed evidence, P(B∣A)P(B \mid A)P(B∣A) is its likelihood under hypothesis AAA. Here it asks how likely a flag would be if the signup really were abusive. Bayes rule combines that likelihood with the prior to obtain the posterior.

The denominator P(B)P(B)P(B) means "how often does a flag happen at all?"

It includes true flags and false alarms:

P(B)=P(B∣A)P(A)+P(B∣not A)P(not A)P(B) = P(B \mid A)P(A) + P(B \mid \text{not }A)P(\text{not }A)P(B)=P(B∣A)P(A)+P(B∣not A)P(not A)

The denominator P(B)P(B)P(B) is the marginal probability of a flag. It averages over both kinds of signups, weighted by how common each kind is.

Plug in the numbers:

P(B)=0.95×0.01+0.05×0.99=0.059P(B) = 0.95 \times 0.01 + 0.05 \times 0.99 = 0.059P(B)=0.95×0.01+0.05×0.99=0.059

Now compute the posterior:

P(A∣B)=0.95×0.010.059≈0.161P(A \mid B) = \frac{0.95 \times 0.01}{0.059} \approx 0.161P(A∣B)=0.0590.95×0.01​≈0.161

Same answer as the count table. Bayes rule is the compact form of the same accounting: track where the flagged signups came from before you divide.

Independence means no update

Evidence helps when it changes the event rate.

When P(B)>0P(B) > 0P(B)>0, seeing BBB leaves the probability of AAA unchanged if the events are independent:

P(A∣B)=P(A)P(A \mid B) = P(A)P(A∣B)=P(A)

Start with a broken abuse detector:

If the signup is...Broken detector flags it
Abusive20 percent
Clean20 percent

Out of 10,000 signups, this detector produces:

Source of flagged signupCount
abusive and flagged20
clean and flagged1,980
all flagged signups2,000

The flagged pile is still 1 percent abusive:

202000=0.01\frac{20}{2000} = 0.01200020​=0.01

The flag created work, but it didn't create information. Useful ML signals are useful because they change the rate of the event you care about.

Code makes the "no update" claim testable. If both groups are flagged at the same rate, that rate cancels from Bayes rule.

independence-means-no-update.py
1def posterior_if_flagged(prior, true_positive_rate, false_positive_rate): 2 true_flags = true_positive_rate * prior 3 false_flags = false_positive_rate * (1 - prior) 4 return true_flags / (true_flags + false_flags) 5 6prior = 0.01 7posterior = posterior_if_flagged(prior, 0.20, 0.20) 8 9print(f"prior: {prior:.3f}") 10print(f"posterior: {posterior:.3f}") 11print(f"update: {posterior - prior:+.3f}") 12 13assert abs(posterior - prior) < 1e-12
Independent evidence
1prior: 0.010 2posterior: 0.010 3update: +0.000

Same detector, different world

Keep the detector fixed:

Detector propertyValue
true-positive rate0.95
false-positive rate0.05

Now change only the population.

Abuse base ratePosterior after flagWhat changed?
1 percentabout 16 percentclean signups dominate the flagged pile
10 percentabout 68 percenttrue flags become a much larger share
50 percentabout 95 percentboth classes are equally common before evidence

The model didn't change. The world around the model changed.

The same classifier can behave differently across products, countries, languages, traffic sources, or time periods. A score without a population is like a map without a scale. The score may look precise, but it isn't enough to act carefully.

Build it: compute posterior risk

The code should read like the table:

  1. Check that each input is a valid probability.
  2. Count true flags as true_positive * prior.
  3. Count false flags as false_positive * (1 - prior).
  4. Divide true flags by all flags.

Put this in probability_demo.py. It's small enough to audit line by line, but already exposes the input validation and evidence guard that production code needs.

probability_demo.py
1def check_probability(x, name): 2 if not 0 <= x <= 1: 3 raise ValueError(f"{name} must be between 0 and 1") 4 5def flagged_posterior(prior, true_positive, false_positive): 6 check_probability(prior, "prior") 7 check_probability(true_positive, "true_positive") 8 check_probability(false_positive, "false_positive") 9 10 true_flags = true_positive * prior 11 false_flags = false_positive * (1 - prior) 12 all_flags = true_flags + false_flags 13 14 if all_flags == 0: 15 raise ValueError("evidence probability must be greater than 0") 16 17 return true_flags / all_flags 18 19def main(): 20 priors = [0.01, 0.10, 0.50] 21 22 for prior in priors: 23 posterior = flagged_posterior(prior, 0.95, 0.05) 24 print(prior, round(posterior, 3)) 25 26 try: 27 flagged_posterior(1.4, 0.95, 0.05) 28 except ValueError as error: 29 print(error) 30 31if __name__ == "__main__": 32 main()
Posterior sweep
10.01 0.161 20.1 0.679 30.5 0.95 4prior must be between 0 and 1

The detector stayed fixed. The prior changed, so the posterior changed.

A threshold changes two probabilities at once

A deployed detector rarely emits only flag or not flag. It emits a score, and a threshold creates the flag. Raising a threshold normally sends fewer signups to reviewers and may make the flagged pile cleaner, but it can also miss more abusive signups.

Use an illustrative measurement table from the same 1 percent abuse population. These rates would need to be measured on labeled data in a real system.

Threshold policyP(flagged∣abuse)P(\text{flagged} \mid \text{abuse})P(flagged∣abuse)P(flagged∣clean)P(\text{flagged} \mid \text{clean})P(flagged∣clean)
broad review0.950.05
stricter review0.800.01

The following code models this trade-off by simulating the review queues for both policies:

threshold-tradeoff.py
1def review_metrics(prior, recall, false_positive_rate): 2 true_flags = recall * prior 3 false_flags = false_positive_rate * (1 - prior) 4 review_rate = true_flags + false_flags 5 if review_rate == 0: 6 raise ValueError("review rate must be greater than 0") 7 precision = true_flags / review_rate 8 return review_rate, precision 9 10prior = 0.01 11policies = [ 12 ("broad review", 0.95, 0.05), 13 ("stricter review", 0.80, 0.01), 14] 15 16for name, recall, false_positive_rate in policies: 17 review_rate, precision = review_metrics(prior, recall, false_positive_rate) 18 print( 19 f"{name:16} review={review_rate:6.2%} " 20 f"abuse_in_queue={precision:6.2%} recall={recall:6.2%}" 21 ) 22 23try: 24 review_metrics(prior, recall=0.0, false_positive_rate=0.0) 25except ValueError as error: 26 print("empty queue:", error)
Threshold tradeoff
1broad review review= 5.90% abuse_in_queue=16.10% recall=95.00% 2stricter review review= 1.79% abuse_in_queue=44.69% recall=80.00% 3empty queue: review rate must be greater than 0

The stricter policy reduces review load and improves the fraction of reviewed signups that are abuse, which is precision, but it catches fewer abuse cases, which lowers recall. This is a precision-recall tradeoff. Probability exposes the tradeoff; product cost and safety policy choose among the options. A threshold can also send no signups to review. In that case, queue precision is undefined because no flagged pile exists, so code should reject or explicitly represent the empty queue instead of dividing by zero.

Probability in ML work

The abuse example is one surface.

ML systemEventEvidenceBase-rate question
abuse detectorsignup is abusiverisk score above thresholdhow common is abuse for this traffic slice?
code searchretrieved chunk is relevantembedding similarity above thresholdhow many returned chunks answer the developer query?
label judgepredicted label is correctautomated judge marks it correcthow often does the judge agree with human reviewers?
training-job anomaly detectorrun breaches SLAanomaly score is highhow common are breaches for this job family or region?

The same habit works everywhere:

  1. Name the event.
  2. Name the population.
  3. Measure the base rate.
  4. Name the evidence.
  5. Ask how the evidence changes the event rate.

Skip those steps and a score starts pretending to be a conclusion.

A score isn't automatically a probability

So far, the evidence was a thresholded flag and all rates were given. A model may instead emit a score such as 0.80. That score earns the interpretation "80 percent probability of abuse" only if similarly scored signups are abusive about 80 percent of the time. That property is calibration. Modern neural classifiers can be accurate while still producing poorly calibrated confidence scores, which is why score calibration is measured rather than assumed.[4]Reference 4On Calibration of Modern Neural Networkshttps://arxiv.org/abs/1706.04599

This toy bucket demonstrates the check. Every signup was assigned a score near 0.80, but only half of these labeled signups were abusive.

check-one-score-bucket.py
1predicted_risk = [0.80, 0.82, 0.78, 0.81, 0.79, 0.80] 2observed_abuse = [1, 0, 1, 0, 1, 0] 3 4advertised_risk = sum(predicted_risk) / len(predicted_risk) 5observed_rate = sum(observed_abuse) / len(observed_abuse) 6gap = advertised_risk - observed_rate 7 8print(f"average predicted risk: {advertised_risk:.0%}") 9print(f"observed abuse rate: {observed_rate:.0%}") 10print(f"calibration gap: {gap:.0%}")
Calibration bucket
1average predicted risk: 80% 2observed abuse rate: 50% 3calibration gap: 30%

Six signups can't prove how a deployed model is calibrated. The calculation names the question. Statistics will teach how much evidence you need before trusting the measured gap.

A token model multiplies probabilities

We used one binary event, but an LLM produces a categorical distribution over the next token. For a known target token, training cares about the probability assigned to that target. With a one-hot target, that token's cross-entropy contribution is its negative log probability, -log(p).[3]Reference 3Deep Learning.https://www.deeplearningbook.org/

This tiny snippet prints the loss for varying target probabilities:

negative-log-probability.py
1import math 2 3for target_probability in [0.90, 0.50, 0.01]: 4 loss = -math.log(target_probability) 5 print(f"target probability={target_probability:>4.2f} loss={loss:>5.3f}")
Token surprise
1target probability=0.90 loss=0.105 2target probability=0.50 loss=0.693 3target probability=0.01 loss=4.605

Low probability for the observed token costs more because the observation was more surprising under the model.

For a sequence, the model's joint probability follows the chain rule: multiply the conditional probability of each next token given the previous tokens.

P(t1,t2,…,tn)=P(t1)∏i=2nP(ti∣t1,…,ti−1)P(t_1, t_2, \ldots, t_n) = P(t_1) \prod_{i=2}^{n} P(t_i \mid t_1, \ldots, t_{i-1})P(t1​,t2​,…,tn​)=P(t1​)i=2∏n​P(ti​∣t1​,…,ti−1​)

Here, tit_iti​ is the token at position iii, and the expression to the right of the bar is the preceding context. Multiplying many small probabilities eventually underflows in floating-point arithmetic. Logs convert the product into a stable sum.

We can verify this underflow by multiplying 200 token probabilities in Python:

keep-sequence-probabilities-in-log-space.py
1import math 2 3token_probability = 0.01 4token_count = 200 5 6raw_product = token_probability ** token_count 7log_probability = token_count * math.log(token_probability) 8 9print(f"raw product in float: {raw_product}") 10print(f"log probability: {log_probability:.1f}") 11print(f"finite in log space: {math.isfinite(log_probability)}")
Log-space sequence probability
1raw product in float: 0.0 2log probability: -921.0 3finite in log space: True

The 0.0 doesn't mean the sequence was impossible. It means ordinary floating-point multiplication lost a representable nonzero number. Later language-modeling chapters use this same log-space habit for cross-entropy and perplexity.

Common mistakes

Most probability bugs are question bugs.

SymptomMistakeBetter move
"The flag means 95 percent abusive."reversed the conditional probabilitieswrite both questions in plain English
Posterior feels too highignored rare base ratestart from counts before formulas
Denominator is all signupsforgot conditioningdenominator should be the evidence pile
One threshold used everywhereignored population shiftrecompute base rates per product slice
Model confidence treated as truthevent never nameddefine the event and compare to labels
Code returns a number for impossible evidencedivided by zero-probability evidencereject undefined cases loudly
A score of 0.80 is used as 80 percent risk without checking labelscalibration was assumedbucket predictions and compare score with observed rate
A long sequence gets probability 0.0 in codesmall probabilities were multiplied directlysum log probabilities instead

The debugging question is short:

Among which cases am I counting?

If you can answer that, the formula usually becomes much easier.

Try it yourself

Use the same 10,000-signup population.

  1. Recompute P(abusive)P(\text{abusive})P(abusive) from the first table.
  2. Recompute the 495 false flags by hand.
  3. Change the false-positive rate from 0.05 to 0.01.
  4. Compute the new posterior.
  5. Explain why the posterior changed even though the true-positive rate stayed 0.95.
  6. Compare the broad and stricter review policies: which catches more abuse, and which sends a cleaner queue to reviewers?
  7. Compute -log(0.80) and -log(0.10). Which observed token is more surprising to a model?

Then translate the lesson to product search:

Product search versionYour answer
eventproduct is truly relevant
populationcandidate products returned for a query class
evidencesimilarity score above threshold
priorrelevance rate before thresholding
posteriorrelevance rate after thresholding

Don't memorize the abuse numbers. Carry the counting habit into any model score.

Solution checks

Check your work after you try the practice.

Practice itemAnswer
Prior100/10000=0.01100 / 10000 = 0.01100/10000=0.01
Original false flags9900×0.05=4959900 \times 0.05 = 4959900×0.05=495
New false flags9900×0.01=999900 \times 0.01 = 999900×0.01=99
New posterior95/(95+99)≈0.4995 / (95 + 99) \approx 0.4995/(95+99)≈0.49
Why it changedthe false-alarm pile got smaller, so true flags became a larger share of all flags
Threshold tradeoffbroad review catches more abuse; stricter review produces a cleaner, smaller queue
Token loss comparison−log⁡(0.80)≈0.223-\log(0.80) \approx 0.223−log(0.80)≈0.223 and −log⁡(0.10)≈2.303-\log(0.10) \approx 2.303−log(0.10)≈2.303; the 0.10 target is more surprising

If your explanation starts with Bayes rule, translate it back into piles. A good answer can move between counts, words, code, and notation.

A complete probability claim

Probability turns model scores into named claims.

For an API platform, a useful probability statement sounds like this:

Event: API-key request is abusive. Population: new workspace key requests in the last 30 days. Evidence: abuse detector score above threshold. Decision: send to manual review when posterior risk is above 20 percent.

That sentence is longer than "score is high," but it's much safer.

Complete the lesson

Mastery Check

Answer every question, then check your score. Score 75% or higher to mark this lesson complete.

1.In 10,000 signups, 100 are abusive. A detector flags 95% of abusive signups and 5% of clean signups. What is P(abusive | flagged)?
2.A detector flags 20% of abusive signups and 20% of clean signups. If the abuse prior is 1%, what is P(abusive | flagged)?
3.At a 1% abuse prior, a broad threshold has 95% recall and a 5% false-positive rate. A strict threshold has 80% recall and a 1% false-positive rate. What changes?
4.A model multiplies 200 token probabilities of 0.01, and the floating-point product is 0.0 while the summed log probability is finite. What should the evaluator do?
5.Six signups have scores averaging 80%, but only 3 of the 6 are abusive. What does this bucket check show?
6.A retrieval system defines the event as product relevance and the evidence as a similarity score above threshold. Which counting worlds define its prior and posterior?
7.In retrieval, 20% of candidates are relevant. A threshold selects 80% of relevant candidates and 10% of irrelevant candidates. Which Bayes calculation is correct?
8.A detector keeps 95% recall and a 5% false-positive rate, but the abuse prior rises from 1% to 10%. Why does P(abusive | flagged) rise?
9.A flagged-posterior function is called with (prior=1.4, true_positive=0.95, false_positive=0.05) and with (prior=0.01, true_positive=0, false_positive=0). What should its contract require?
10.An observed target token has probability 0.10 in one prediction and 0.80 in another. Which prediction has the larger negative log probability?

10 questions remaining.

Next Step
Continue to Statistics and Uncertainty

Probability taught you how to reason inside a known population. Statistics asks how much you should trust probabilities estimated from finite data.

PreviousAdam, Momentum, Schedulers
Share this article
XFacebookLinkedInBlueskyRedditHacker NewsEmail
References

Machine Learning: A Probabilistic Perspective.

Murphy, K. P. · 2012

Pattern Recognition and Machine Learning.

Bishop, C. M. · 2006

Deep Learning.

Goodfellow, I., Bengio, Y., Courville, A. · 2016

On Calibration of Modern Neural Networks

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. · 2017

Discussion

Questions and insights from fellow learners.

Discussion loads when you reach this section.