Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
Access-change request REQ-10234 is sitting in the queue. May the workflow auto-process a permission change under policy line P-7, or must a human review it first?
The engine has to choose before the case is closed, so the true outcome isn't available yet. It needs a score that orders risk and a rule that turns that score into an action.
The previous chapter fitted latency from retrieved-context length with a design matrix, a loss, and a gradient-descent loop. We reuse that machinery, but now predict whether a request requires review. The score estimates a label; a separate policy chooses an action.
Logistic regression maps two evidence features to a score in (0, 1), then asks you to justify treating that score as a probability.[1]
Fit eight fictional requests, derive the gradient, implement it in NumPy, and check against scikit-learn.[2][3] A separate eight-request validation slice exposes three different questions: what does a threshold cost, how well do scores rank cases, and do they describe frequencies? These constructed rows teach arithmetic; they aren't evidence for releasing a permission workflow.
From a latency number to a review decision
The previous target was continuous: "how many milliseconds will this evidence cost?" A residual of 2 ms still meant something. Now the label is binary: 1 means an access-change request needs human review before any permission change; 0 means policy evidence supports automatic handling.
For this fictional cost model, an incorrect automatic approval costs $120 in permission rollback and investigation. An unnecessary review costs $18 of queue time. Correct decisions have zero incremental cost. Those assumptions will choose the threshold later, not the loss used to fit the model. A real permission workflow must still enforce its authorization rules independently of this classifier.
Ordinary least squares can fit 0/1 labels as a linear probability model, but it has limitations:
- Unbounded predictions: can fall outside [0, 1]. Values such as
-0.2aren't valid probabilities. - Sensitivity to influential rows: unusual feature values and residuals can change the fitted boundary. The direction depends on the data; an extreme positive doesn't universally move a threshold rightward.
- Conditional variance: for Bernoulli outcomes, , which can vary with features. Constant variance isn't required to compute OLS or for unbiasedness under a correct conditional mean, but it is part of the classical BLUE guarantee and conventional standard-error formulas.
Squared error on bounded probabilities is also a legitimate scoring loss, the Brier score. Logistic regression gives us a Bernoulli likelihood with a linear logit; it isn't the only way to estimate probabilities.
We need a function that squeezes any real score smoothly into (0, 1). The same separation appears across AI systems: a prompt-injection guardrail, toxicity filter, or agent action router emits a score, and a threshold turns that score into an intervention.
Squeeze the score into (0, 1)
Before choosing a threshold, we need a score that behaves like a probability: it must stay between 0 and 1, move upward when evidence favors review, and change smoothly enough for gradient descent. The sigmoid (or logistic) function does that:
Four properties matter for the later gradient:
- As ,
- As ,
- At ,
- Its derivative is , which is positive for finite and peaks at when . That tidy form is why the gradient of log-loss later collapses so cleanly.
Predict the shape before calculating. A large negative score should look almost like zero, z = 0 should land at one half, and a large positive score should approach one. Then compute five concrete values by hand:
| z | e^{-z} | 1 + e^{-z} | σ(z) |
|---|---|---|---|
| -4.0 | 54.6 | 55.6 | 0.018 |
| -1.0 | 2.718 | 3.718 | 0.269 |
| 0.0 | 1.0 | 2.0 | 0.500 |
| 1.0 | 0.368 | 1.368 | 0.731 |
| 4.0 | 0.018 | 1.018 | 0.982 |
1import numpy as np
2
3def sigmoid(z: np.ndarray) -> np.ndarray:
4 return 1.0 / (1.0 + np.exp(-z))
5
6scores = np.array([-4.0, -1.0, 0.0, 1.0, 4.0])
7print("scores:", scores)
8print("probabilities:", sigmoid(scores).round(3))1scores: [-4. -1. 0. 1. 4.]
2probabilities: [0.018 0.269 0.5 0.731 0.982]The code confirms the hand calculation. The curve is steep near zero and flat in both tails, which will matter when an update tries to move an already confident request.

The curve saturates quickly. A score of +4 produces p = 0.982; whether that number represents an honest frequency still needs validation on labels the fit did not see.
For finite real logits, sigmoid is strictly between 0 and 1 mathematically. Floating-point evaluation can round extreme results to 0 or 1. The first cell uses small logits for inspection; the trainer below uses a stable implementation and computes loss directly from logits.
Compute by hand. What decision does it suggest for an access-change request if the human-review threshold is 0.5?
Answer
. At threshold 0.5, route the request to human review. Calling this an 88.1% event frequency requires a calibration audit on held-out labels.
The model remains linear inside the sigmoid:
The probability is . The positive class is needs human review, so the sign and size of carry the evidence before the sigmoid bounds it.
The log-odds and logit link
The connection comes from odds, the ratio of the positive-label probability to the negative-label probability. An odds ratio would compare two odds, for example at two feature settings. Here the events are review required versus review not required, not the action the threshold will take:
When , the odds are (4 to 1 in favor of review). Taking the natural logarithm produces the logit, or log-odds:[1]
Modeling the log-odds as a linear function is the logit link function. Inverting that link gives the sigmoid activation directly:
At , the odds are 1:1 and . Adding one unit to multiplies the odds by .
The linear decision boundary hyperplane
Use and for the feature coefficients, keeping the intercept separate. At , review is selected when . The equal-score boundary is
For , we can write this line as:
For nonzero , it is the normal vector pointing toward . The signed orthogonal distance in these feature coordinates is:
Points sitting far from the hyperplane carry large , driving into the saturated tails near 0 or 1. Changing the decision threshold to any shifts the decision boundary to the parallel hyperplane:
Lowering the threshold to requires only , shifting the decision plane into negative logit space to route more borderline requests to human review.
If but , the line is vertical. If , the classifier is constant and has no separating hyperplane. Distance depends on feature units, so don't interpret it as a physical risk measure.
Eight historical requests
Each historical request has two scaled features produced from retrieved evidence. Read them as competing signals, not as finished decisions:
x1= ambiguity score: conflicting audit logs, missing requester context, or unclear scope descriptionx2= auto-policy support score: strength of evidence that policy lineP-7permits automatic handling
The label y = 1 when the completed case required human review before a permission change, and y = 0 when automatic handling was acceptable. The model must learn how those two signals combine.
These eight rows are constructed training data; the validation rows introduced later are separate:
| Request | x1 ambiguity | x2 auto-policy support | y needs review |
|---|---|---|---|
| R1 | 1.0 | 1 | 0 |
| R2 | 1.5 | 2 | 0 |
| R3 | 2.0 | 3 | 1 |
| R4 | 3.5 | 6 | 1 |
| R5 | 4.0 | 7 | 1 |
| R6 | 4.5 | 5 | 1 |
| R7 | 2.8 | 4 | 0 |
| R8 | 3.2 | 8 | 0 |
The training fixture stays small so you can verify each number. Before seeing any weights, predict their signs: ambiguity should raise review risk, while policy support should lower it. R3 needed review despite only moderate ambiguity. R8 had substantial ambiguity, but strong policy support made automatic processing acceptable.
Add the intercept column of ones and build the design matrix X (8 × 3) and the label vector y.
Log-loss (binary cross-entropy) is the right loss for probabilities
If a request didn't need review (y = 0), a score of p = 0.93 gives the observed outcome only 1 - 0.93 = 0.07 probability. Its negative log is -log(0.07) ≈ 2.659. A score of p = 0.07 instead assigns that outcome probability 0.93, for a much smaller loss, -log(0.93) ≈ 0.073. A confident wrong prediction pays heavily.
Deriving binary cross-entropy from maximum likelihood
Each historical request has a binary label . We model given features as a Bernoulli random trial with success probability . The Bernoulli probability mass function for a single example is:
When , leaving . When , leaving . Assuming conditional independence across the requests in our training set, the likelihood of observing the entire dataset is the joint product:
Multiplying hundreds of probabilities in swiftly triggers floating-point underflow. Taking the natural logarithm converts the product into an additive sum of log-probabilities:
Maximizing this likelihood is equivalent to minimizing the Negative Log-Likelihood (NLL). Normalizing by gives the empirical Binary Cross-Entropy (BCE) loss:[4][5]
For one hard label, define its target distribution and predicted distribution . Then and because this per-row target is a point mass. That doesn't make the population label distribution deterministic: our balanced training labels have nonzero marginal entropy. A sampled label isn't the full true conditional distribution. Expected log loss is minimized by that true conditional probability when the model can represent it.
Why not reuse Mean Squared Error?
We can fit with squared error, but composing it with sigmoid changes the optimization:
- Nonconvexity in weights: squared error after sigmoid is generally nonconvex. BCE with a linear logit is convex. Its Hessian is , positive definite at finite weights when has full column rank; redundant columns instead give flat directions.
- Small gradients on confident mistakes: for squared error , the logit derivative is . For half-squared error it is . At , the former is about , while BCE's derivative is about . BCE preserves a strong learning signal on this confidently wrong prediction.
Strict convexity gives at most one finite minimizer, not a guarantee that it exists. With complete or quasi-complete class separation, an unregularized optimum can require coefficients diverging to infinity. On perfectly separable data, loss can approach zero without attaining it at finite weights. Regularization can constrain that growth. Neither convexity nor low training loss establishes future accuracy.
Work one row by hand. For R4 (x1=3.5, x2=6, y=1) with a starting guess w = [0, 0, 0]:
z = 0 + 0·3.5 + 0·6 = 0p = σ(0) = 0.5L = -[1 * log(0.5) + 0] = -log(0.5) ≈ 0.693
Now suppose the model improves in the correct direction and predicts for the same row:
L = -log(0.82) ≈ 0.198- The loss dropped because the model became more confident in the correct direction, which is the behavior you want.
For code, compute the same loss from logits rather than taking log(sigmoid(z)). The expression log(1 + exp(z)) - y*z is the same loss, but evaluating exp(z) directly can overflow. NumPy's logaddexp(0, z) - y*z evaluates that softplus term stably even when z is very large in either direction.
The next check keeps the arithmetic in logit space. Predict which row should be cheap, which should be expensive, and whether extreme logits remain finite.
1import numpy as np
2
3def log_loss_from_logits(y: np.ndarray, z: np.ndarray) -> float:
4 return float(np.mean(np.logaddexp(0.0, z) - y * z))
5
6print("correct positive:", round(log_loss_from_logits(np.array([1.0]), np.array([2.0])), 3))
7print("confident wrong:", round(log_loss_from_logits(np.array([1.0]), np.array([-4.0])), 3))
8extreme = log_loss_from_logits(np.array([0.0, 1.0]), np.array([1000.0, -1000.0]))
9print("extreme logits finite:", np.isfinite(extreme))1correct positive: 0.127
2confident wrong: 4.018
3extreme logits finite: TrueThe log-loss gradient collapses to (p - y) · x
The loss tells us how wrong a prediction is. The gradient tells each weight how to change the score. For R4, the starting probability error is 0.5 - 1 = -0.5. Multiplying by [1, 3.5, 6] gives [-0.5, -1.75, -3.0]. Subtracting this gradient raises the next score for that row. Why does this simple multiplication give the derivative?
Chain rule step by step
We have with and linear score :
- Loss with respect to probability:
- Probability with respect to score:
- Score with respect to weight:
Multiplying the first two terms produces a cancellation: the sigmoid derivative in the numerator eliminates the denominator from the cross-entropy derivative:
Compare the two label cases directly:
| Label | Loss | Derivative with respect to score | Interpretation |
|---|---|---|---|
| Prediction is too low unless is near 1; gradient pushes upward | |||
| Prediction is too high unless is near 0; gradient pushes downward |
Both cases yield . Combining this with gives the single-sample weight gradient:
Across the entire dataset of samples, stacking rows into design matrix , predicted probabilities , and label vector gives the batch gradient:[4][5]
The bridge to linear regression
Notice the mathematical unity between this classifier and the linear regression model:
- Linear regression (MSE): , where
- Logistic regression (BCE): , where
Both use features transposed against prediction errors. The factor of 2 differs because the previous lesson used MSE; using half-MSE would make the prefactors match. Loss normalization changes gradient magnitudes and therefore the learning rate needed for the same update.
The logit is the Bernoulli canonical link in generalized linear models. The score-error term is simple, but weight updates still depend on feature values, averaging, step size, and any penalties or sample weights.
Plug in numbers for the starting point on R4 (p=0.5, y=1, x=[1, 3.5, 6]):
The gradient points uphill, toward increasing loss. Gradient descent subtracts it, scaled by the learning rate, to move downhill.
The sign is the learning signal. If and , is negative, so subtracting the gradient increases weights attached to positive features and raises the future score for similar requests. If and , is positive, so subtracting the gradient lowers those weights and reduces the score.
Repeat that arithmetic for each row and average the gradients. You now have one logistic-regression update. The code keeps the linear-regression loop and swaps in the sigmoid prediction and (p - y) * x gradient.
1import numpy as np
2
3def sigmoid(z: float) -> float:
4 return float(1.0 / (1.0 + np.exp(-z)))
5
6def loss(w: np.ndarray, x: np.ndarray, y: float) -> float:
7 z = x @ w
8 return float(np.logaddexp(0.0, z) - y * z)
9
10x = np.array([1.0, 3.5, 6.0])
11y = 1.0
12w = np.zeros(3)
13analytic = (sigmoid(x @ w) - y) * x
14eps = 1e-6
15numeric = np.array([
16 (loss(w + eps * np.eye(3)[j], x, y) - loss(w - eps * np.eye(3)[j], x, y)) / (2 * eps)
17 for j in range(3)
18])
19print("analytic:", analytic.round(3))
20print("numeric :", numeric.round(3))
21print("match:", np.allclose(analytic, numeric, atol=1e-6))1analytic: [-0.5 -1.75 -3. ]
2numeric : [-0.5 -1.75 -3. ]
3match: True
An access-change request has but the model predicts . Is positive or negative, and what does gradient descent do to the score for similar requests?
Answer
, so the gradient is positive in the feature direction. Gradient descent subtracts that gradient, lowering the score for similar requests next time.
Fit the eight rows in NumPy
Start with the full eight-row fit. The loop is the gradient-descent trainer from the latency chapter, with a sigmoid prediction and a (p - y) * x gradient. Before running it, predict the first update: the positive feature weights should rise from zero, while the intercept should stay near zero because the batch starts with equal positive and negative pressure.
1import numpy as np
2
3def sigmoid(z: np.ndarray) -> np.ndarray:
4 # exp(-log(1 + exp(-z))) avoids overflow without clipping the score.
5 return np.exp(-np.logaddexp(0.0, -z))
6
7def log_loss_from_logits(y_true: np.ndarray, z: np.ndarray) -> float:
8 return float(np.mean(np.logaddexp(0.0, z) - y_true * z))
9
10def fit_logistic(X: np.ndarray, y: np.ndarray, lr: float = 0.2, epochs: int = 3000, verbose: bool = True) -> np.ndarray:
11 n, d = X.shape
12 w = np.zeros(d)
13 for epoch in range(epochs):
14 z = X @ w
15 p = sigmoid(z)
16 grad = (1 / n) * (X.T @ (p - y))
17 w -= lr * grad
18 if verbose and epoch in (0, 200, 500, 1000, 2999):
19 loss = log_loss_from_logits(y, X @ w)
20 print(f"after_epoch={epoch + 1:4d} loss={loss:.4f} w={np.round(w, 3)}")
21 return w
22
23def predict_proba(X: np.ndarray, w: np.ndarray) -> np.ndarray:
24 return sigmoid(X @ w)
25
26# Historical access-change requests: [ambiguity, auto-policy support]
27X_raw = np.array([
28 [1.0, 1.0], [1.5, 2.0], [2.0, 3.0], [3.5, 6.0],
29 [4.0, 7.0], [4.5, 5.0], [2.8, 4.0], [3.2, 8.0],
30])
31y = np.array([0, 0, 1, 1, 1, 1, 0, 0])
32
33X = np.c_[np.ones((len(X_raw), 1)), X_raw] # add intercept column
34
35w = fit_logistic(X, y, lr=0.2, epochs=3000)
36p = predict_proba(X, w)
37print("\nFinal weights (w0, w1, w2):", w.round(3))
38print("Probabilities:", p.round(3))
39print("weights_match_reference", np.allclose(w, [-4.894, 3.399, -0.959], atol=0.01))1after_epoch= 1 loss=0.6830 w=[0. 0.069 0.075]
2after_epoch= 201 loss=0.4648 w=[-2.205 1.615 -0.462]
3after_epoch= 501 loss=0.4206 w=[-3.524 2.452 -0.686]
4after_epoch=1001 loss=0.4092 w=[-4.356 3.023 -0.85 ]
5after_epoch=3000 loss=0.4072 w=[-4.894 3.399 -0.959]
6
7Final weights (w0, w1, w2): [-4.894 3.399 -0.959]
8Probabilities: [0.079 0.153 0.274 0.777 0.88 0.996 0.687 0.156]
9weights_match_reference TrueThe negative intercept starts low when both signals are absent. Ambiguity's positive coefficient raises review risk, while the negative policy-support coefficient lowers it, holding the other feature fixed. Both feature weights rose on the first step because both features were larger among positive rows on average. Later steps separate their conditional effects; an initial gradient sign isn't the final coefficient sign.
Each progress line reports loss and weights after the same update, so snapshots stay comparable. The loss falls while the weights settle near the reference vector.
At , R3 needs review but scores 0.274, while R7 doesn't but scores 0.687. A straight boundary can't create an isolated pocket. In real data, also inspect labels and missing features before interpreting errors as proof that the relationship needs a nonlinear model.

Round the fitted weights to . For REQ-10234, suppose ambiguity is and auto-policy support is . Compute its score and action at threshold 0.5.
Answer
, so . Route it to review at threshold 0.5. Rounding the weights explains why this differs slightly from the later validation score 0.624 for the same features.
The fitted weights explain the training table. They don't pick a routing policy; that job moves to requests the optimizer never saw.
Select a threshold on held-out requests
The fitted model emits a score in (0, 1). A routing policy needs a threshold t: send the request to review when p >= t. Selecting t on the training rows would reward overfitting.
Use a validation slice with cases the optimizer never saw. Here a false negative means automatic processing when review was required and costs $120. A false positive means unnecessary human review and costs $18.
The code below tests three cuts on the same eight held-out rows. It counts caught reviews (tp), unnecessary reviews (fp), and missed reviews (fn). Precision measures the fraction of the review queue that truly needed review; recall measures the fraction of required reviews caught. Before running it, expect the lowest cut to catch more required reviews and create more queue work.
1import numpy as np
2
3def sigmoid(z: np.ndarray) -> np.ndarray:
4 return 1.0 / (1.0 + np.exp(-z))
5
6def metrics(y: np.ndarray, p: np.ndarray, threshold: float) -> tuple[float, float, float, int]:
7 pred = (p >= threshold).astype(int)
8 tp = int(((pred == 1) & (y == 1)).sum())
9 fp = int(((pred == 1) & (y == 0)).sum())
10 fn = int(((pred == 0) & (y == 1)).sum())
11 precision = tp / (tp + fp) if tp + fp else 0.0
12 recall = tp / (tp + fn) if tp + fn else 0.0
13 f1 = 2 * precision * recall / (precision + recall) if precision + recall else 0.0
14 cost = 120 * fn + 18 * fp
15 return precision, recall, f1, cost
16
17w = np.array([-4.894, 3.399, -0.959])
18X_val_raw = np.array([
19 [1.2, 2.0], [2.1, 4.0], [2.4, 3.0], [3.0, 5.0],
20 [3.6, 4.0], [3.2, 7.0], [4.2, 5.0], [2.8, 6.0],
21])
22y_val = np.array([0, 0, 1, 0, 1, 0, 1, 1])
23X_val = np.c_[np.ones(len(X_val_raw)), X_val_raw]
24p_val = sigmoid(X_val @ w)
25
26for t in (0.20, 0.50, 0.80):
27 precision, recall, f1, cost = metrics(y_val, p_val, t)
28 print(f"t={t:.2f} precision={precision:.3f} recall={recall:.3f} f1={f1:.3f} cost=${cost}")
29
30print("p_val:", np.round(p_val, 3))
31print("cheapest tested cut:", min((0.20, 0.50, 0.80), key=lambda t: metrics(y_val, p_val, t)[3]))1t=0.20 precision=0.667 recall=1.000 f1=0.800 cost=$36
2t=0.50 precision=0.750 recall=0.750 f1=0.750 cost=$138
3t=0.80 precision=1.000 recall=0.500 f1=0.667 cost=$240
4p_val: [0.061 0.169 0.595 0.624 0.971 0.325 0.99 0.244]
5cheapest tested cut: 0.2Among these three cuts, t = 0.20 costs least. Compared with 0.50, it prevents one missed review and adds one unnecessary review: a net saving of $120 - $18 = $102. Once you select the threshold, freeze it and evaluate on an untouched test set. Validation cost was used to choose the policy, so it isn't an independent estimate of the selected policy's performance.[6]
There is also a probability-based decision rule, if the scores are calibrated. At p = 0.20, automatic processing has expected error cost 0.20 × $120 = $24; review has expected error cost 0.80 × $18 = $14.40. In general, review is cheaper when , or . That theoretical cut follows from our zero-cost-correct-decisions assumption. It needn't minimize observed cost on eight noisy cases, and it isn't justified by uncalibrated scores. For a live policy, estimate the full cost table and check reviewer capacity.
At threshold 0.50, validation has FP=1 and FN=1. At threshold 0.20, it has FP=2 and FN=0. If an automatic-processing mistake costs $120 and an unnecessary review costs $18, which threshold wins on validation?
Answer
At 0.50, cost is 1 * $18 + 1 * $120 = $138. At 0.20, cost is 2 * $18 + 0 * $120 = $36. Select 0.20 for this policy experiment, then monitor cost and reviewer capacity on fresh traffic.
Precision, recall, F1, and class imbalance
For any threshold, fill the 2x2 confusion matrix. At t = 0.20, ask one concrete question for each validation request: did the route match the completed case?

At this cut the four cells represent concrete operational outcomes:
- True Positive (TP = 4): required reviews (
V8, V3, V5, V7) routed to review. Correct labels don't prove that authorization or human review succeeds. - False Positive (FP = 2): requests not requiring review (
V6, V4) enter the queue, costing$18each under our assumptions. - True Negative (TN = 2): requests not requiring review (
V1, V2) take the automatic route, with zero incremental cost in this toy table. - False Negative (FN = 0): a request requiring review would take the automatic route, costing
$120in this toy model. This label alone doesn't establish that a breach occurred.
| pred auto | pred review | |
|---|---|---|
| actual 0 | TN 2 (V1, V2) | FP 2 (V6, V4) |
| actual 1 | FN 0 | TP 4 (V8, V3, V5, V7) |
Six requests enter the queue, four of which need review: precision is 4/6 = 2/3. All four required reviews are caught: recall is 4/4 = 1. Their F1 is 2 × (2/3) × 1 / (2/3 + 1) = 0.8.
Read the matrix in the direction of the action:
- Precision =
TP / (TP + FP): among requests routed to review, what fraction required it? Precision measures queue purity and alert trustworthiness. - Recall =
TP / (TP + FN): among requests requiring review, what fraction reached it? Recall measures safety coverage and catch rate. - F1 =
2 * precision * recall / (precision + recall): the harmonic mean of precision and recall.
Precision is undefined when the selected queue is empty; recall is undefined when there are no positive labels. Our threshold helper uses zero for those summaries, a convention rather than an observed rate. Keep the class and queue counts alongside metrics. The reciprocal harmonic-mean expression below assumes positive precision and recall; use the count-based form or a declared zero-division convention at the boundaries.
Why F1 uses the harmonic mean
Why use the harmonic mean rather than the arithmetic average ?
Consider a degenerate detector that flags only a single, glaringly obvious request where ambiguity is extreme (). If that one request is positive, precision is . But across 100 actual risky requests, recall is only .
The arithmetic mean would be , making an utterly useless detector look moderately successful. The harmonic mean inverts this:
Because the harmonic mean is pulled strongly toward the smaller value, it plummets whenever either precision or recall fails. A high F1 requires both rates to remain high.
The accuracy trap under class imbalance
The four cells also explain why accuracy alone is dangerous when positive cases are rare. Try the extreme policy below: send every request to automatic processing.
1import numpy as np
2
3y_true = np.array([1] * 4 + [0] * 96)
4predict_auto_process_everything = np.zeros(100, dtype=int)
5accuracy = (predict_auto_process_everything == y_true).mean()
6recall = ((predict_auto_process_everything == 1) & (y_true == 1)).sum() / (y_true == 1).sum()
7print("accuracy:", round(float(accuracy), 3))
8print("review recall:", round(float(recall), 3))1accuracy: 0.96
2review recall: 0.0The policy gets 96 of 100 rows right by exploiting the negative majority, yet it catches none of the four requests that required review. Accuracy answers a different question from review recall. When positive events are rare, a trivial majority-class baseline scores high accuracy while failing completely at its job.
ROC and precision-recall curves measure ranking
To draw an ROC (Receiver Operating Characteristic) curve, vary the threshold from high to low and plot recall (true-positive rate) on the y-axis against false-positive rate, FP / (FP + TN), on the x-axis. At t = 0.20, those coordinates are (2/4, 4/4) = (0.5, 1.0). Both classes must be present to estimate both rates.
AUC answers a ranking question: how often does a random positive score rank above a random negative score? Ties get half credit. AUC can tell you whether ordering is useful, but it never chooses a business threshold for you.
Before running the pairwise check, count the combinations: four positive and four negative validation rows should produce 16 comparisons.
1import numpy as np
2
3p_val = np.array([0.061, 0.169, 0.595, 0.624, 0.971, 0.325, 0.990, 0.244])
4y_val = np.array([0, 0, 1, 0, 1, 0, 1, 1])
5positive = p_val[y_val == 1]
6negative = p_val[y_val == 0]
7pair_scores = [(pos > neg) + 0.5 * (pos == neg) for pos in positive for neg in negative]
8auc = float(np.mean(pair_scores))
9print("positive-negative pairs:", len(pair_scores))
10print("validation AUC:", round(auc, 3))1positive-negative pairs: 16
2validation AUC: 0.812This slice gives 13 ranking wins out of 16. The ordering is useful, even though one fixed threshold can still make an expensive mistake.

ROC-AUC versus PR-AUC under class imbalance
ROC curves measure threshold-independent discriminative ranking. The two axes normalize exclusively within each ground-truth class:
Population ROC rates depend on the class-conditional score distributions, not directly on the proportion of positives. Changing prevalence while preserving those distributions leaves the population ROC unchanged. Arbitrarily adding benign requests doesn't guarantee this: their score distribution may differ, and finite-sample estimates also vary. Uniformly replicating every negative row preserves an empirical ROC exactly; adding only easy low-score negatives doesn't.
Test that distinction on our eight scores. No model is retrained. Compare an unchanged negative-score mix with 100 extra negatives whose scores are all zero:
1import numpy as np
2
3p = np.array([0.061, 0.169, 0.595, 0.624, 0.971, 0.325, 0.990, 0.244])
4y = np.array([0, 0, 1, 0, 1, 0, 1, 1])
5positive = p[y == 1]
6negative = p[y == 0]
7
8def audit(labels: np.ndarray, scores: np.ndarray) -> tuple[float, float, float]:
9 pos = scores[labels == 1]
10 neg = scores[labels == 0]
11 pairs = (pos[:, None] > neg[None, :]) + 0.5 * (pos[:, None] == neg[None, :])
12 selected = scores >= 0.20
13 precision = float(np.mean(labels[selected]))
14 return float(np.mean(neg >= 0.20)), float(np.mean(pairs)), precision
15
16same_mix_y = np.r_[np.ones(len(positive), dtype=int), np.zeros(25 * len(negative), dtype=int)]
17same_mix_p = np.r_[positive, np.repeat(negative, 25)]
18easy_y = np.r_[y, np.zeros(100, dtype=int)]
19easy_p = np.r_[p, np.zeros(100)]
20
21for name, labels, scores in (
22 ("original", y, p),
23 ("same-negative-mix", same_mix_y, same_mix_p),
24 ("extra-easy-negatives", easy_y, easy_p),
25):
26 fpr, auc, precision = audit(labels, scores)
27 print(name, "fpr", round(fpr, 4), "auc", round(auc, 4), "precision", round(precision, 4))1original fpr 0.5 auc 0.8125 precision 0.6667
2same-negative-mix fpr 0.5 auc 0.8125 precision 0.0741
3extra-easy-negatives fpr 0.0192 auc 0.9928 precision 0.6667The unchanged mix lowers precision by adding false positives in the same proportion. The easy-negative addition instead inflates AUC while leaving the selected queue unchanged. State how an evaluation set was sampled before interpreting its ranking metrics.
That invariance becomes a blind spot when positives are rare. A classifier with recall 0.9 and false-positive rate 0.1 selects 90 true positives and 10 false positives from 100 positives and 100 negatives: precision 90 / 100 = 0.90. With only 100 positives among 10,000 requests (a 1:99 ratio), the exact same rates produce 90 true positives and 990 false positives: precision collapses to 90 / (90 + 990) ≈ 0.083.
Under the assumed unchanged rates, the ROC point stays (0.1, 0.9), but 91.7% of the queue consists of false positives. A precision-recall curve exposes that burden.[1] For rare events, report AP together with counts, recall, cost, and reviewer capacity at the chosen threshold. Neither ROC-AUC nor an area under a PR curve alone measures operating effort.

Average precision (AP) summarizes the precision-recall curve with a stepwise weighted average of precision at each positive-label recall increase. People sometimes label AP as PR-AUC, but trapezoidal area under the same curve can produce a different number, so name the statistic you compute.[7]
Here the four positive rows enter at ranks 1, 2, 4, 6, with precisions 1, 1, 3/4, 4/6. Each adds one quarter of the total recall, so AP is their mean, about 0.854. This shortcut assumes distinct scores. Tied scores must enter together at one threshold; use average_precision_score for that general case.[7]
1import numpy as np
2
3p_val = np.array([0.061, 0.169, 0.595, 0.624, 0.971, 0.325, 0.990, 0.244])
4y_val = np.array([0, 0, 1, 0, 1, 0, 1, 1])
5
6# For distinct scores, AP is the mean precision at each positive row's rank.
7order = np.argsort(-p_val)
8y_sorted = y_val[order]
9tp_cum = np.cumsum(y_sorted)
10fp_cum = np.cumsum(1 - y_sorted)
11precision_at = tp_cum / (tp_cum + fp_cum)
12ap = float(precision_at[y_sorted == 1].mean())
13print("validation average precision:", round(ap, 3))
14print("validation ROC-AUC: ", 0.812)1validation average precision: 0.854
2validation ROC-AUC: 0.812On this balanced slice AP and ROC-AUC are close. When positives become rare, report both. ROC-AUC can stay high while AP falls with positive prevalence and with the quality of ranking near the top of the queue.
Calibration: do scores represent frequencies?
A model can rank access-change requests well while producing over-confident or under-confident numbers. If requests scored near 0.80 require review only half the time, showing 80% to an operator misstates risk.
Logistic regression is often described as "already a probability model" because it optimizes log loss. That interpretation depends on the linear log-odds assumption and the data.
Shrinkage (scikit-learn's default C = 1.0, an L2 penalty) pulls coefficients toward zero, which can narrow the score range and move predictions toward the base rate. Even this unregularized eight-row fit isn't a frequency guarantee. Audit calibration on held-out labels.[1]
A reliability diagram compares average score with observed positive frequency per bin. Our lowest bin has three requests with mean score 0.158, and one actually needs review (1/3 ≈ 0.333). Its absolute gap is about 0.175; because it contains 3/8 of the data, it contributes about 0.066 to expected calibration error (ECE).
We use binary positive-class ECE: bin the review probabilities, not the confidence in whichever class won. With denoting the rows in bin , sum these weighted gaps:
ECE depends on sample size, prevalence, and bin choices. It has no universal pass threshold. Pair it with a Brier score, the mean squared difference between predicted probabilities and binary outcomes, when you need a bin-independent probability-quality summary. Brier score measures more than calibration: discrimination and outcome uncertainty matter too, so a lower Brier score doesn't by itself prove better calibration.[8]
State the binning scheme and measure on held-out labels. Fit a post-hoc map on data separate from the classifier's fitting data, then verify the map on fresh labels. Platt scaling fits a one-dimensional logistic curve; isotonic regression fits a nondecreasing step function; temperature scaling divides logits by a learned positive scalar.[9] A strictly increasing map preserves binary ranking, while isotonic can introduce ties and change AUC.[8] Reassess the numerical operating threshold on suitable validation data after calibration; don't tune it on the final test labels. Positive temperature scaling alone leaves the binary 0.5 decision boundary unchanged.
The next output makes the ECE bookkeeping visible rather than hiding it behind one number.
1import numpy as np
2
3p_val = np.array([0.061, 0.169, 0.595, 0.624, 0.971, 0.325, 0.990, 0.244])
4y_val = np.array([0, 0, 1, 0, 1, 0, 1, 1])
5edges = np.linspace(0.0, 1.0, 5)
6ece = 0.0
7for lo, hi in zip(edges[:-1], edges[1:]):
8 mask = (p_val >= lo) & ((p_val < hi) if hi < 1.0 else (p_val <= hi))
9 if mask.any():
10 avg_p = p_val[mask].mean()
11 frac_pos = y_val[mask].mean()
12 ece += mask.mean() * abs(avg_p - frac_pos)
13 closing = "]" if hi == 1.0 else ")"
14 print(f"[{lo:.2f}, {hi:.2f}{closing} n={mask.sum()} avg_p={avg_p:.3f} frac_pos={frac_pos:.3f}")
15print("validation ECE:", round(float(ece), 3))
16print("validation Brier:", round(float(np.mean((p_val - y_val)**2)), 3))1[0.00, 0.25) n=3 avg_p=0.158 frac_pos=0.333
2[0.25, 0.50) n=1 avg_p=0.325 frac_pos=0.000
3[0.50, 0.75) n=2 avg_p=0.609 frac_pos=0.500
4[0.75, 1.00] n=2 avg_p=0.980 frac_pos=1.000
5validation ECE: 0.139
6validation Brier: 0.158In this sample, the first bin's observed review frequency exceeds its mean prediction (0.333 versus 0.158); the second is lower (0 versus 0.325). Those tiny bin counts don't establish population miscalibration. Calling the first bin "under-confident" would also obscure its high predicted confidence in the negative class. ECE uses an absolute gap regardless of direction.
Eight validation requests are enough for arithmetic, not a release gate. To see why, flip one disputed label and measure both metrics:
1import numpy as np
2
3def f1_at_half(y: np.ndarray, p: np.ndarray) -> float:
4 pred = p >= 0.5
5 tp = ((pred == 1) & (y == 1)).sum()
6 fp = ((pred == 1) & (y == 0)).sum()
7 fn = ((pred == 0) & (y == 1)).sum()
8 return float(2 * tp / (2 * tp + fp + fn))
9
10def ece(y: np.ndarray, p: np.ndarray) -> float:
11 edges = np.linspace(0.0, 1.0, 5)
12 total = 0.0
13 for lo, hi in zip(edges[:-1], edges[1:]):
14 mask = (p >= lo) & ((p < hi) if hi < 1.0 else (p <= hi))
15 if mask.any():
16 total += mask.mean() * abs(p[mask].mean() - y[mask].mean())
17 return float(total)
18
19p = np.array([0.061, 0.169, 0.595, 0.624, 0.971, 0.325, 0.990, 0.244])
20base = np.array([0, 0, 1, 0, 1, 0, 1, 1])
21one_flip = base.copy()
22one_flip[3] = 1
23for name, labels in (("base", base), ("one flip", one_flip)):
24 print(f"{name:8s} f1={f1_at_half(labels, p):.3f} ece={ece(labels, p):.3f}")1base f1=0.750 ece=0.139
2one flip f1=0.889 ece=0.209One label moves both results sharply. Treat this slice as a teaching audit, not a release gate. Two questions remain: what does each threshold cost, and do the scores describe frequencies? First, compare routing cost across thresholds:

Cost chooses an action. It doesn't say whether 0.87 means 87%. The reliability diagram asks that second question:

Validation AUC is 0.812 and four-bin ECE is 0.139, both measured on only eight requests. A product manager wants to display "This access request has an 87% chance of manual review." What two issues need checking?
Answer
The target is whether review is required; the actual route follows a threshold deterministically. Even a calibrated score wouldn't mean an 87% chance of that routing action. Eight labels also provide too little evidence for a precise user-facing probability of the target event. Collect more relevant labels, reserve calibration/verification data, and test reliability before displaying such estimates.
Verify the NumPy fit with scikit-learn
A from-scratch implementation gains a useful check by matching a maintained implementation under equivalent settings.[3] The September 21, 2026 library check uses NumPy 2.5.3 and scikit-learn 1.9.1. LogisticRegression defaults to L2 regularization (C=1.0, l1_ratio=0). Since 1.8, C=np.inf selects an unpenalized fit and penalty is deprecated; use C and l1_ratio with a compatible solver.[10]
Set C=np.inf here so the library solve is comparable to our unregularized gradient-descent fit. The solver is L-BFGS rather than full-batch gradient descent, so the weights should be close, not bit-identical.
1import numpy as np
2from sklearn.linear_model import LogisticRegression
3from sklearn.metrics import f1_score, roc_auc_score
4
5X_raw = np.array([
6 [1.0, 1.0], [1.5, 2.0], [2.0, 3.0], [3.5, 6.0],
7 [4.0, 7.0], [4.5, 5.0], [2.8, 4.0], [3.2, 8.0],
8])
9y = np.array([0, 0, 1, 1, 1, 1, 0, 0])
10X = np.c_[np.ones((len(X_raw), 1)), X_raw]
11w_ours = np.array([-4.894, 3.399, -0.959])
12X_val_raw = np.array([
13 [1.2, 2.0], [2.1, 4.0], [2.4, 3.0], [3.0, 5.0],
14 [3.6, 4.0], [3.2, 7.0], [4.2, 5.0], [2.8, 6.0],
15])
16y_val = np.array([0, 0, 1, 0, 1, 0, 1, 1])
17X_val = np.c_[np.ones(len(X_val_raw)), X_val_raw]
18
19clf = LogisticRegression(C=np.inf, solver="lbfgs", max_iter=5000, fit_intercept=False)
20clf.fit(X, y)
21w_sk = clf.coef_[0]
22print("sklearn weights:", np.round(w_sk, 3))
23print("our weights :", w_ours)
24print("weights close :", np.allclose(w_sk, w_ours, atol=0.05))
25p_val = clf.predict_proba(X_val)[:, 1]
26print("validation F1 @ 0.20:", round(f1_score(y_val, (p_val >= 0.20).astype(int)), 3))
27print("validation AUC:", round(roc_auc_score(y_val, p_val), 3))1sklearn weights: [-4.914 3.414 -0.964]
2our weights : [-4.894 3.399 -0.959]
3weights close : True
4validation F1 @ 0.20: 0.8
5validation AUC: 0.812Close weights and matching metrics support our implementation on this fixture; they aren't a proof for other inputs. Compare objectives, logits, rank, and stopping tolerances when a match fails. The validation metrics describe eight constructed rows, not release performance.
Multiclass routing: the softmax extension
Binary logistic regression decides between two actions. Live policy engines often require more granular dispatch. Suppose our access-control engine sorts requests into three tiers:
- Class 0:
Auto-Approve(standard permissions under established policy) - Class 1:
Standard Review(moderate risk ticket queue) - Class 2:
Security Escalation(severe risk requiring immediate review and credential freeze)
These are classifier labels. Permission checks and any credential actions are separately authorized workflow steps.
Linear heads and the softmax function
Instead of a single scalar logit, we assign each of the classes its own weight vector (including intercept):
In matrix form, the logit vector for one request is , where .
To convert unconstrained real logits into a valid probability distribution where each and , we apply the softmax function:[4][11]
Notice that when , softmax simplifies directly back to the binary sigmoid. Setting class logits to and :
Binary logistic regression is two-class softmax regression with the reference class logit pinned to zero.
Categorical cross-entropy and its gradient
For a one-hot ground-truth vector ( for the true class, otherwise), the likelihood of a single observation under the categorical distribution is . Taking the negative log-likelihood across independent requests yields the Categorical Cross-Entropy loss:
Differentiating categorical cross-entropy with respect to class weight vector gives:
Stacking all weight vectors into weight matrix produces the batch gradient:
where , contains the softmax probabilities, and is the one-hot target matrix. The exact same error-residual form, features dotted with predicted minus actual, extends directly from scalar linear regression to multiclass classification.
Adding the same coefficient vector to every class head adds an identical feature-dependent value to every logit and leaves softmax unchanged. Thus unrestricted class weights aren't identifiable from probabilities alone. A reference class or an appropriate constraint removes that redundancy; convexity doesn't imply a unique unrestricted weight matrix.
Numerical stability: the log-sum-exp trick
Direct exponentials overflow around logits 709.8 in float64 or 88.7 in float32. Subtracting preserves softmax mathematically and stabilizes those exponentials:
With representable finite differences, , so the largest exponential is 1. Tiny probabilities can still underflow or round to zero; invalid inputs and extreme subtraction overflow need separate handling. Compute cross-entropy from shifted logits, not clipped probabilities. Clipping would cap the penalty for an extremely wrong prediction.
Fit a three-class routing classifier on six historical requests using NumPy:
1import numpy as np
2
3def softmax(Z: np.ndarray) -> np.ndarray:
4 shift_Z = Z - np.max(Z, axis=1, keepdims=True)
5 exp_Z = np.exp(shift_Z)
6 return exp_Z / np.sum(exp_Z, axis=1, keepdims=True)
7
8def categorical_cross_entropy_from_logits(Y_onehot: np.ndarray, Z: np.ndarray) -> float:
9 shifted = Z - np.max(Z, axis=1, keepdims=True)
10 log_normalizer = np.log(np.sum(np.exp(shifted), axis=1, keepdims=True))
11 log_probabilities = shifted - log_normalizer
12 return float(-np.mean(np.sum(Y_onehot * log_probabilities, axis=1)))
13
14# 6 requests: [ambiguity, auto-policy support] + intercept column
15# Target routing classes: 0: Auto-Approve, 1: Standard Review, 2: Escalation
16X_raw = np.array([
17 [1.0, 8.0],
18 [1.5, 7.0],
19 [2.5, 4.0],
20 [3.2, 5.0],
21 [4.8, 1.5],
22 [5.0, 1.0],
23])
24X = np.c_[np.ones(len(X_raw)), X_raw]
25y_classes = np.array([0, 0, 1, 1, 2, 2])
26K = 3
27N, d = X.shape
28Y_onehot = np.eye(K)[y_classes]
29
30W = np.zeros((d, K))
31lr = 0.2
32for epoch in range(1000):
33 P = softmax(X @ W)
34 grad = (1 / N) * (X.T @ (P - Y_onehot))
35 W -= lr * grad
36
37P_final = softmax(X @ W)
38preds = np.argmax(P_final, axis=1)
39loss = categorical_cross_entropy_from_logits(Y_onehot, X @ W)
40print("final loss:", round(loss, 4))
41print("predicted classes:", preds.tolist())
42print("all correct:", bool((preds == y_classes).all()))
43print("sample probabilities row 0 (auto):", P_final[0].round(3).tolist())
44print("sample probabilities row 4 (escalate):", P_final[4].round(3).tolist())
45
46extreme_logits = np.array([[1000.0, 0.0, -1000.0]])
47extreme_target = np.array([[0.0, 0.0, 1.0]])
48extreme_loss = categorical_cross_entropy_from_logits(extreme_target, extreme_logits)
49assert np.isfinite(extreme_loss) and np.isclose(extreme_loss, 2000.0)
50print("extreme wrong log loss:", round(extreme_loss, 1))1final loss: 0.0055
2predicted classes: [0, 0, 1, 1, 2, 2]
3all correct: True
4sample probabilities row 0 (auto): [1.0, 0.0, 0.0]
5sample probabilities row 4 (escalate): [0.0, 0.005, 0.995]
6extreme wrong log loss: 2000.0Suppose a model produces logits for classes [Auto, Review, Escalate]. If you add 100 to all three logits, , what happens to the resulting softmax probabilities?
Answer
The probabilities don't change. Adding a constant to every logit scales both numerator and denominator by , which cancels out. Softmax depends only on the relative logit differences .
"All correct" scores the six training rows. The printed [1.0, 0.0, 0.0] is rounded, not calibrated certainty. The extreme wrong prediction costs 2,000 nats; clipping its target probability at 1e-15 would misleadingly report only about 34.54. This separable toy fit also doesn't establish finite maximum-likelihood coefficients.
The logistic and softmax models give linear log-odds scores followed by normalized exponential curves. Before moving to trees, compare that objective with two other uses of feature geometry: a separating margin and a covariance kernel.
Margins, kernels, and predictive uncertainty
A support vector machine (SVM) can classify the same access requests with a linear score . Its question differs from logistic regression's: can the classes be separated with a wide geometric buffer around the boundary?
In the usual soft-margin SVM, support vectors lie on or inside the margin, or on the wrong side of the decision boundary. Rows comfortably beyond the margin don't contribute hinge-loss pressure.[12]
Encode needs review as and review not required as . The functional margin is : positive means the request is classified correctly. Dividing by the feature-weight norm gives the signed geometric margin for a linear model.
A standard hinge loss charges:
A positive request with score 0.4 is classified correctly, but its margin is only 0.4, so hinge loss is 0.6. A positive score of 1.4 clears the unit margin and incurs zero hinge loss.
The logistic loss on the second request is about 0.220; hinge loss is zero after the unit functional margin. A usual soft-margin objective balances violations with weight regularization, for example . LinearSVC defaults to squared hinge and also regularizes the intercept through liblinear. Name the loss, penalty, and normalization when comparing implementations.[12]
A linear SVM still can't carve out an isolated exception that requires a curved boundary. A kernel method replaces an explicit feature expansion with a pairwise similarity function. The radial basis function (RBF) kernel is one common choice:
Here is the length scale. Identical points have similarity 1. With , points one unit apart have similarity 0.607, and points three units apart have similarity 0.011.
An RBF SVM can draw nonlinear boundaries around feature neighborhoods without materializing an infinite-dimensional feature map. Its decision score still isn't a calibrated class probability. Apply a separate held-out calibration procedure before displaying percentages, and scale features before treating Euclidean distance as meaningful.[12]
A Gaussian process (GP) can use that same RBF kernel for a different job: define covariance between unknown function values and return a predictive mean together with uncertainty. Switch briefly to a regression example, not the binary review label: observe a continuous value at input 0. Assume zero prior mean, unit prior variance, length scale 1, and Gaussian observation-noise variance 0.04. At the observed input, the posterior mean is 2/1.04 ≈ 1.923 and latent variance is 1 - 1/1.04 ≈ 0.0385. At a new input the same conditioning calculation gives:[13]
Near the observation, the mean stays close to 2 and uncertainty shrinks. Far away, similarity vanishes, the mean returns toward its zero prior, and standard deviation rises toward 1. This variance describes the latent function; predicting a new noisy observation adds 0.04 to it. GP classification requires a non-Gaussian likelihood and doesn't use these regression equations unchanged.
This is predictive uncertainty under the stated kernel, noise, and prior assumptions, not proof that every interval will be calibrated on shifted traffic. Exact dense GP regression also requires roughly fitting time and storage for observations, so it fits smaller data problems better than large unapproximated serving datasets.[13]
The following dependency-free comparison makes both distinctions concrete: it prints the two loss objectives and the one-observation GP posterior:
1from math import exp, log1p, sqrt
2
3examples = [
4 ("required-review", 1, 0.4),
5 ("confident-review", 1, 1.4),
6 ("missed-review", 1, -0.2),
7 ("safe-auto", -1, -1.2),
8]
9
10for name, label, score in examples:
11 margin = label * score
12 hinge = max(0.0, 1.0 - margin)
13 logistic = log1p(exp(-margin))
14 print(f"{name:16} margin={margin:4.1f} hinge={hinge:.3f} logistic={logistic:.3f}")
15
16print("distance kernel gp_mean gp_std")
17for distance in [0.0, 1.0, 3.0]:
18 similarity = exp(-(distance**2) / 2.0)
19 mean = 2.0 * similarity / 1.04
20 standard_deviation = sqrt(1.0 - similarity**2 / 1.04)
21 print(f"{distance:8.1f} {similarity:.3f} {mean:.3f} {standard_deviation:.3f}")1required-review margin= 0.4 hinge=0.600 logistic=0.513
2confident-review margin= 1.4 hinge=0.000 logistic=0.220
3missed-review margin=-0.2 hinge=1.200 logistic=0.798
4safe-auto margin= 1.2 hinge=0.000 logistic=0.263
5distance kernel gp_mean gp_std
6 0.0 1.000 1.923 0.196
7 1.0 0.607 1.166 0.804
8 3.0 0.011 0.021 1.000The margin rows compare two loss terms; fitted models also depend on their penalties. In our one-observation stationary-RBF example, uncertainty rises with distance from that observation. With other kernels, priors, and several observations, uncertainty needn't be a simple monotone function of distance.
| Method | What it optimizes or predicts | What still needs checking |
|---|---|---|
| Logistic regression | Log loss and a sigmoid-based class score | Calibration, class balance, and the decision threshold |
| Linear SVM | A separating margin under hinge or squared-hinge loss | Feature scaling, regularization, and separate probability calibration |
| RBF-kernel SVM | A nonlinear margin based on distance to support vectors | Kernel bandwidth, overfitting, and held-out operating cost |
| Gaussian process | A posterior mean and variance under a chosen covariance kernel | Prior assumptions, label noise, shift, and cubic exact-training cost |
If a run diverges from its expected numbers, diagnose the boundary between data, objective, and implementation before blaming the optimizer.
When the numbers look wrong
| Symptom | Evidence to inspect | Next check |
|---|---|---|
| Scores cluster near the base rate | Signal strength, inherent uncertainty, penalties, and convergence | Compare with a constant baseline on fresh labels before adding features or epochs |
| Lower threshold catches more reviews | Queue counts, false-positive burden, cost table, and capacity | Choose a policy on validation and score it on an untouched test set |
| Good AUC but large held-out bin gaps | Bin populations, label quality, sampling, and subgroup/traffic differences | Fit a justified calibrator on separate data, then verify reliability |
| NumPy and library coefficients differ | Equivalent designs/objectives, intercept handling, scaling, rank, and stopping tolerances | Compare logits and objective values as well as weights; use C=np.inf for this unpenalized comparison |
| Loss grows after an update | Finite inputs, gradient calculation, step size, and the loss implementation | Trace the first failing update; compare a smaller rate and stable logit-space loss |
Try it yourself
Use the same training and validation rows. Follow the order: verify the update, stress the optimizer, choose a policy, then test how fragile the probability story is. Fit only on training; choose policy only on validation.
- Start with
w = [0, 0, 0]and compute the gradient vector on the full batch by hand for the initial step (lr = 0.2). Verify it matches the first printed line of the NumPy run:w = [0, 0.069, 0.075]. - Change the learning rate to
2.0. Print the first ten losses instead of only the final loss. Does each update improve the objective? Explain any increases in terms of step size and feature scale. - Sweep thresholds
[0.2, 0.35, 0.5, 0.65, 0.8]on validation and print precision, recall, F1, and cost. Which threshold minimizes the stated cost? - Bin validation probabilities into four equal-width bins, compute ECE by hand, and compare with code output. Repeat with eight bins and explain why the number changes.
- Flip one validation label and measure F1 and ECE again. Write the argument for gathering a larger, later-in-time validation set before publishing probability claims.
Check your calculations: the initial mean gradient is [0, -0.34375, -0.375]. The five thresholds give costs $36, $138, $138, $240, $240, respectively. Four-bin ECE is about 0.139; eight-bin ECE is about 0.154 using the displayed three-decimal scores. The changed bins expose gaps that previously canceled within a bin. Neither value settles calibration on such a small sample.
On the training table, R3 still scores only 0.274 at the fitted weights: moderate ambiguity, moderate policy support, but it needed review. A linear score has no way to carve out that exception without also moving nearby automatic requests.
Next, decision trees split the same eight-request table with axis-aligned rules, then forests and boosting combine those rules. The later cross-validation lesson deepens the evidence required before shipping a model.