Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
You're investigating a production incident during a high-traffic release. Pull three customer support tickets into an incident table. Each row tracks two binary signals: did the user report high latency, and did they request an emergency rollback? Ticket t2 flags both issues at once. Intuitively, the joint latency-plus-rollback pattern dominates your dataset, while the contrast between latency-only and rollback-only looks like a secondary detail.
What happens if you compress this table to save memory or speed up vector retrieval? If you discard the weaker contrast direction, a latency-only ticket and a rollback-only ticket collapse into identical representations. Your search or routing pipeline suddenly can't tell an API timeout from a broken deployment.
Singular value decomposition (SVD) finds perpendicular input directions and measures the output length produced by each unit direction. For this uncentered table, those strengths describe magnitude relative to zero, rather than variation around the average ticket. Choosing which directions to keep turns linear algebra into a compression decision whose cost we can calculate.
In the vectors, matrices, and tensors lesson, we worked with rows, columns, and dot products. Keep those operations close: three tickets and two features give us everything we need to compute, visualize, and diagnose a full SVD by hand.
A matrix can hide repeated directions
Write those flags as a 3×2 matrix. Each row is one ticket; each column records a latency or rollback signal:
Look at t2 and predict its score along two diagonals. Because it contains both signals, it should reinforce [latency + rollback]; its contribution to [latency - rollback] should cancel.
First compare the raw patterns [1, 1] and [1, -1]. Ticket t2 = [1, 1] scores 1 + 1 = 2 on the first and 1 - 1 = 0 on the second. Both pattern vectors have length . Divide by that length to compare unit vectors, whose length is 1:
Multiply A by either direction. Each entry is one ticket's score for that pattern. The common scores are , so their squared lengths add to . For contrast, the scores are and the squared lengths add to . Taking square roots gives strengths and . NumPy checks each step:
1import numpy as np
2
3A = np.array([
4 [1.0, 0.0], # latency
5 [0.0, 1.0], # rollback
6 [1.0, 1.0], # latency and rollback
7])
8common = np.array([1.0, 1.0]) / np.sqrt(2)
9contrast = np.array([1.0, -1.0]) / np.sqrt(2)
10
11common_scores = A @ common
12contrast_scores = A @ contrast
13
14print("A_shape", A.shape)
15print("common_scores", common_scores.round(4).tolist())
16print("contrast_scores", contrast_scores.round(4).tolist())
17print("strengths", round(np.linalg.norm(common_scores), 4), round(np.linalg.norm(contrast_scores), 4))1A_shape (3, 2)
2common_scores [0.7071, 0.7071, 1.4142]
3contrast_scores [0.7071, -0.7071, 0.0]
4strengths 1.7321 1.0The common direction has strength ; the contrast direction has strength . Nothing has been compressed yet. This is a measurement of which feature mix the rows support most strongly.

Coordinates come from a basis
The common and contrast directions are linear combinations of the original feature axes: scale each axis, then add. The span is every vector reachable by those combinations.
These directions are orthogonal: their dot product is , so they meet at a right angle. Neither is a multiple of the other, which makes them linearly independent. Together they form a basis: a set of independent vectors that spans the entire two-dimensional feature space. Because they're also unit vectors, this basis is orthonormal.
To practice changing coordinates, allow mention counts instead of yes-or-no flags for a new row . Its common score is ; its contrast score is .
Because this basis is orthonormal, its coordinates are dot products with the two unit directions:
Each coefficient is an orthogonal projection coordinate. Multiply a coefficient by its unit direction to recover the part of x carried by that pattern:
| Component | Coefficient × unit direction | Contribution |
|---|---|---|
| common | [2, 2] | |
| contrast | [1, -1] | |
| sum | [3, 1] |
So . The A @ v scores above measure every ticket along one direction; these coordinates rebuild one ticket from its directions. Later, projecting onto fewer SVD directions will keep selected coordinates and deliberately discard the rest.
We chose this basis by inspection. For real data, you need the matrix to reveal it. A.T @ A makes that choice visible without changing the feature space.
The Gram matrix reveals those directions by hand
For a data matrix A, compare feature columns with one another by multiplying on the feature side:
The first diagonal entry is the latency column's dot product with itself: . The off-diagonal entry is the columns' overlap: . G has shape [2, 2] because its rows and columns both index features.
A nonzero vector v satisfying is an eigenvector. Multiplying by G keeps it on the same line through the origin. Here the eigenvalue is nonnegative because G is a Gram matrix, so it stretches or shrinks the vector without reversing it. For a unit eigenvector, is squared strength because
when v has length 1.
For this matrix, the common direction is scaled by 3 and the contrast direction by 1:

1import numpy as np
2
3A = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
4G = A.T @ A
5common = np.array([1.0, 1.0]) / np.sqrt(2)
6contrast = np.array([1.0, -1.0]) / np.sqrt(2)
7
8print("G", G.tolist())
9print("G_common", (G @ common).round(4).tolist())
10print("three_common", (3 * common).round(4).tolist())
11print("G_contrast_equals_contrast", np.allclose(G @ contrast, contrast))1G [[2.0, 1.0], [1.0, 2.0]]
2G_common [2.1213, 2.1213]
3three_common [2.1213, 2.1213]
4G_contrast_equals_contrast TrueWhy do eigenvalues of A.T @ A become squared singular values?
Answer
For a unit direction v, the number ||A @ v|| is how much A stretches that direction. Squaring the length gives (A @ v).T @ (A @ v), which equals v.T @ A.T @ A @ v. If v is an eigenvector, that value is its eigenvalue. SVD reports the stretch, so it takes the square root.
SVD names the directions and strengths
Any real matrix, including this 3×2 incident table, has a singular value decomposition (SVD):[1]
The incident matrix A has shape [3, 2]. Its reduced SVD keeps min(3, 2) = 2 directions. Reduced SVD still includes zero singular values if the matrix has dependent columns; it hasn't made a compression decision yet.[2]
| Factor | Shape | Meaning here |
|---|---|---|
[2, 2] | directions in feature space, such as common signal versus contrast | |
[2, 2] | non-negative strengths, ordered largest first | |
[3, 2] | how each ticket participates in each direction |
Here, the right singular vectors are the normalized common and contrast directions. Their singular values tell us how strongly A stretches those directions:
For each positive singular value, the corresponding left singular vector is . For the common pattern, dividing its three scores by gives . For contrast, . Columns of U have unit length; the scaled columns U * S contain the actual ticket scores. Put all three factors back together and they reconstruct every original entry.
For computation, call the SVD routine directly. It returns U, the one-dimensional singular-value array S, and Vt, which is . Before running it, predict the shapes: reduced U should have one row per ticket, and Vt should have one row per retained feature direction.
1import numpy as np
2
3A = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
4U, S, Vt = np.linalg.svd(A, full_matrices=False)
5reconstructed = U @ np.diag(S) @ Vt
6
7print("shapes", U.shape, S.shape, Vt.shape)
8print("singular_values", S.round(4).tolist())
9print("reconstruction_error", float(np.max(np.abs(A - reconstructed))))
10print("reconstructs_A", np.allclose(A, reconstructed))1shapes (3, 2) (2,) (2, 2)
2singular_values [1.7321, 1.0]
3reconstruction_error 4.440892098500626e-16
4reconstructs_A TrueThe sign of a singular vector may differ between correct implementations. If matching directions in U and V flip sign together, their product stays the same. Repeated singular values also permit rotations within the tied subspace. Compare singular values and reconstruction, not raw vectors. The tiny printed reconstruction error is floating-point roundoff, not discarded information.
Rank counts independent patterns
The rank of a matrix is the number of non-zero singular values. It counts how many independent directions the matrix really needs, not how many columns happen to be present.
Now make the failure concrete. Suppose every rollback feature exactly repeats the latency feature:
The second column adds no new information. B still has two columns, but only one independent direction, so its rank is 1. If you had to predict a new row from one column, the duplicate couldn't improve that prediction.
1import numpy as np
2
3B = np.array([
4 [1.0, 1.0],
5 [0.0, 0.0],
6 [1.0, 1.0],
7])
8singular_values = np.linalg.svd(B, compute_uv=False)
9
10print("shape", B.shape)
11print("singular_values", singular_values.round(6).tolist())
12print("rank", int(np.linalg.matrix_rank(B)))1shape (3, 2)
2singular_values [2.0, 0.0]
3rank 1A feature table with rank below its column count deserves inspection. It may contain a deliberate encoding, or two features may duplicate each other and waste compute. Rank exposes the structural fact; domain meaning tells you whether that fact is a bug.
In general, rank is the dimension of a matrix's column space, the span of its columns, and also the dimension of its row space.
Exact rank is a mathematical definition. Floating-point calculations needn't produce a perfect zero for a dependent direction, so np.linalg.matrix_rank treats sufficiently small singular values as zero using a tolerance. Its default accounts for roundoff and matrix dimensions; a measurement-noise threshold may need to be larger. Numerical rank depends on that threshold.[3]
Truncation trades detail for a smaller matrix
Return to full-rank incident matrix A. Keeping both singular directions reconstructs it exactly. What would one direction lose? Keeping only the strongest creates a rank-one approximation:
Among all rank-one matrices, this approximation has the smallest Frobenius reconstruction error: the square root of the sum of squared entry errors.[4] Equivalently, it minimizes that squared sum. Notice what the guarantee covers: reconstructing A. It doesn't promise that a downstream classifier or retriever keeps its quality.
1import numpy as np
2
3A = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
4U, S, Vt = np.linalg.svd(A, full_matrices=False)
5rank_one = U[:, :1] @ np.diag(S[:1]) @ Vt[:1, :]
6
7error = np.linalg.norm(A - rank_one, ord="fro")
8
9print("rank_one")
10print(rank_one.round(3))
11print("frobenius_error", round(float(error), 4))
12print("discarded_singular_value", round(float(S[1]), 4))1rank_one
2[[0.5 0.5]
3 [0.5 0.5]
4 [1. 1. ]]
5frobenius_error 1.0
6discarded_singular_value 1.0The discarded direction has strength 1. Look at the top two rows of A₁: a latency-only ticket and a rollback-only ticket both become [0.5, 0.5]. Their contrast is gone. Four entries each have absolute error 0.5, so the Frobenius error is .
A₁ is still a 3×2 matrix. The smaller representation is its one coordinate per ticket, A @ Vt[:1].T, plus the saved feature direction needed for reconstruction. Rebuilding the full matrix doesn't save storage. This distinction matters when compressing embeddings.

PCA is SVD after centering each feature
So far, zero meant "no signal." Principal component analysis (PCA) changes the question: which directions describe how tickets vary around their average row?
That question requires centering first. Given a matrix X of tickets by numeric features, subtract each column's average:
For our three tickets, both column means are . Subtraction changes t0 from [1, 0] to [1/3, -2/3]: it has more latency and less rollback than the average ticket. The centered rows are [1/3, -2/3], [-2/3, 1/3], and [1/3, 1/3].
Run SVD on those rows and the right singular vectors give principal-component directions. They're also eigenvectors of the sample covariance matrix centered.T @ centered / (n - 1), where n is the number of rows. Its entries measure how feature columns vary together. For this example, centering makes the contrast direction strongest. Check which two tickets differ most from each other before running the code.[5]
1import numpy as np
2
3X = np.array([
4 [1.0, 0.0],
5 [0.0, 1.0],
6 [1.0, 1.0],
7])
8centered = X - X.mean(axis=0)
9_, S, Vt = np.linalg.svd(centered, full_matrices=False)
10pc1 = Vt[0]
11if pc1[0] < 0:
12 pc1 = -pc1
13
14retained = S[0] ** 2 / np.sum(S ** 2)
15
16print("column_means", X.mean(axis=0).round(3).tolist())
17print("centered_means_are_zero", np.allclose(centered.mean(axis=0), 0))
18print("pc1", pc1.round(4).tolist())
19print("pc1_variance_fraction", round(float(retained), 4))1column_means [0.667, 0.667]
2centered_means_are_zero True
3pc1 [0.7071, -0.7071]
4pc1_variance_fraction 0.75Uncentered SVD favored common signal because the rows sit mostly in the positive diagonal direction from the origin. Centered PCA favors contrast because the latency-only and rollback-only tickets are the most separated around their mean. Both results are correct: we asked different questions of the same rows.
PCA finds variance, not relevance. It also centers without automatically rescaling feature units. If one feature is measured in milliseconds and another in seconds, choose the scaling deliberately before fitting. For a new row, reuse the training column means and the fitted directions; don't fit a fresh PCA basis per query.[5]
Avoid materializing a centered sparse matrix
PCA centers every feature mathematically. In a sparse term-document matrix, subtracting a nonzero column mean turns each implicit zero into . Explicitly building that centered matrix destroys sparsity. This doesn't make PCA impossible on sparse input: scikit-learn supports sparse inputs with selected solvers, including arpack. The implementation can handle centering without constructing the full centered data table.[5]
A 100,000 by 10,000 matrix has one billion cells. At 1% density, a sparse format stores about 10 million values plus indices. The dense values alone take about 4 GB as float32, or about 8 GB as float64. Those are order-of-magnitude values; index widths and format add their own cost.
If directions relative to the origin are what you want, TruncatedSVD works directly on sparse counts or TF-IDF values (term counts reweighted by how common a term is across documents). Unlike PCA, it doesn't center; this is the usual latent semantic analysis setup.[6] The choice changes the question, not just memory use. Count the nonzero entries before and after centering a small dense array with roughly 1% nonzero values:
1import numpy as np
2
3rng = np.random.default_rng(42)
4shape = (1000, 100) # 100,000 possible entries
5X = (rng.random(shape) < 0.01).astype(float)
6centered = X - X.mean(axis=0)
7
8print("explicit_values_before", int(np.count_nonzero(X)))
9print("explicit_values_after", int(np.count_nonzero(centered)))
10print("column_means_all_nonzero", bool(np.all(X.mean(axis=0) != 0)))
11
12assert np.count_nonzero(X) < 1500
13assert np.count_nonzero(centered) == 1000001explicit_values_before 1047
2explicit_values_after 100000
3column_means_all_nonzero TrueThe example is deliberately small. For a large sparse matrix, don't materialize X or X - mean; use a sparse-aware decomposition path.
Least squares projects a target onto a feature span
SVD also supports the fitting problem behind linear regression. Suppose the three tickets have target severities
The two columns of A are the available latency and rollback features. The first two rows demand weights 1 and 2, but those weights predict 3 for the last row rather than its target 2. No combination matches all three targets. Choose weights that make the squared residual as small as possible:
This is least squares.[1] The fitted vector is the orthogonal projection of y onto the span of A's columns. In other words, choose the point in that feature span closest to the target. For this small example, the hand calculation uses the Gram matrix already computed:
That gives and residual . The residual is perpendicular to both feature columns. Move anywhere else inside their span and the squared error can only stay the same or grow.
NumPy's least-squares solver checks the arithmetic. It finds the weights that minimize the Euclidean residual norm; that also minimizes its square. Inspect the prediction and residual, not just a single loss number:[7]
1import numpy as np
2
3A = np.array([[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]])
4y = np.array([1.0, 2.0, 2.0])
5
6weights, _, rank, singular_values = np.linalg.lstsq(A, y, rcond=None)
7prediction = A @ weights
8residual = y - prediction
9
10print("weights", weights.round(4).tolist())
11print("prediction", prediction.round(4).tolist())
12print("residual", residual.round(4).tolist())
13print("squared_error", round(float(np.sum(residual ** 2)), 4))
14print("residual_orthogonal", bool(np.allclose(A.T @ residual, 0)))
15print("rank", int(rank), "singular_values", singular_values.round(4).tolist())1weights [0.6667, 1.6667]
2prediction [0.6667, 1.6667, 2.3333]
3residual [0.3333, 0.3333, -0.3333]
4squared_error 0.3333
5residual_orthogonal True
6rank 2 singular_values [1.7321, 1.0]The normal equations make the geometry easy to see, but forming can square a condition number. For a large or ill-conditioned design matrix, use a direct least-squares or SVD-based routine instead of building the Gram matrix yourself. That numerical warning is next.
Condition number warns about solving through weak directions
SVD also reveals a failure mode. If a matrix stretches one direction strongly and almost erases another, a solve that tries to reverse that transformation must amplify whatever landed in the weak direction.
For example, a diagonal matrix with entries 1 and 0.0001 maps [1, 1] to [1, 0.0001]. Reversing it divides the second output by 0.0001. An output error of only 0.0001 therefore adds 1 to the recovered second coordinate.
The ratio between the largest and smallest stretch measures this imbalance. It's the 2-norm condition number:
when the smallest singular value is nonzero. For an invertible square matrix, a condition number near 1 means directions have comparable scales; it doesn't mean the matrix is close to the identity. A large condition number warns about sensitivity of inverse-like operations. If a singular value is zero, a square matrix has no inverse and its condition number is infinite. Whitening (rescaling projected components toward unit variance) also divides by their strengths, so weak directions need care.[8] The next example isolates that sensitivity:
1import numpy as np
2
3A = np.diag([1.0, 0.0001])
4x_true = np.array([1.0, 1.0])
5y = A @ x_true
6y_with_noise = y + np.array([0.0, 0.0001])
7
8x_clean = np.linalg.solve(A, y)
9x_noisy = np.linalg.solve(A, y_with_noise)
10
11print("condition_number", float(np.linalg.cond(A)))
12print("clean_solution", x_clean.tolist())
13print("noisy_solution", x_noisy.tolist())
14print("small_output_noise_doubled_second_coordinate", x_noisy[1] == 2.0)1condition_number 10000.0
2clean_solution [1.0, 1.0]
3noisy_solution [1.0, 2.0]
4small_output_noise_doubled_second_coordinate TrueStoring vectors in a matrix with a weak direction doesn't automatically make cosine or dot-product retrieval unstable. Amplification appears when a pipeline performs an inverse-like operation, such as a solve or whitening transform, along that weak direction.
A matrix has singular values 1 and 0.0001. Why can a solve be sensitive even when the input perturbation is small?
Answer
Its condition number is 10,000. Reversing the weak direction divides by 0.0001, so a small perturbation along that direction can become a large change in the solution.
Forming AᵀA can square the condition number
A.T @ A made our hand calculation easy, but it isn't the safe general implementation of SVD. For a full-column-rank matrix, forming it squares the 2-norm condition number:
A condition number of 10,000 becomes 100,000,000 in the Gram matrix. The full-column-rank assumption matters: a wide matrix can have a finite ratio of singular values while A.T @ A is singular. Weak directions can lose useful precision before the decomposition begins. Use a direct SVD implementation in numerical code; keep A.T @ A as a teaching device for the relationship between eigenvectors and singular vectors.
1import numpy as np
2
3A = np.diag([1.0, 0.0001])
4G = A.T @ A
5print("cond_A", float(np.linalg.cond(A)))
6print("cond_AtA", float(np.linalg.cond(G)))
7print("ratio", float(np.linalg.cond(G) / np.linalg.cond(A)))1cond_A 10000.0
2cond_AtA 100000000.0
3ratio 10000.0Why is np.linalg.svd(A) a better computational default than finding eigenvectors of A.T @ A yourself?
Answer
The Gram-matrix route is excellent for understanding the formula, but it squares the condition number. Direct SVD routines are designed to avoid throwing away as much numerical precision on weak directions.
Choosing a cutoff uses squared singular values
For centered data, squared singular values are proportional to the variation explained by each principal direction. For an uncentered count or embedding matrix, they still measure Frobenius energy captured by each direction. In either case, the cumulative ratio gives a reconstruction-based cutoff, where r is the number of singular values. Predict the smallest k that should pass a 95% target from the spectrum below, then check it:
1import numpy as np
2
3singular_values = np.array([12.4, 8.7, 3.1, 0.9, 0.3])
4fractions = np.cumsum(singular_values ** 2) / np.sum(singular_values ** 2)
5target = 0.95
6k = int(np.searchsorted(fractions, target) + 1)
7
8print("retained_by_component", fractions.round(4).tolist())
9print("smallest_k_for_95_percent", k)
10print("retained_at_k", round(float(fractions[k - 1]), 4))1retained_by_component [0.6408, 0.9562, 0.9962, 0.9996, 1.0]
2smallest_k_for_95_percent 2
3retained_at_k 0.9562This rule answers a reconstruction question. For a retrieval system, follow it with ranking metrics on held-out queries. Matrix energy isn't a proxy for relevance by itself.
Compress a retrieval matrix and check the ranking
Represent four incident tickets with three term features: latency, rollback, and trace. The query is mostly about a latency trace with a smaller rollback signal. We'll project tickets and query onto the top two right-singular directions, then compare rankings before and after compression with cosine similarity.
Cosine similarity divides each dot product by both vector lengths, so it compares direction rather than raw magnitude. Before running the fixture, predict whether dropping only the weakest direction should change the top ticket.
1import numpy as np
2
3tickets = np.array([
4 [3.0, 0.0, 1.0], # latency trace
5 [0.0, 3.0, 1.0], # rollback trace
6 [2.0, 1.0, 0.5], # latency with rollback mention
7 [1.0, 2.0, 1.5], # rollback with latency mention
8])
9query = np.array([2.5, 0.5, 1.0])
10
11def cosine_scores(matrix: np.ndarray, vector: np.ndarray) -> np.ndarray:
12 row_norms = np.linalg.norm(matrix, axis=1)
13 vector_norm = np.linalg.norm(vector)
14 if np.any(row_norms == 0) or vector_norm == 0:
15 raise ValueError("cosine similarity requires non-zero vectors")
16 return (matrix @ vector) / (row_norms * vector_norm)
17
18_, singular_values, Vt = np.linalg.svd(tickets, full_matrices=False)
19basis = Vt[:2].T
20retained_energy = np.sum(singular_values[:2] ** 2) / np.sum(singular_values ** 2)
21
22original_scores = cosine_scores(tickets, query)
23compressed_scores = cosine_scores(tickets @ basis, query @ basis)
24original_order = np.argsort(original_scores)[::-1]
25compressed_order = np.argsort(compressed_scores)[::-1]
26
27print("singular_values", singular_values.round(4).tolist())
28print("retained_energy", round(float(retained_energy), 4))
29print("original_order", original_order.tolist())
30print("compressed_order", compressed_order.tolist())
31print("top_ticket_preserved", bool(original_order[0] == compressed_order[0]))
32print("max_score_change", round(float(np.max(np.abs(original_scores - compressed_scores))), 4))
33
34try:
35 cosine_scores(tickets, np.zeros(3))
36except ValueError as error:
37 print("zero_vector_guard", str(error))1singular_values [4.7011, 3.1677, 0.6044]
2retained_energy 0.9888
3original_order [0, 2, 3, 1]
4compressed_order [0, 2, 3, 1]
5top_ticket_preserved True
6max_score_change 0.0221
7zero_vector_guard cosine similarity requires non-zero vectorsFor this tiny fixture, dropping the tail direction changes scores a little and keeps the same order. The explicit norm guard matters too: cosine similarity is undefined for a zero-length query or ticket vector. A serving pipeline should reject, quarantine, or replace that vector instead of returning NaN scores.
The fixture is a sanity check, not a credible evaluation. On a real retrieval set, measure Recall@k (the fraction of relevant tickets recovered in the top k returned candidates) and inspect failures before accepting reduced dimensions.
Low-rank factors in LoRA and embedding retrieval
| Later task | Matrix idea used here | What must still be validated |
|---|---|---|
| Latent semantic analysis for text retrieval | Truncated SVD maps term-document structure into fewer directions.[9] | Ranking quality on relevant queries |
| Embedding dimensionality reduction | Centered PCA or a task-specific projection drops feature directions | Recall, latency, and storage tradeoffs |
| Linear regression fitting | Least squares projects targets onto the span of feature columns | Held-out error and conditioning |
| Low-rank adaptation (LoRA) | A LoRA update is represented as two thin matrices whose product has limited rank.[10] | Fine-tuned model quality |
| Training diagnostics | Weak or differently scaled directions help explain why optimization can be difficult | Loss curves and held-out performance |
LoRA doesn't compute the SVD of a full trained update during fine-tuning. It learns thin factors directly. SVD gives you the language for "low rank"; later training lessons explain how those factors are optimized.
Practice with measured answers
Use the same fixture to test each idea under a small change. Start with the symmetric matrix, then stress the interpretation rather than the arithmetic:
- Start with the symmetric matrix : compute , its two eigenvalues, singular values, and condition number.
- In the retrieval exercise, keep only one singular direction. Did the first result stay the same? How large did the maximum score change become?
- Add a large constant offset
[100, 100]to every row of the PCA example. Compare SVD before and after centering. Which version describes differences between rows? - Fit
y = [1, 2, 2]with the columns ofAusing least squares. What are the weights, prediction, and residual?
Check your reasoning against these answers:
- , eigenvalues are
16and4, singular values are4and2, and . - In one dimension, nonzero scalars with the same sign have cosine similarity
1. All four tickets and this query project to the same sign, so all four scores are1up to roundoff. The order of those ties isn't evidence of relevance and can vary with numerical precision or sorting behavior. The maximum score change is approximately0.7113. - Uncentered SVD is pulled toward
[1, 1]by the offset. Centering removes it and recovers the contrast direction[1, -1]/sqrt(2), up to sign, just as before. - The weights come out to
[2/3, 5/3], yielding prediction[2/3, 5/3, 7/3]with residual[1/3, 1/3, -1/3]. That residual is orthogonal to both columns ofA.
Our three-ticket matrix leaves one concrete warning: its weaker contrast direction is exactly what distinguishes a latency-only ticket from a rollback-only ticket. SVD tells you how much reconstruction error discarding it costs. Only the downstream task can tell you whether losing that distinction is acceptable.