Skip to main content

MATH 345: Linear Algebra and Optimization

Section 6.5 Singular value decomposition

SectionΒ 6.3 found singular values from the symmetric matrix \(A^TA\text{.}\) SectionΒ 6.4 used the same construction for a centered data matrix \(Z\text{,}\) since
\begin{equation*} C=\frac1nZ^TZ. \end{equation*}
This section packages input directions, output directions, and stretch factors into one factorization.

Subsection Definition and properties

The singular value decomposition is a rectangular analogue of orthogonal diagonalization. Orthogonal diagonalization uses an orthonormal eigenvector basis for a symmetric square matrix. SVD uses an orthonormal basis in the input space and an orthonormal basis in the output space, and it applies to every real matrix.

Definition 6.5.1. Singular value decomposition.

Let \(A\) be an \(m\times n\) matrix, and let \(p=\min(m,n)\text{.}\) A singular value decomposition, or SVD, of \(A\) is a factorization
\begin{equation*} A=U\Sigma V^T, \end{equation*}
where \(U\) is an \(m\times m\) orthogonal matrix, \(V\) is an \(n\times n\) orthogonal matrix, and \(\Sigma\) is an \(m\times n\) diagonal matrix whose diagonal entries
\begin{equation*} \sigma_1\geq \sigma_2\geq \cdots \geq \sigma_p\geq 0 \end{equation*}
are nonnegative.
The numbers \(\sigma_i\) are the singular values of \(A\text{.}\) The columns
\begin{equation*} \mathbf{u}_1,\ldots,\mathbf{u}_m \end{equation*}
of \(U\) are left singular vectors. The columns
\begin{equation*} \mathbf{v}_1,\ldots,\mathbf{v}_n \end{equation*}
of \(V\) are right singular vectors.
If \(r\) is the number of positive singular values, then
\begin{equation*} A\mathbf{v}_i=\sigma_i\mathbf{u}_i \qquad \text{for }1\leq i\leq r, \end{equation*}
and
\begin{equation*} A\mathbf{v}_i=\mathbf{0} \qquad \text{for }i>r. \end{equation*}

Note 6.5.2. Notation for SVD.

Some references write the same factorization as
\begin{equation*} A=P\Sigma Q^T. \end{equation*}
In this book we use
\begin{equation*} A=U\Sigma V^T \end{equation*}
because this matches the NumPy output names U, s, Vt = np.linalg.svd(A). The roles are the same: \(U\) stores left singular vectors, \(V\) stores right singular vectors, and \(\Sigma\) stores singular values. The definition above uses a full SVD; NumPy with full_matrices=False returns the reduced factors needed for the matrix product.

Input directions, stretch, and output directions.

Write
\begin{equation*} U= \begin{bmatrix} \mathbf{u}_1 \amp \cdots \amp \mathbf{u}_m \end{bmatrix}, \qquad V= \begin{bmatrix} \mathbf{v}_1 \amp \cdots \amp \mathbf{v}_n \end{bmatrix}. \end{equation*}
Read
\begin{equation*} A=U\Sigma V^T \end{equation*}
from right to left.
  1. The vector \(V^T\mathbf{x}\) gives the coordinates of \(\mathbf{x}\) in the orthonormal basis of right singular vectors.
  2. The matrix \(\Sigma\) multiplies the \(i\)-th coordinate by \(\sigma_i\text{.}\) A zero singular value erases that direction.
  3. The matrix \(U\) combines the resulting coordinates using the left singular vectors.
By TheoremΒ 4.1.13, if \(r\) singular values are positive, then
\begin{equation*} A\mathbf{x} = \sum_{i=1}^r \sigma_i (\mathbf{v}_i\cdot\mathbf{x}) \mathbf{u}_i. \end{equation*}
The dot product \(\mathbf{v}_i\cdot\mathbf{x}\) measures how much of the input lies in the \(i\)-th right singular direction. The singular value \(\sigma_i\) stretches that coordinate, and \(\mathbf{u}_i\) gives the corresponding output direction.
Since \(U\) and \(V\) are orthogonal, they preserve lengths by LemmaΒ 5.3.3. The changes in length occur in \(\Sigma\text{.}\) Geometrically, the right singular vectors give principal input directions, the singular values give the corresponding stretch factors, and the left singular vectors give principal output directions. The following figure illustrates this picture for a rectangular matrix.
The unit sphere in R3 maps to an ellipse in R2 under a 2 by 3 matrix.
The left side shows a blue unit sphere with coordinate axes labeled \(x_1\text{,}\) \(x_2\text{,}\) and \(x_3\text{.}\) The right side shows a green ellipse in a two-dimensional coordinate plane with axes labeled \(y_1\) and \(y_2\text{.}\) Two black line segments mark the principal semiaxes of the ellipse, ending at \((18,6)\) and \((3,-9)\text{.}\)
Figure 6.5.3. A matrix map from \(\mathbb{R}^3\) to \(\mathbb{R}^2\) corresponding to the matrix
\begin{equation*} A = \begin{bmatrix} 4 \amp 11 \amp 14 \\ 8 \amp 7 \amp -2 \end{bmatrix}\text{.} \end{equation*}
The left panel shows the unit sphere in \(\mathbb{R}^3\text{.}\) The right panel shows its image under \(A\) in \(\mathbb{R}^2\text{,}\) an ellipse with principal semiaxis endpoints \((18,6)\) and \((3,-9)\text{.}\) Adapted from David C. Lay, Steven R. Lay, and Judi J. McDonald, Linear Algebra and Its Applications, 6th edition, Pearson, Β© 2021, Pearson.

Why is this true?.

Since \(\mathbf{v}_i\) is a unit eigenvector,
\begin{align*} \|A\mathbf{v}_i\|^2\\ \amp=\mathbf{v}_i^TA^TA\mathbf{v}_i\\ \amp=\lambda_i. \end{align*}
Thus, when \(\lambda_i>0\text{,}\)
\begin{equation*} \left\| \frac{1}{\sigma_i}A\mathbf{v}_i \right\|=1. \end{equation*}
If \(i\ne j\text{,}\) then
\begin{align*} (A\mathbf{v}_i)\cdot(A\mathbf{v}_j)\\ \amp=\mathbf{v}_i^TA^TA\mathbf{v}_j\\ \amp=\lambda_j(\mathbf{v}_i\cdot\mathbf{v}_j)\\ \amp=0. \end{align*}
Therefore the resulting left singular vectors are orthonormal. If \(\lambda_i=0\text{,}\) then
\begin{equation*} \|A\mathbf{v}_i\|^2=0, \end{equation*}
so \(A\mathbf{v}_i=\mathbf{0}\text{.}\)

Why is this true?.

The matrix \(A^TA\) is symmetric, so TheoremΒ 5.3.11 gives an orthonormal eigenbasis. Its eigenvalues are nonnegative because
\begin{equation*} \mathbf{x}^TA^TA\mathbf{x} = \|A\mathbf{x}\|^2 \geq 0. \end{equation*}
The previous lemma shows that the vectors
\begin{equation*} \mathbf{u}_1,\ldots,\mathbf{u}_r \end{equation*}
are orthonormal and satisfy
\begin{equation*} A\mathbf{v}_i=\sigma_i\mathbf{u}_i. \end{equation*}
Use LemmaΒ 4.2.4 to complete them to an orthonormal basis of \(\R^m\text{.}\)
Comparing columns gives
\begin{equation*} AV = \begin{bmatrix} \sigma_1\mathbf{u}_1 \amp \cdots \amp \sigma_r\mathbf{u}_r \amp \mathbf{0} \amp \cdots \amp \mathbf{0} \end{bmatrix} = U\Sigma. \end{equation*}
Since \(V\) is orthogonal, multiplying by \(V^T\) gives
\begin{equation*} A=U\Sigma V^T. \end{equation*}

Note 6.5.6. Non-uniqueness of the SVD.

The SVD is not unique. A matched pair of singular vectors can be multiplied by \(-1\) without changing the product:
\begin{equation*} \sigma_i\mathbf{u}_i\mathbf{v}_i^T = \sigma_i(-\mathbf{u}_i)(-\mathbf{v}_i)^T. \end{equation*}
Repeated positive singular values also allow orthogonal changes of basis within the corresponding singular subspaces. When a singular value is zero, there can also be many valid choices for the corresponding orthonormal basis vectors. The roles are more important than the specific signs.
Here is one complete SVD computation, broken into the same steps as the theorem.

Activity 6.5.7. Computing an SVD, part 1: right singular vectors (U6-LO6).

Let
\begin{equation*} A = \begin{bmatrix} 1 \amp -1 \\ -2 \amp 2 \\ 2 \amp -2 \end{bmatrix}\text{.} \end{equation*}
  1. Compute \(A^TA\text{.}\)
  2. Find the eigenvalues of \(A^TA\text{.}\)
  3. Find an orthonormal eigenbasis
    \begin{equation*} \mathbf{v}_1,\mathbf{v}_2 \end{equation*}
    for \(A^TA\text{,}\) with the positive eigenvalue first.
  4. Find the singular values of \(A\text{.}\)
Solution.
We compute
\begin{align*} A^TA\\ \amp = \begin{bmatrix} 1 \amp -2 \amp 2\\ -1 \amp 2 \amp -2 \end{bmatrix} \begin{bmatrix} 1 \amp -1\\ -2 \amp 2\\ 2 \amp -2 \end{bmatrix}\\ \amp = \begin{bmatrix} 9 \amp -9\\ -9 \amp 9 \end{bmatrix}. \end{align*}
The characteristic polynomial is
\begin{align*} \det(A^TA-\lambda I)\\ \amp = \det \begin{bmatrix} 9-\lambda \amp -9\\ -9 \amp 9-\lambda \end{bmatrix}\\ \amp = \lambda(\lambda-18). \end{align*}
Thus the eigenvalues are
\begin{equation*} \lambda_1=18, \qquad \lambda_2=0. \end{equation*}
For \(\lambda_1=18\text{,}\) one eigenvector is
\begin{equation*} \begin{bmatrix} 1\\ -1 \end{bmatrix}, \end{equation*}
so we choose the unit vector
\begin{equation*} \mathbf{v}_1= \frac{1}{\sqrt2} \begin{bmatrix} 1\\ -1 \end{bmatrix}. \end{equation*}
For \(\lambda_2=0\text{,}\) one eigenvector is
\begin{equation*} \begin{bmatrix} 1\\ 1 \end{bmatrix}, \end{equation*}
so we choose
\begin{equation*} \mathbf{v}_2= \frac{1}{\sqrt2} \begin{bmatrix} 1\\ 1 \end{bmatrix}. \end{equation*}
The singular values are the square roots of the eigenvalues of \(A^TA\text{:}\)
\begin{equation*} \sigma_1=\sqrt{18}=3\sqrt2, \qquad \sigma_2=0. \end{equation*}

Activity 6.5.8. Computing an SVD, part 2: output directions (U6-LO6, U2-LO3).

Use the vectors from the previous activity.
  1. Compute \(A\mathbf{v}_1\text{.}\)
  2. Compute \(A\mathbf{v}_2\text{.}\)
  3. Compute
    \begin{equation*} \mathbf{u}_1=\frac{1}{\sigma_1}A\mathbf{v}_1. \end{equation*}
  4. What does \(A\mathbf{v}_2=\mathbf{0}\) say about the matrix map?
Solution.
First,
\begin{align*} A\mathbf{v}_1\\ \amp = \begin{bmatrix} 1 \amp -1\\ -2 \amp 2\\ 2 \amp -2 \end{bmatrix} \frac{1}{\sqrt2} \begin{bmatrix} 1\\ -1 \end{bmatrix}\\ \amp = \begin{bmatrix} \sqrt2\\ -2\sqrt2\\ 2\sqrt2 \end{bmatrix}. \end{align*}
Since \(\sigma_1=3\sqrt2\text{,}\)
\begin{align*} \mathbf{u}_1\\ \amp = \frac{1}{3\sqrt2} \begin{bmatrix} \sqrt2\\ -2\sqrt2\\ 2\sqrt2 \end{bmatrix}\\ \amp = \begin{bmatrix} 1/3\\ -2/3\\ 2/3 \end{bmatrix}. \end{align*}
Also,
\begin{align*} A\mathbf{v}_2\\ \amp = \begin{bmatrix} 1 \amp -1\\ -2 \amp 2\\ 2 \amp -2 \end{bmatrix} \frac{1}{\sqrt2} \begin{bmatrix} 1\\ 1 \end{bmatrix}\\ \amp = \begin{bmatrix} 0\\ 0\\ 0 \end{bmatrix}. \end{align*}
Thus \(\mathbf{v}_2\) is a right null direction. The matrix map forgets that input direction. Since only one singular value is positive, the rank of \(A\) is \(1\text{.}\)

Activity 6.5.9. Checking a completed SVD (U6-LO6).

Continue with the matrix and singular vectors from the previous two activities. One orthonormal completion of
\begin{equation*} \mathbf{u}_1= \begin{bmatrix} 1/3\\ -2/3\\ 2/3 \end{bmatrix}. \end{equation*}
gives
\begin{equation*} U= \begin{bmatrix} 1/3 \amp 2\sqrt2/3 \amp 0\\ -2/3 \amp \sqrt2/6 \amp \sqrt2/2\\ 2/3 \amp -\sqrt2/6 \amp \sqrt2/2 \end{bmatrix}. \end{equation*}
Also set
\begin{equation*} \Sigma= \begin{bmatrix} 3\sqrt2 \amp 0\\ 0 \amp 0\\ 0 \amp 0 \end{bmatrix}, \qquad V= \frac1{\sqrt2} \begin{bmatrix} 1 \amp 1\\ -1 \amp 1 \end{bmatrix}. \end{equation*}
  1. Identify the shapes of \(U\text{,}\) \(\Sigma\text{,}\) and \(V\text{.}\)
  2. Check that
    \begin{equation*} U^TU=I_3, \qquad V^TV=I_2. \end{equation*}
  3. Use
    \begin{equation*} A\mathbf{v}_1=3\sqrt2\,\mathbf{u}_1, \qquad A\mathbf{v}_2=\mathbf{0} \end{equation*}
    to explain why
    \begin{equation*} AV=U\Sigma \end{equation*}
    without multiplying all the matrices entry by entry.
  4. Conclude that
    \begin{equation*} A=U\Sigma V^T. \end{equation*}
  5. Which column of \(U\) comes from a positive singular value? Which columns arise from an orthonormal completion?
Solution.
The shapes are
\begin{equation*} U\in\R^{3\times3}, \qquad \Sigma\in\R^{3\times2}, \qquad V\in\R^{2\times2}. \end{equation*}
Direct multiplication gives
\begin{equation*} U^TU=I_3, \qquad V^TV=I_2, \end{equation*}
so \(U\) and \(V\) are orthogonal.
The columns of \(AV\) are
\begin{equation*} A\mathbf{v}_1 \qquad\text{and}\qquad A\mathbf{v}_2. \end{equation*}
From the previous activities,
\begin{equation*} AV = \begin{bmatrix} 3\sqrt2\,\mathbf{u}_1 \amp \mathbf{0} \end{bmatrix} = U\Sigma. \end{equation*}
Since \(V\) is orthogonal,
\begin{equation*} V^{-1}=V^T. \end{equation*}
Multiplying \(AV=U\Sigma\) on the right by \(V^T\) gives
\begin{equation*} A=U\Sigma V^T. \end{equation*}
The first column \(\mathbf{u}_1\) comes from the positive singular value. The remaining columns complete \(\mathbf{u}_1\) to an orthonormal basis of \(\R^3\text{.}\) That completion is not unique.
Unit 2 introduced row spaces, column spaces, and null spaces in DefinitionΒ 2.4.33 and DefinitionΒ 2.3.13. The null space of \(A^T\) is called the left null space of \(A\text{.}\) Together, these are the four fundamental subspaces of a matrix.

Definition 6.5.10. Four fundamental subspaces.

Let \(A\) be an \(m\times n\) matrix. Its four fundamental subspaces are
\begin{equation*} \operatorname{row}(A)\subseteq\R^n, \qquad \operatorname{col}(A)\subseteq\R^m, \end{equation*}
\begin{equation*} \operatorname{null}(A)\subseteq\R^n, \qquad \operatorname{null}(A^T)\subseteq\R^m. \end{equation*}
More explicitly,
\begin{equation*} \operatorname{row}(A) = \spans\{\text{rows of }A\}, \end{equation*}
\begin{equation*} \operatorname{col}(A) = \spans\{\text{columns of }A\}, \end{equation*}
\begin{equation*} \operatorname{null}(A) = \{\mathbf{x}\in\R^n:A\mathbf{x}=\mathbf{0}\}, \end{equation*}
and
\begin{equation*} \operatorname{null}(A^T) = \{\mathbf{y}\in\R^m:A^T\mathbf{y}=\mathbf{0}\}. \end{equation*}

Why is this true?.

\begin{equation*} A\mathbf{x} = \sum_{i=1}^r \sigma_i (\mathbf{v}_i\cdot\mathbf{x}) \mathbf{u}_i. \end{equation*}
Thus every output lies in
\begin{equation*} \spans\{\mathbf{u}_1,\ldots,\mathbf{u}_r\}. \end{equation*}
Conversely,
\begin{equation*} \mathbf{u}_i = \frac1{\sigma_i}A\mathbf{v}_i \end{equation*}
belongs to \(\operatorname{col}(A)\) for \(1\leq i\leq r\text{.}\) Therefore the first \(r\) left singular vectors form a basis for the column space.
Write an input in the right singular basis:
\begin{equation*} \mathbf{x} = c_1\mathbf{v}_1+\cdots+c_n\mathbf{v}_n. \end{equation*}
Then
\begin{equation*} A\mathbf{x} = \sigma_1c_1\mathbf{u}_1+\cdots+ \sigma_rc_r\mathbf{u}_r. \end{equation*}
Since the displayed left singular vectors are independent, \(A\mathbf{x}=\mathbf{0}\) exactly when
\begin{equation*} c_1=\cdots=c_r=0. \end{equation*}
Hence the remaining right singular vectors form a basis for \(\operatorname{null}(A)\text{.}\)
Apply the same argument to
\begin{equation*} A^T=V\Sigma^TU^T. \end{equation*}
This gives the row space and left null space statements. Finally, the column space has \(r\) basis vectors, so
\begin{equation*} \operatorname{rank}(A)=r. \end{equation*}

Activity 6.5.12. Reading the computed SVD (U6-LO6, U6-LO7).

Use the SVD computed above:
\begin{equation*} A=U\Sigma V^T, \end{equation*}
where
\begin{equation*} V= \begin{bmatrix} 1/\sqrt2 \amp 1/\sqrt2\\ -1/\sqrt2 \amp 1/\sqrt2 \end{bmatrix}, \end{equation*}
\begin{equation*} U= \begin{bmatrix} 1/3 \amp (2/3)\sqrt2 \amp 0\\ -2/3 \amp (1/6)\sqrt2 \amp (1/2)\sqrt2\\ 2/3 \amp -(1/6)\sqrt2 \amp (1/2)\sqrt2 \end{bmatrix}, \end{equation*}
and
\begin{equation*} \Sigma= \begin{bmatrix} 3\sqrt2 \amp 0\\ 0 \amp 0\\ 0 \amp 0 \end{bmatrix}. \end{equation*}
  1. What is \(\operatorname{rank}(A)\text{?}\)
  2. Give an orthonormal basis for \(\operatorname{row}(A)\text{.}\)
  3. Give an orthonormal basis for \(\operatorname{null}(A)\text{.}\)
  4. Give an orthonormal basis for \(\operatorname{col}(A)\text{.}\)
  5. Give an orthonormal basis for \(\operatorname{null}(A^T)\text{.}\)
  6. Which input direction is forgotten?
  7. Which input direction is transmitted, and by what stretch factor?
Solution.
There is one positive singular value, so
\begin{equation*} \operatorname{rank}(A)=1. \end{equation*}
The first right singular vector gives a basis for the row space:
\begin{equation*} \operatorname{row}(A) = \operatorname{span} \left\{ \frac{1}{\sqrt2} \begin{bmatrix} 1\\ -1 \end{bmatrix} \right\}. \end{equation*}
The zero singular value gives the null-space direction:
\begin{equation*} \operatorname{null}(A) = \operatorname{span} \left\{ \frac{1}{\sqrt2} \begin{bmatrix} 1\\ 1 \end{bmatrix} \right\}. \end{equation*}
The first left singular vector gives a basis for the column space:
\begin{equation*} \operatorname{col}(A) = \operatorname{span} \left\{ \begin{bmatrix} 1/3\\ -2/3\\ 2/3 \end{bmatrix} \right\}. \end{equation*}
The remaining left singular vectors give a basis for \(\operatorname{null}(A^T)\text{:}\)
\begin{equation*} \operatorname{null}(A^T) = \operatorname{span} \left\{ \begin{bmatrix} (2/3)\sqrt2\\ (1/6)\sqrt2\\ -(1/6)\sqrt2 \end{bmatrix}, \begin{bmatrix} 0\\ (1/2)\sqrt2\\ (1/2)\sqrt2 \end{bmatrix} \right\}. \end{equation*}
The forgotten input direction is
\begin{equation*} \frac{1}{\sqrt2} \begin{bmatrix} 1\\ 1 \end{bmatrix}. \end{equation*}
The transmitted input direction is
\begin{equation*} \frac{1}{\sqrt2} \begin{bmatrix} 1\\ -1 \end{bmatrix}, \end{equation*}
and its stretch factor is
\begin{equation*} 3\sqrt2. \end{equation*}

Note 6.5.13. SVD and orthogonal diagonalization.

The SVD also gives orthogonal diagonalizations of two symmetric matrices:
\begin{equation*} A^TA = V\Sigma^T\Sigma V^T \end{equation*}
and
\begin{equation*} AA^T = U\Sigma\Sigma^TU^T. \end{equation*}
Thus the right singular vectors are eigenvectors of \(A^TA\text{,}\) the left singular vectors are eigenvectors of \(AA^T\text{,}\) and the positive eigenvalues of both symmetric matrices are
\begin{equation*} \sigma_1^2,\ldots,\sigma_r^2. \end{equation*}
Orthogonal diagonalization writes a symmetric square matrix as
\begin{equation*} B=QDQ^T, \end{equation*}
using the same orthonormal basis in its domain and codomain; see DefinitionΒ 5.3.5 and TheoremΒ 5.3.11. The diagonal entries of \(D\) are signed eigenvalues.
The SVD
\begin{equation*} A=U\Sigma V^T \end{equation*}
uses one orthonormal basis in the input space and another in the output space. Its diagonal entries are nonnegative stretch factors. This two-basis version applies to every real matrix, including rectangular and nonsymmetric matrices.
It gives a second description of the matrix map
\begin{equation*} \mathbf{x}\longmapsto A\mathbf{x}: \end{equation*}
measure the input in special directions, stretch or erase those coordinates, and express the result in special output directions. This is a precise version of the Unit 2 question: what information does a matrix map keep, weaken, or forget?

Subsection PCA from the SVD

SectionΒ 6.4 found principal directions by diagonalizing the covariance matrix
\begin{equation*} C=\frac1nZ^TZ. \end{equation*}
The SVD of the centered data matrix \(Z\) contains the same directions and the same captured-variation values. It also connects principal-component scores with left singular vectors.

Why is this true?.

Since \(U\) is orthogonal,
\begin{equation*} U^TU=I. \end{equation*}
Therefore
\begin{align*} Z^TZ\\ \amp= (U\Sigma V^T)^T(U\Sigma V^T)\\ \amp= V\Sigma^TU^TU\Sigma V^T\\ \amp= V\Sigma^T\Sigma V^T. \end{align*}
Dividing by \(n\) gives the displayed orthogonal diagonalization of \(C\text{.}\) The columns of \(V\) are therefore covariance eigenvectors, with eigenvalues \(\sigma_j^2/n\text{.}\)
Also,
\begin{align*} ZV\\ \amp= U\Sigma V^TV\\ \amp= U\Sigma. \end{align*}
Thus the columns of \(U\Sigma\) are the principal-component score vectors.

Keeping the first \(k\) principal directions.

Let \(p=\min(n,d)\) and \(1\leq k\leq p\text{.}\) Set
\begin{equation*} V_k= \begin{bmatrix} \mathbf{v}_1 \amp \cdots \amp \mathbf{v}_k \end{bmatrix}, \qquad T_k=ZV_k, \end{equation*}
where \(V_k\in\R^{d\times k}\) and \(T_k\in\R^{n\times k}\text{.}\)
The centered-data reconstruction using the first \(k\) principal directions is
\begin{equation*} \widehat Z_k = T_kV_k^T = ZV_kV_k^T. \end{equation*}
Since the columns of \(V_k\) are orthonormal, each row of \(\widehat Z_k\) is the orthogonal projection of the corresponding centered data vector onto
\begin{equation*} \spans \{\mathbf{v}_1,\ldots,\mathbf{v}_k\}; \end{equation*}
If \(U_k\in\R^{n\times k}\) contains the first \(k\) left singular vectors and
\begin{equation*} \Sigma_k = \operatorname{diag}(\sigma_1,\ldots,\sigma_k) \in\R^{k\times k}, \end{equation*}
then
\begin{equation*} \widehat Z_k = U_k\Sigma_kV_k^T. \end{equation*}
The reconstruction has rank at most \(k\text{.}\) Each \(d\)-dimensional centered data point is represented by its \(k\) scores and reconstructed from those scores.
The fraction of total centered variation captured is
\begin{equation*} \frac{\lambda_1+\cdots+\lambda_k} {\lambda_1+\cdots+\lambda_d} = \frac{\sigma_1^2+\cdots+\sigma_k^2} {\sigma_1^2+\cdots+\sigma_p^2}, \end{equation*}
provided the centered data are not all zero.

Activity 6.5.15. The small PCA data set through the SVD (U6-LO5, U6-LO6, U6-LO7, U4-LO3).

\begin{equation*} Z= \begin{bmatrix} 2 \amp 1\\ 1 \amp 2\\ -2 \amp -1\\ -1 \amp -2 \end{bmatrix}, \qquad V= \frac1{\sqrt2} \begin{bmatrix} 1 \amp 1\\ 1 \amp -1 \end{bmatrix}, \end{equation*}
and the covariance eigenvalues are
\begin{equation*} \lambda_1=\frac92, \qquad \lambda_2=\frac12. \end{equation*}
  1. Use
    \begin{equation*} \lambda_j=\frac{\sigma_j^2}{4} \end{equation*}
    to find the singular values of \(Z\text{.}\)
  2. Compute the principal-coordinate matrix
    \begin{equation*} T=ZV. \end{equation*}
  3. Let \(\mathbf{t}_1,\mathbf{t}_2\) be the columns of \(T\text{.}\) Compute
    \begin{equation*} \mathbf{u}_1=\frac{\mathbf{t}_1}{\sigma_1}, \qquad \mathbf{u}_2=\frac{\mathbf{t}_2}{\sigma_2}, \end{equation*}
    and verify that \(\mathbf{u}_1,\mathbf{u}_2\) are orthonormal.
  4. Set
    \begin{equation*} U_2= \begin{bmatrix} \mathbf{u}_1 \amp \mathbf{u}_2 \end{bmatrix}, \qquad \Sigma_2= \begin{bmatrix} \sigma_1 \amp 0\\ 0 \amp \sigma_2 \end{bmatrix}. \end{equation*}
    Explain why
    \begin{equation*} Z=U_2\Sigma_2V^T. \end{equation*}
  5. Compute the reconstruction using only the first principal direction:
    \begin{equation*} \widehat Z_1 = \mathbf{t}_1\mathbf{v}_1^T. \end{equation*}
  6. Interpret the rows of \(\widehat Z_1\text{.}\) How does its first row compare with the projection computed in the earlier PCA activity?
  7. Use the singular values to compute the fraction of total centered variation captured by the first principal direction.
Solution.
Since \(n=4\text{,}\)
\begin{equation*} \sigma_1 = \sqrt{4\lambda_1} = \sqrt{18} = 3\sqrt2, \qquad \sigma_2 = \sqrt{4\lambda_2} = \sqrt2. \end{equation*}
We compute
\begin{equation*} T=ZV = \frac1{\sqrt2} \begin{bmatrix} 3 \amp 1\\ 3 \amp -1\\ -3 \amp -1\\ -3 \amp 1 \end{bmatrix}. \end{equation*}
Thus
\begin{equation*} \mathbf{t}_1 = \frac1{\sqrt2} \begin{bmatrix} 3\\ 3\\ -3\\ -3 \end{bmatrix}, \qquad \mathbf{t}_2 = \frac1{\sqrt2} \begin{bmatrix} 1\\ -1\\ -1\\ 1 \end{bmatrix}. \end{equation*}
Dividing by the singular values gives
\begin{equation*} \mathbf{u}_1 = \frac12 \begin{bmatrix} 1\\ 1\\ -1\\ -1 \end{bmatrix}, \qquad \mathbf{u}_2 = \frac12 \begin{bmatrix} 1\\ -1\\ -1\\ 1 \end{bmatrix}. \end{equation*}
Both vectors have norm \(1\text{,}\) and
\begin{equation*} \mathbf{u}_1\cdot\mathbf{u}_2=0. \end{equation*}
Therefore they are orthonormal.
We have
\begin{equation*} T=U_2\Sigma_2. \end{equation*}
Since
\begin{equation*} T=ZV \end{equation*}
and \(V\) is orthogonal,
\begin{equation*} Z = TV^T = U_2\Sigma_2V^T. \end{equation*}
This factorization uses the two left singular vectors needed to reconstruct \(Z\text{.}\)
Keeping only the first principal direction gives
\begin{equation*} \widehat Z_1 = \mathbf{t}_1\mathbf{v}_1^T = \begin{bmatrix} 3/2 \amp 3/2\\ 3/2 \amp 3/2\\ -3/2 \amp -3/2\\ -3/2 \amp -3/2 \end{bmatrix}. \end{equation*}
Each row is the projection of the corresponding centered data vector onto
\begin{equation*} \spans\{\mathbf{v}_1\}. \end{equation*}
In particular, the first row is
\begin{equation*} \begin{bmatrix} 3/2 \amp 3/2 \end{bmatrix}, \end{equation*}
which is the transpose of the projected centered vector
\begin{equation*} \widehat{\mathbf{z}}_1 = \begin{bmatrix} 3/2\\ 3/2 \end{bmatrix} \end{equation*}
found in the earlier PCA activity.
Finally,
\begin{equation*} \frac{\sigma_1^2} {\sigma_1^2+\sigma_2^2} = \frac{18}{18+2} = \frac9{10}. \end{equation*}
Thus the first principal direction captures \(90\%\) of the total centered variation.