Section C.4 Unit 4 applied, geometric, and computational interpretation
These solution sketches correspond to the additional applied and computational interpretation exercises at the end of Unit 4.
Subsection Orthogonal coordinates and projection
Solution to Orthogonal coordinates and projection.
The columns have length \(1\text{,}\) and
\begin{equation*}
\mathbf{q}_1^T\mathbf{q}_2
=
\frac12(1-1)
=
0.
\end{equation*}
Therefore
\begin{equation*}
Q^TQ=I_2.
\end{equation*}
The coordinate vector is
\begin{equation*}
\mathbf{c}
=
Q^T\mathbf{x}
=
\begin{bmatrix}
\mathbf{q}_1^T\mathbf{x}\\
\mathbf{q}_2^T\mathbf{x}
\end{bmatrix}
=
\begin{bmatrix}
2\sqrt{2}\\
\sqrt{2}
\end{bmatrix}.
\end{equation*}
Combining the columns of \(Q\) gives
\begin{equation*}
\widehat{\mathbf{x}}
=
Q\mathbf{c}
=
2\sqrt{2}\mathbf{q}_1
+
\sqrt{2}\mathbf{q}_2
=
\begin{bmatrix}
3\\
1\\
0
\end{bmatrix}.
\end{equation*}
The residual is
\begin{equation*}
\mathbf{r}
=
\mathbf{x}-\widehat{\mathbf{x}}
=
\begin{bmatrix}
0\\
0\\
2
\end{bmatrix}.
\end{equation*}
\begin{equation*}
Q^T\mathbf{r}
=
\begin{bmatrix}
0\\
0
\end{bmatrix}.
\end{equation*}
The multiplication \(Q^T\mathbf{x}\) scores the orthonormal directions. The multiplication \(Q\mathbf{c}\) combines those directions using the resulting coordinates. The vector \(\mathbf{c}\) contains coordinates. The projected vector is
\begin{equation*}
\widehat{\mathbf{x}}
=
Q\mathbf{c}.
\end{equation*}
Subsection One GramβSchmidt column step
Solution to One GramβSchmidt column step.
The projection coefficient is
\begin{equation*}
r_{12}
=
\mathbf{q}_1^T\mathbf{a}_2
=
\frac{2}{\sqrt{2}}
=
\sqrt{2}.
\end{equation*}
Therefore
\begin{equation*}
\mathbf{p}_2
=
r_{12}\mathbf{q}_1
=
\sqrt{2}
\frac{1}{\sqrt{2}}
\begin{bmatrix}
1\\
1\\
0
\end{bmatrix}
=
\begin{bmatrix}
1\\
1\\
0
\end{bmatrix}.
\end{equation*}
The residual is
\begin{equation*}
\mathbf{v}_2
=
\mathbf{a}_2-\mathbf{p}_2
=
\begin{bmatrix}
1\\
-1\\
1
\end{bmatrix}.
\end{equation*}
Its length is
\begin{equation*}
r_{22}
=
\|\mathbf{v}_2\|
=
\sqrt{3}.
\end{equation*}
Thus
\begin{equation*}
\mathbf{q}_2
=
\frac{1}{\sqrt{3}}
\begin{bmatrix}
1\\
-1\\
1
\end{bmatrix}.
\end{equation*}
The orthogonality check gives
\begin{equation*}
\mathbf{q}_1^T\mathbf{q}_2
=
\frac{1}{\sqrt{6}}
(1-1+0)
=
0.
\end{equation*}
If \(r_{22}=0\text{,}\) then \(\mathbf{v}_2=\mathbf{0}\text{.}\) The column \(\mathbf{a}_2\) would already be in the span of the earlier directions and would not produce a new unit direction.
Subsection Orthogonal-basis coefficients
The vectors are pairwise orthogonal. The coefficients are
\begin{align*}
\frac{\mathbf{x} \cdot \mathbf{f}_1}{\mathbf{f}_1 \cdot \mathbf{f}_1} \amp= 2,\\
\frac{\mathbf{x} \cdot \mathbf{f}_2}{\mathbf{f}_2 \cdot \mathbf{f}_2} \amp= 1,\\
\frac{\mathbf{x} \cdot \mathbf{f}_3}{\mathbf{f}_3 \cdot \mathbf{f}_3} \amp= 2.
\end{align*}
So \(\mathbf{x} = 2\mathbf{f}_1 + \mathbf{f}_2 + 2\mathbf{f}_3\text{.}\) Orthogonality makes the cross terms vanish.
Subsection Projection onto a line
\(\mathbf{x} \cdot \mathbf{u} = 5\) and \(\mathbf{u} \cdot \mathbf{u} = 5\text{,}\) so \(\operatorname{proj}_L(\mathbf{x}) = \mathbf{u} = \begin{bmatrix}1\\2\end{bmatrix}\text{.}\) The residual is \(\begin{bmatrix}2\\-1\end{bmatrix}\text{.}\) Its dot product with \(\mathbf{u}\) is \(2 - 2 = 0\text{.}\) The residual is the part of \(\mathbf{x}\) perpendicular to the line.
Subsection Column-space orthogonal complement
\(A^T\mathbf{r} = \begin{bmatrix}0\\0\end{bmatrix}\text{.}\) Therefore \(\mathbf{r}\) is orthogonal to each column of \(A\text{.}\) The orthogonal complement is \(\operatorname{span}\{\begin{bmatrix}-1\\-1\\1\end{bmatrix}\}\text{.}\) This is exactly the null space of \(A^T\text{.}\)
Subsection Projection onto a plane
A normal vector is \(\mathbf{n} = \begin{bmatrix}1\\1\\1\end{bmatrix}\text{.}\) Since \(\mathbf{p} \cdot \mathbf{n} = 7\) and \(\mathbf{n} \cdot \mathbf{n} = 3\text{,}\)
\begin{equation*}
\operatorname{proj}_U(\mathbf{p}) = \mathbf{p} - (7/3)\mathbf{n}
=
\begin{bmatrix}-4/3\\-1/3\\5/3\end{bmatrix}.
\end{equation*}
The entries add to 0, so the projected point lies in \(U\text{.}\)
Subsection Unreachable target, closest output
The output of \(A\) has the form \((x_1, x_2, x_1+x_2)\text{.}\) The first two target entries force \(x_1 = 2\) and \(x_2 = 1\text{,}\) which would make the third entry \(3\text{,}\) not \(5\text{.}\)
\(A^T A = \begin{bmatrix}2\amp 1\\1\amp 2\end{bmatrix}\text{,}\) \(A^T\mathbf{b} = \begin{bmatrix}7\\6\end{bmatrix}\text{.}\)
Solving gives \(\hat{\mathbf{x}} = \begin{bmatrix}8/3\\5/3\end{bmatrix}\text{.}\) Then
\begin{equation*}
A\hat{\mathbf{x}} =
\begin{bmatrix}8/3\\5/3\\13/3\end{bmatrix},
\end{equation*}
\begin{equation*}
\mathbf{r} =
\begin{bmatrix}-2/3\\-2/3\\2/3\end{bmatrix},
\end{equation*}
and \(A^T\mathbf{r} = \begin{bmatrix}0\\0\end{bmatrix}\text{.}\)
Subsection Regression design matrix and residual
\(X = \begin{bmatrix}1\amp 0\\1\amp 1\\1\amp 2\end{bmatrix}\text{,}\) \(\mathbf{y} = \begin{bmatrix}1\\2\\2\end{bmatrix}\text{.}\)
\(X^T X = \begin{bmatrix}3\amp 3\\3\amp 5\end{bmatrix}\text{,}\) \(X^T\mathbf{y} = \begin{bmatrix}5\\6\end{bmatrix}\text{.}\)
Solving gives \(\boldsymbol{\beta} = \begin{bmatrix}7/6\\1/2\end{bmatrix}\text{.}\) The fitted values are \(\begin{bmatrix}7/6\\5/3\\13/6\end{bmatrix}\text{.}\) The residual is \(\begin{bmatrix}-1/6\\1/3\\-1/6\end{bmatrix}\text{,}\) and \(X^T\mathbf{r} = \begin{bmatrix}0\\0\end{bmatrix}\text{.}\)
Subsection Code interpretation: residual orthogonality
The coefficient vector is
\begin{equation*}
\widehat{\mathbf{x}}
=
\begin{bmatrix}
8/3\\
5/3
\end{bmatrix}.
\end{equation*}
The fitted vector is
\begin{equation*}
\widehat{\mathbf{b}}
=
A\widehat{\mathbf{x}}
=
\begin{bmatrix}
8/3\\
5/3\\
13/3
\end{bmatrix},
\end{equation*}
and the residual is
\begin{equation*}
\mathbf{r}
=
\mathbf{b}-\widehat{\mathbf{b}}
=
\begin{bmatrix}
-2/3\\
-2/3\\
2/3
\end{bmatrix}.
\end{equation*}
The expression
A.T @ r checks
\begin{equation*}
A^T\mathbf{r}=\mathbf{0},
\end{equation*}
so it checks that the residual is orthogonal to every column of \(A\text{.}\) The tiny entries come from floating-point roundoff.
The residual is not zero:
\begin{equation*}
\|\mathbf{r}\|
=
\frac{2}{\sqrt{3}}
\approx
1.15470.
\end{equation*}
This does not mean least squares failed. The target is not reachable, so the closest reachable output has a nonzero residual.
Subsection Debug: wrong residual check
The matrix \(A\) has shape \(m\times n\text{,}\) the transpose \(A^T\) has shape \(n\times m\text{,}\) and \(\mathbf{r}\) has length \(m\text{.}\) Thus
\begin{equation*}
A^T\mathbf{r}
\end{equation*}
is defined and has length \(n\text{.}\) The product \(A\mathbf{r}\) is usually not defined because \(A\) expects an input of length \(n\text{,}\) not \(m\text{.}\)
Even if the dimensions happened to match, \(A\mathbf{r}\) would form a linear combination of the columns of \(A\text{.}\) It would not collect the dot products of \(\mathbf{r}\) with those columns.
A numerical Boolean check is
np.allclose(A.T @ r, np.zeros(A.shape[1]))
Subsection Rank and nonunique coefficients
\(A\mathbf{z} = \mathbf{0}\text{,}\) so \(\mathbf{z}\) is a null-space direction. Therefore
\begin{equation*}
A(\hat{\mathbf{x}} + 5\mathbf{z}) = A\hat{\mathbf{x}} + 5A\mathbf{z} = A\hat{\mathbf{x}}.
\end{equation*}
The coefficient vector can change while the fitted vector stays fixed. This happens because the third column is the sum of the first two columns.
Subsection \(QR\) least-squares code reading
Since \(A\) has shape \(3\times2\text{,}\) reduced \(QR\) gives
\begin{equation*}
Q:3\times2,
\qquad
R:2\times2.
\end{equation*}
The first Boolean check verifies
\begin{equation*}
Q^TQ=I_2,
\end{equation*}
and the second verifies
\begin{equation*}
QR=A.
\end{equation*}
The solve computes the reduced system
\begin{equation*}
R\widehat{\mathbf{x}}
=
Q^T\mathbf{b}.
\end{equation*}
Both methods return
\begin{equation*}
\widehat{\mathbf{x}}
=
\begin{bmatrix}
8/3\\
5/3
\end{bmatrix}
\end{equation*}
up to roundoff, so the final comparison is
True.The agreement does not mean that \(A\mathbf{x}=\mathbf{b}\) has an exact solution. The residual is nonzero.
A \(QR\) factorization permits a simultaneous sign change in one column of \(Q\) and the corresponding row of \(R\text{.}\) Numerical software may therefore return signs different from a hand calculation while still satisfying
\begin{equation*}
Q^TQ=I
\qquad\text{and}\qquad
QR=A.
\end{equation*}
Subsection Projection and attention: score, then combine
For the projection pipeline,
\begin{equation*}
\mathbf{c}
=
Q^T\mathbf{x}
=
\begin{bmatrix}
2\\
1
\end{bmatrix}.
\end{equation*}
Therefore
\begin{equation*}
\widehat{\mathbf{x}}
=
Q\mathbf{c}
=
\begin{bmatrix}
2\\
1\\
0
\end{bmatrix},
\qquad
\mathbf{r}
=
\begin{bmatrix}
0\\
0\\
3
\end{bmatrix}.
\end{equation*}
The residual check gives
\begin{equation*}
Q^T\mathbf{r}
=
\begin{bmatrix}
0\\
0
\end{bmatrix}.
\end{equation*}
For the attention pipeline,
\begin{equation*}
\mathbf{s}
=
K\mathbf{q}
=
\begin{bmatrix}
2\\
1
\end{bmatrix},
\qquad
\boldsymbol{\alpha}
=
\begin{bmatrix}
2/3\\
1/3
\end{bmatrix}.
\end{equation*}
The output is
\begin{equation*}
\mathbf{h}
=
\boldsymbol{\alpha}^TV
=
\frac23
\begin{bmatrix}
3\amp 0
\end{bmatrix}
+
\frac13
\begin{bmatrix}
0\amp 6
\end{bmatrix}
=
\begin{bmatrix}
2\amp 2
\end{bmatrix}.
\end{equation*}
Both pipelines use dot products or matrix products to produce scores and use linear combinations to combine vectors.
Projection scores the orthonormal directions in \(Q\) and combines those same directions. Its residual is orthogonal to their span. Attention scores key vectors and combines value vectors, which may be different vectors.
