Skip to main content

MATH 345: Linear Algebra and Optimization

Section C.4 Unit 4 applied, geometric, and computational interpretation

These solution sketches correspond to the additional applied and computational interpretation exercises at the end of Unit 4.

Subsection Orthogonal coordinates and projection

The columns have length \(1\text{,}\) and
\begin{equation*} \mathbf{q}_1^T\mathbf{q}_2 = \frac12(1-1) = 0. \end{equation*}
Therefore
\begin{equation*} Q^TQ=I_2. \end{equation*}
The coordinate vector is
\begin{equation*} \mathbf{c} = Q^T\mathbf{x} = \begin{bmatrix} \mathbf{q}_1^T\mathbf{x}\\ \mathbf{q}_2^T\mathbf{x} \end{bmatrix} = \begin{bmatrix} 2\sqrt{2}\\ \sqrt{2} \end{bmatrix}. \end{equation*}
Combining the columns of \(Q\) gives
\begin{equation*} \widehat{\mathbf{x}} = Q\mathbf{c} = 2\sqrt{2}\mathbf{q}_1 + \sqrt{2}\mathbf{q}_2 = \begin{bmatrix} 3\\ 1\\ 0 \end{bmatrix}. \end{equation*}
The residual is
\begin{equation*} \mathbf{r} = \mathbf{x}-\widehat{\mathbf{x}} = \begin{bmatrix} 0\\ 0\\ 2 \end{bmatrix}. \end{equation*}
Since both columns of \(Q\) have third coordinate \(0\text{,}\)
\begin{equation*} Q^T\mathbf{r} = \begin{bmatrix} 0\\ 0 \end{bmatrix}. \end{equation*}
The multiplication \(Q^T\mathbf{x}\) scores the orthonormal directions. The multiplication \(Q\mathbf{c}\) combines those directions using the resulting coordinates. The vector \(\mathbf{c}\) contains coordinates. The projected vector is
\begin{equation*} \widehat{\mathbf{x}} = Q\mathbf{c}. \end{equation*}

Subsection One Gram–Schmidt column step

The projection coefficient is
\begin{equation*} r_{12} = \mathbf{q}_1^T\mathbf{a}_2 = \frac{2}{\sqrt{2}} = \sqrt{2}. \end{equation*}
Therefore
\begin{equation*} \mathbf{p}_2 = r_{12}\mathbf{q}_1 = \sqrt{2} \frac{1}{\sqrt{2}} \begin{bmatrix} 1\\ 1\\ 0 \end{bmatrix} = \begin{bmatrix} 1\\ 1\\ 0 \end{bmatrix}. \end{equation*}
The residual is
\begin{equation*} \mathbf{v}_2 = \mathbf{a}_2-\mathbf{p}_2 = \begin{bmatrix} 1\\ -1\\ 1 \end{bmatrix}. \end{equation*}
Its length is
\begin{equation*} r_{22} = \|\mathbf{v}_2\| = \sqrt{3}. \end{equation*}
\begin{equation*} \mathbf{q}_2 = \frac{1}{\sqrt{3}} \begin{bmatrix} 1\\ -1\\ 1 \end{bmatrix}. \end{equation*}
The orthogonality check gives
\begin{equation*} \mathbf{q}_1^T\mathbf{q}_2 = \frac{1}{\sqrt{6}} (1-1+0) = 0. \end{equation*}
If \(r_{22}=0\text{,}\) then \(\mathbf{v}_2=\mathbf{0}\text{.}\) The column \(\mathbf{a}_2\) would already be in the span of the earlier directions and would not produce a new unit direction.

Subsection Orthogonal-basis coefficients

The vectors are pairwise orthogonal. The coefficients are
\begin{align*} \frac{\mathbf{x} \cdot \mathbf{f}_1}{\mathbf{f}_1 \cdot \mathbf{f}_1} \amp= 2,\\ \frac{\mathbf{x} \cdot \mathbf{f}_2}{\mathbf{f}_2 \cdot \mathbf{f}_2} \amp= 1,\\ \frac{\mathbf{x} \cdot \mathbf{f}_3}{\mathbf{f}_3 \cdot \mathbf{f}_3} \amp= 2. \end{align*}
So \(\mathbf{x} = 2\mathbf{f}_1 + \mathbf{f}_2 + 2\mathbf{f}_3\text{.}\) Orthogonality makes the cross terms vanish.

Subsection Projection onto a line

\(\mathbf{x} \cdot \mathbf{u} = 5\) and \(\mathbf{u} \cdot \mathbf{u} = 5\text{,}\) so \(\operatorname{proj}_L(\mathbf{x}) = \mathbf{u} = \begin{bmatrix}1\\2\end{bmatrix}\text{.}\) The residual is \(\begin{bmatrix}2\\-1\end{bmatrix}\text{.}\) Its dot product with \(\mathbf{u}\) is \(2 - 2 = 0\text{.}\) The residual is the part of \(\mathbf{x}\) perpendicular to the line.

Subsection Column-space orthogonal complement

\(A^T\mathbf{r} = \begin{bmatrix}0\\0\end{bmatrix}\text{.}\) Therefore \(\mathbf{r}\) is orthogonal to each column of \(A\text{.}\) The orthogonal complement is \(\operatorname{span}\{\begin{bmatrix}-1\\-1\\1\end{bmatrix}\}\text{.}\) This is exactly the null space of \(A^T\text{.}\)

Subsection Projection onto a plane

A normal vector is \(\mathbf{n} = \begin{bmatrix}1\\1\\1\end{bmatrix}\text{.}\) Since \(\mathbf{p} \cdot \mathbf{n} = 7\) and \(\mathbf{n} \cdot \mathbf{n} = 3\text{,}\)
\begin{equation*} \operatorname{proj}_U(\mathbf{p}) = \mathbf{p} - (7/3)\mathbf{n} = \begin{bmatrix}-4/3\\-1/3\\5/3\end{bmatrix}. \end{equation*}
The entries add to 0, so the projected point lies in \(U\text{.}\)

Subsection Unreachable target, closest output

The output of \(A\) has the form \((x_1, x_2, x_1+x_2)\text{.}\) The first two target entries force \(x_1 = 2\) and \(x_2 = 1\text{,}\) which would make the third entry \(3\text{,}\) not \(5\text{.}\)
\(A^T A = \begin{bmatrix}2\amp 1\\1\amp 2\end{bmatrix}\text{,}\) \(A^T\mathbf{b} = \begin{bmatrix}7\\6\end{bmatrix}\text{.}\)
Solving gives \(\hat{\mathbf{x}} = \begin{bmatrix}8/3\\5/3\end{bmatrix}\text{.}\) Then
\begin{equation*} A\hat{\mathbf{x}} = \begin{bmatrix}8/3\\5/3\\13/3\end{bmatrix}, \end{equation*}
\begin{equation*} \mathbf{r} = \begin{bmatrix}-2/3\\-2/3\\2/3\end{bmatrix}, \end{equation*}
and \(A^T\mathbf{r} = \begin{bmatrix}0\\0\end{bmatrix}\text{.}\)

Subsection Regression design matrix and residual

\(X = \begin{bmatrix}1\amp 0\\1\amp 1\\1\amp 2\end{bmatrix}\text{,}\) \(\mathbf{y} = \begin{bmatrix}1\\2\\2\end{bmatrix}\text{.}\)
\(X^T X = \begin{bmatrix}3\amp 3\\3\amp 5\end{bmatrix}\text{,}\) \(X^T\mathbf{y} = \begin{bmatrix}5\\6\end{bmatrix}\text{.}\)
Solving gives \(\boldsymbol{\beta} = \begin{bmatrix}7/6\\1/2\end{bmatrix}\text{.}\) The fitted values are \(\begin{bmatrix}7/6\\5/3\\13/6\end{bmatrix}\text{.}\) The residual is \(\begin{bmatrix}-1/6\\1/3\\-1/6\end{bmatrix}\text{,}\) and \(X^T\mathbf{r} = \begin{bmatrix}0\\0\end{bmatrix}\text{.}\)

Subsection Code interpretation: residual orthogonality

The index [0] selects the least-squares coefficient vector returned by np.linalg.lstsq.
The coefficient vector is
\begin{equation*} \widehat{\mathbf{x}} = \begin{bmatrix} 8/3\\ 5/3 \end{bmatrix}. \end{equation*}
The fitted vector is
\begin{equation*} \widehat{\mathbf{b}} = A\widehat{\mathbf{x}} = \begin{bmatrix} 8/3\\ 5/3\\ 13/3 \end{bmatrix}, \end{equation*}
and the residual is
\begin{equation*} \mathbf{r} = \mathbf{b}-\widehat{\mathbf{b}} = \begin{bmatrix} -2/3\\ -2/3\\ 2/3 \end{bmatrix}. \end{equation*}
The expression A.T @ r checks
\begin{equation*} A^T\mathbf{r}=\mathbf{0}, \end{equation*}
so it checks that the residual is orthogonal to every column of \(A\text{.}\) The tiny entries come from floating-point roundoff.
The residual is not zero:
\begin{equation*} \|\mathbf{r}\| = \frac{2}{\sqrt{3}} \approx 1.15470. \end{equation*}
This does not mean least squares failed. The target is not reachable, so the closest reachable output has a nonzero residual.

Subsection Debug: wrong residual check

The matrix \(A\) has shape \(m\times n\text{,}\) the transpose \(A^T\) has shape \(n\times m\text{,}\) and \(\mathbf{r}\) has length \(m\text{.}\) Thus
\begin{equation*} A^T\mathbf{r} \end{equation*}
is defined and has length \(n\text{.}\) The product \(A\mathbf{r}\) is usually not defined because \(A\) expects an input of length \(n\text{,}\) not \(m\text{.}\)
Even if the dimensions happened to match, \(A\mathbf{r}\) would form a linear combination of the columns of \(A\text{.}\) It would not collect the dot products of \(\mathbf{r}\) with those columns.
The correct condition is
\begin{equation*} A^T\mathbf{r}=\mathbf{0}. \end{equation*}
A numerical Boolean check is
np.allclose(A.T @ r, np.zeros(A.shape[1]))

Subsection Rank and nonunique coefficients

\(A\mathbf{z} = \mathbf{0}\text{,}\) so \(\mathbf{z}\) is a null-space direction. Therefore
\begin{equation*} A(\hat{\mathbf{x}} + 5\mathbf{z}) = A\hat{\mathbf{x}} + 5A\mathbf{z} = A\hat{\mathbf{x}}. \end{equation*}
The coefficient vector can change while the fitted vector stays fixed. This happens because the third column is the sum of the first two columns.

Subsection \(QR\) least-squares code reading

Since \(A\) has shape \(3\times2\text{,}\) reduced \(QR\) gives
\begin{equation*} Q:3\times2, \qquad R:2\times2. \end{equation*}
The first Boolean check verifies
\begin{equation*} Q^TQ=I_2, \end{equation*}
and the second verifies
\begin{equation*} QR=A. \end{equation*}
The solve computes the reduced system
\begin{equation*} R\widehat{\mathbf{x}} = Q^T\mathbf{b}. \end{equation*}
Both methods return
\begin{equation*} \widehat{\mathbf{x}} = \begin{bmatrix} 8/3\\ 5/3 \end{bmatrix} \end{equation*}
up to roundoff, so the final comparison is True.
The agreement does not mean that \(A\mathbf{x}=\mathbf{b}\) has an exact solution. The residual is nonzero.
A \(QR\) factorization permits a simultaneous sign change in one column of \(Q\) and the corresponding row of \(R\text{.}\) Numerical software may therefore return signs different from a hand calculation while still satisfying
\begin{equation*} Q^TQ=I \qquad\text{and}\qquad QR=A. \end{equation*}

Subsection Projection and attention: score, then combine

For the projection pipeline,
\begin{equation*} \mathbf{c} = Q^T\mathbf{x} = \begin{bmatrix} 2\\ 1 \end{bmatrix}. \end{equation*}
Therefore
\begin{equation*} \widehat{\mathbf{x}} = Q\mathbf{c} = \begin{bmatrix} 2\\ 1\\ 0 \end{bmatrix}, \qquad \mathbf{r} = \begin{bmatrix} 0\\ 0\\ 3 \end{bmatrix}. \end{equation*}
The residual check gives
\begin{equation*} Q^T\mathbf{r} = \begin{bmatrix} 0\\ 0 \end{bmatrix}. \end{equation*}
For the attention pipeline,
\begin{equation*} \mathbf{s} = K\mathbf{q} = \begin{bmatrix} 2\\ 1 \end{bmatrix}, \qquad \boldsymbol{\alpha} = \begin{bmatrix} 2/3\\ 1/3 \end{bmatrix}. \end{equation*}
The output is
\begin{equation*} \mathbf{h} = \boldsymbol{\alpha}^TV = \frac23 \begin{bmatrix} 3\amp 0 \end{bmatrix} + \frac13 \begin{bmatrix} 0\amp 6 \end{bmatrix} = \begin{bmatrix} 2\amp 2 \end{bmatrix}. \end{equation*}
Both pipelines use dot products or matrix products to produce scores and use linear combinations to combine vectors.
Projection scores the orthonormal directions in \(Q\) and combines those same directions. Its residual is orthogonal to their span. Attention scores key vectors and combines value vectors, which may be different vectors.