Skip to main content
Contents
Dark Mode Prev Up Next
\(\newcommand{\N}{\mathbb{N}}
\newcommand{\Z}{\mathbb{Z}}
\newcommand{\Q}{\mathbb{Q}}
\newcommand{\R}{\mathbb{R}}
\newcommand{\dimens}{\operatorname{dim}}
\DeclareMathOperator{\row}{\operatorname{row}}
\DeclareMathOperator{\col}{\operatorname{col}}
\newcommand{\im}{\operatorname{im}}
\newcommand{\nulls}{\operatorname{null}}
\newcommand{\minor}{\operatorname{minor}}
\newcommand{\spans}{\operatorname{span}}
\newcommand{\nullity}{\operatorname{nullity}}
\newcommand{\kers}{\operatorname{ker}}
\newcommand{\proj}{\operatorname{proj}}
\newcommand{\diag}{\operatorname{diag}}
\newcommand{\Tr}{\operatorname{Tr}}
\newcommand{\rank}{\operatorname{rank}}
\newcommand{\lt}{<}
\newcommand{\gt}{>}
\newcommand{\amp}{&}
\definecolor{fillinmathshade}{gray}{0.9}
\newcommand{\fillinmath}[1]{\mathchoice{\colorbox{fillinmathshade}{$\displaystyle \phantom{\,#1\,}$}}{\colorbox{fillinmathshade}{$\textstyle \phantom{\,#1\,}$}}{\colorbox{fillinmathshade}{$\scriptstyle \phantom{\,#1\,}$}}{\colorbox{fillinmathshade}{$\scriptscriptstyle\phantom{\,#1\,}$}}}
\)
Section 4.6 Exercises
Learning outcomes. The labels below identify the learning outcomes for each exercise group. Individual problems may also involve earlier outcomes.
Subsection Orthogonality and orthogonal bases
Learning outcomes. U4-LO1, U4-LO2.
Subsection Orthogonal projection and Gram-Schmidt
Learning outcomes. U4-LO2, U4-LO3.
Nicholson 8.1.1(d) math.libretexts.org/Bookshelves/Linear_Algebra/Linear_Algebra_with_Applications_(Nicholson)/08%3A_Orthogonality/8.01%3A_Orthogonal_Complements_and_Projections/8.1E%3A_Orthogonal_Complements_and_Projections_Exercises
Nicholson 8.1.2(b,d) math.libretexts.org/Bookshelves/Linear_Algebra/Linear_Algebra_with_Applications_(Nicholson)/08%3A_Orthogonality/8.01%3A_Orthogonal_Complements_and_Projections/8.1E%3A_Orthogonal_Complements_and_Projections_Exercises
Nicholson 8.1.4(b) math.libretexts.org/Bookshelves/Linear_Algebra/Linear_Algebra_with_Applications_(Nicholson)/08%3A_Orthogonality/8.01%3A_Orthogonal_Complements_and_Projections/8.1E%3A_Orthogonal_Complements_and_Projections_Exercises
Subsection QR factorization and least squares
Learning outcomes. U4-LO6.
Subsection Additional exercises
These exercises connect the main Unit 4 ideas: orthogonality, orthogonal bases, GramβSchmidt, projections, residuals, least squares, normal equations, regression, QR factorization, and code interpretation. Solutions are collected in Appendix
C.4 .
Additional exercise 4.6.1 . Orthogonal coordinates and projection.
\begin{equation*}
\mathbf{q}_1
=
\frac{1}{\sqrt{2}}
\begin{bmatrix}
1\\
1\\
0
\end{bmatrix},
\qquad
\mathbf{q}_2
=
\frac{1}{\sqrt{2}}
\begin{bmatrix}
1\\
-1\\
0
\end{bmatrix},
\end{equation*}
\begin{equation*}
Q
=
\begin{bmatrix}
\mathbf{q}_1\amp\mathbf{q}_2
\end{bmatrix},
\qquad
\mathbf{x}
=
\begin{bmatrix}
3\\
1\\
2
\end{bmatrix}.
\end{equation*}
\begin{equation*}
Q^TQ=I_2.
\end{equation*}
Compute the coordinate vector
\begin{equation*}
\mathbf{c}=Q^T\mathbf{x}.
\end{equation*}
\begin{equation*}
\widehat{\mathbf{x}}=Q\mathbf{c}.
\end{equation*}
\begin{equation*}
\mathbf{r}
=
\mathbf{x}-\widehat{\mathbf{x}}.
\end{equation*}
\begin{equation*}
Q^T\mathbf{r}=\mathbf{0}.
\end{equation*}
Explain how the calculation follows the pattern βscore, then combine.β
Which vector is the projection of
\(\mathbf{x}\text{:}\) \(\mathbf{c}\) or
\(\widehat{\mathbf{x}}\text{?}\)
Learning outcomes. U4-LO1, U4-LO3.
Additional exercise 4.6.2 . One GramβSchmidt column step.
Suppose the first completed unit direction is
\begin{equation*}
\mathbf{q}_1
=
\frac{1}{\sqrt{2}}
\begin{bmatrix}
1\\
1\\
0
\end{bmatrix},
\end{equation*}
and the next original column is
\begin{equation*}
\mathbf{a}_2
=
\begin{bmatrix}
2\\
0\\
1
\end{bmatrix}.
\end{equation*}
\begin{equation*}
r_{12}
=
\mathbf{q}_1^T\mathbf{a}_2.
\end{equation*}
\begin{equation*}
\mathbf{p}_2
=
r_{12}\mathbf{q}_1.
\end{equation*}
\begin{equation*}
\mathbf{v}_2
=
\mathbf{a}_2-\mathbf{p}_2.
\end{equation*}
\begin{equation*}
r_{22}
=
\|\mathbf{v}_2\|.
\end{equation*}
\begin{equation*}
\mathbf{q}_2
=
\frac{\mathbf{v}_2}{r_{22}}.
\end{equation*}
\begin{equation*}
\mathbf{q}_1^T\mathbf{q}_2=0.
\end{equation*}
What would
\(r_{22}=0\) say about
\(\mathbf{a}_2\text{?}\)
Learning outcomes. U4-LO2, U4-LO6.
Additional exercise 4.6.3 . Orthogonal-basis coefficients.
Let
\begin{equation*}
\mathbf{f}_1 =
\begin{bmatrix}1\\1\\0\end{bmatrix},\quad
\mathbf{f}_2 =
\begin{bmatrix}1\\-1\\0\end{bmatrix},\quad
\mathbf{f}_3 =
\begin{bmatrix}0\\0\\2\end{bmatrix},
\end{equation*}
and let
\begin{equation*}
\mathbf{x} =
\begin{bmatrix}3\\1\\4\end{bmatrix}.
\end{equation*}
Verify that
\(\mathbf{f}_1,\mathbf{f}_2,\mathbf{f}_3\) are pairwise orthogonal.
Use dot products to write
\(\mathbf{x}\) as a linear combination of
\(\mathbf{f}_1,\mathbf{f}_2,\mathbf{f}_3\text{.}\)
Explain why each dot product isolates one coefficient.
Learning outcomes. U4-LO1, U4-LO2.
Additional exercise 4.6.4 . Projection onto a line.
Let
\begin{equation*}
\mathbf{u} = \begin{bmatrix}1\\2\end{bmatrix},\quad
\mathbf{x} = \begin{bmatrix}3\\1\end{bmatrix},
\end{equation*}
and let \(L = \operatorname{span}\{\mathbf{u}\}\text{.}\)
Compute
\(\operatorname{proj}_L(\mathbf{x})\text{.}\)
Compute
\(\mathbf{r} = \mathbf{x} - \operatorname{proj}_L(\mathbf{x})\text{.}\)
Check that
\(\mathbf{r} \cdot \mathbf{u} = 0\text{.}\)
Explain what the residual measures.
Learning outcomes. U4-LO3.
Additional exercise 4.6.5 . Column-space orthogonal complement.
Let
\begin{equation*}
A =
\begin{bmatrix}
1\amp 0\\
0\amp 1\\
1\amp 1
\end{bmatrix}.
\end{equation*}
Compute
\(A^T\mathbf{r}\) for
\(\mathbf{r} = \begin{bmatrix}-1\\-1\\1\end{bmatrix}\text{.}\)
Explain why
\(\mathbf{r}\) is orthogonal to
\(\operatorname{col}(A)\text{.}\)
Describe
\(\operatorname{col}(A)^\perp\text{.}\)
Explain the identity
\(\operatorname{col}(A)^\perp = \operatorname{null}(A^T)\) in this example.
Learning outcomes. U4-LO3, U4-LO4, U2-LO3.
Additional exercise 4.6.6 . Projection onto a plane.
Let \(U\) be the plane \(x + y + z = 0\text{,}\) and let
\begin{equation*}
\mathbf{p} = \begin{bmatrix}1\\2\\4\end{bmatrix}.
\end{equation*}
Give a normal vector
\(\mathbf{n}\) for
\(U\text{.}\)
Compute
\(\operatorname{proj}_U(\mathbf{p})\) by subtracting the component of
\(\mathbf{p}\) in the normal direction.
Check that the projected point lies in
\(U\text{.}\)
Learning outcomes. U4-LO3.
Additional exercise 4.6.7 . Unreachable target, closest output.
Let
\begin{equation*}
A =
\begin{bmatrix}
1\amp 0\\
0\amp 1\\
1\amp 1
\end{bmatrix},
\qquad
\mathbf{b} = \begin{bmatrix}2\\1\\5\end{bmatrix}.
\end{equation*}
Explain why
\(A\mathbf{x} = \mathbf{b}\) has no exact solution.
Form
\(A^T A\) and
\(A^T\mathbf{b}\text{.}\)
Solve the normal equations.
Compute
\(A\hat{\mathbf{x}}\) and
\(\mathbf{r} = \mathbf{b} - A\hat{\mathbf{x}}\text{.}\)
Check
\(A^T\mathbf{r} = \mathbf{0}\text{.}\)
Learning outcomes. U4-LO4, U4-LO5, U2-LO3.
Additional exercise 4.6.8 . Regression design matrix and residual.
Fit a line \(y = c_0 + c_1 t\) to the data
\begin{equation*}
(0,1),\quad (1,2),\quad (2,2).
\end{equation*}
Build the design matrix
\(X\text{.}\)
Form
\(X^T X\) and
\(X^T\mathbf{y}\text{.}\)
Solve the normal equations.
Compute the fitted values and residual.
Check
\(X^T\mathbf{r} = \mathbf{0}\text{.}\)
Learning outcomes. U4-LO5.
Additional exercise 4.6.9 . Code interpretation: residual orthogonality.
import numpy as np
A = np.array([
[1.0, 0.0],
[0.0, 1.0],
[1.0, 1.0],
])
b = np.array([2.0, 1.0, 5.0])
xhat = np.linalg.lstsq(A, b, rcond=None)[0]
bhat = A @ xhat
r = b - bhat
orthogonality = A.T @ r
xhat, bhat, r, orthogonality, np.linalg.norm(r)
Suppose the last two outputs are approximately
array([1.3e-15, -2.2e-16])
What does
[0] select from the value returned by
np.linalg.lstsq?
What mathematical object is stored in
xhat?
What mathematical object is stored in
bhat?
What mathematical object is stored in
r?
What geometric condition does
A.T @ r check?
Why are the entries of
orthogonality tiny rather than exactly zero?
Does the nonzero residual norm mean that least squares failed?
Learning outcomes. U4-LO4, U4-LO5, U4-LO7.
Additional exercise 4.6.10 . Debug: wrong residual check.
A student tries to check residual orthogonality with
Assume that
\(A\) is an
\(m\times n\) matrix and
\(\mathbf{r}\in\mathbb{R}^m\text{.}\)
What are the shapes of
\(A\text{,}\) \(A^T\text{,}\) and
\(\mathbf{r}\text{?}\)
Why is
A @ r usually not defined?
Suppose the dimensions happened to make
A @ r defined. Why would it still be the wrong geometric check?
State the correct mathematical residual-orthogonality condition.
Write a numerical Boolean check using
np.allclose,
A.shape[1], and
np.zeros.
Learning outcomes. U4-LO4, U4-LO7.
Additional exercise 4.6.11 . Rank and nonunique coefficients.
Let
\begin{equation*}
A =
\begin{bmatrix}
1\amp 0\amp 1\\
0\amp 1\amp 1\\
1\amp 1\amp 2
\end{bmatrix},
\end{equation*}
and let \(\mathbf{z} = \begin{bmatrix}1\\1\\-1\end{bmatrix}\text{.}\)
Compute
\(A\mathbf{z}\text{.}\)
Suppose
\(\hat{\mathbf{x}}\) is a least-squares solution. Compare
\(A\hat{\mathbf{x}}\) and
\(A(\hat{\mathbf{x}} + 5\mathbf{z})\text{.}\)
Which can be nonunique: the fitted vector or the coefficient vector?
How does this connect to dependent columns?
Learning outcomes. U4-LO4, U2-LO6.
Additional exercise 4.6.12 . \(QR\) least-squares code reading.
import numpy as np
A = np.array([
[1.0, 0.0],
[0.0, 1.0],
[1.0, 1.0],
])
b = np.array([2.0, 1.0, 5.0])
Q, R = np.linalg.qr(A, mode="reduced")
x_qr = np.linalg.solve(R, Q.T @ b)
x_lstsq = np.linalg.lstsq(A, b, rcond=None)[0]
(
Q.shape,
R.shape,
np.allclose(Q.T @ Q, np.eye(Q.shape[1])),
np.allclose(Q @ R, A),
x_qr,
x_lstsq,
np.allclose(x_qr, x_lstsq),
)
Why does the code request
mode="reduced"?
What should the shapes of
Q and
R be?
np.allclose(Q.T @ Q, np.eye(Q.shape[1]))
What mathematical system is solved by
np.linalg.solve(R, Q.T @ b)
Why should
x_qr and
x_lstsq agree?
Does their agreement mean that
\begin{equation*}
A\mathbf{x}=\mathbf{b}
\end{equation*}
Why may the entries of
Q and
R have signs different from a hand GramβSchmidt calculation?
Learning outcomes. U4-LO5, U4-LO6, U4-LO7.
Additional exercise 4.6.13 . Projection and attention: score, then combine.
\begin{equation*}
Q
=
\begin{bmatrix}
1\amp 0\\
0\amp 1\\
0\amp 0
\end{bmatrix},
\qquad
\mathbf{x}
=
\begin{bmatrix}
2\\
1\\
3
\end{bmatrix}.
\end{equation*}
1. Compute the score vector
\begin{equation*}
\mathbf{c}=Q^T\mathbf{x}.
\end{equation*}
\begin{equation*}
\widehat{\mathbf{x}}=Q\mathbf{c}.
\end{equation*}
\begin{equation*}
\mathbf{r}
=
\mathbf{x}-\widehat{\mathbf{x}},
\end{equation*}
\begin{equation*}
Q^T\mathbf{r}=\mathbf{0}.
\end{equation*}
\begin{equation*}
K
=
\begin{bmatrix}
1\amp 0\\
0\amp 1
\end{bmatrix},
\qquad
\mathbf{q}
=
\begin{bmatrix}
2\\
1
\end{bmatrix}.
\end{equation*}
\begin{equation*}
\mathbf{s}=K\mathbf{q}.
\end{equation*}
A weighting rule divides the positive scores by their sum:
\begin{equation*}
\boldsymbol{\alpha}
=
\frac{\mathbf{s}}{s_1+s_2}.
\end{equation*}
Let the value vectors be the rows of
\begin{equation*}
V
=
\begin{bmatrix}
3\amp 0\\
0\amp 6
\end{bmatrix}.
\end{equation*}
4. Compute
\(\mathbf{s}\) and
\(\boldsymbol{\alpha}\text{.}\)
5. Compute the attention output
\begin{equation*}
\mathbf{h}
=
\boldsymbol{\alpha}^T V.
\end{equation*}
6. Identify the score step and the combine step in each pipeline.
7. What mathematical operations do both pipelines use?
8. What extra condition identifies
\(\widehat{\mathbf{x}}\) as an orthogonal projection?
9. In the attention pipeline, which vectors are scored and which vectors are combined?
Learning outcomes. U4-LO1, U4-LO3, U1-LO8.