Skip to main content
Contents
Dark Mode Prev Up Next
\(\newcommand{\N}{\mathbb{N}}
\newcommand{\Z}{\mathbb{Z}}
\newcommand{\Q}{\mathbb{Q}}
\newcommand{\R}{\mathbb{R}}
\newcommand{\dimens}{\operatorname{dim}}
\DeclareMathOperator{\row}{\operatorname{row}}
\DeclareMathOperator{\col}{\operatorname{col}}
\newcommand{\im}{\operatorname{im}}
\newcommand{\nulls}{\operatorname{null}}
\newcommand{\minor}{\operatorname{minor}}
\newcommand{\spans}{\operatorname{span}}
\newcommand{\nullity}{\operatorname{nullity}}
\newcommand{\kers}{\operatorname{ker}}
\newcommand{\proj}{\operatorname{proj}}
\newcommand{\diag}{\operatorname{diag}}
\newcommand{\Tr}{\operatorname{Tr}}
\newcommand{\rank}{\operatorname{rank}}
\newcommand{\lt}{<}
\newcommand{\gt}{>}
\newcommand{\amp}{&}
\definecolor{fillinmathshade}{gray}{0.9}
\newcommand{\fillinmath}[1]{\mathchoice{\colorbox{fillinmathshade}{$\displaystyle \phantom{\,#1\,}$}}{\colorbox{fillinmathshade}{$\textstyle \phantom{\,#1\,}$}}{\colorbox{fillinmathshade}{$\scriptstyle \phantom{\,#1\,}$}}{\colorbox{fillinmathshade}{$\scriptscriptstyle\phantom{\,#1\,}$}}}
\)
Section 4.5 Coding recap
Subsection Linked notebook
For a quick reference on least-squares commands, transposes, matrix products, and numerical checks such as "close to zero," see the programming appendix sections
B.10 ,
B.4 , and
B.9 .
Subsection Review activities
Activity 4.5.1 . Reading least-squares objects (U4-LO4, U4-LO7).
A = np.array([
[1.0, 0.0],
[0.0, 1.0],
[1.0, 1.0],
])
b = np.array([1.0, 2.0, 4.0])
xhat = np.linalg.lstsq(A, b, rcond=None)[0]
bhat = A @ xhat
r = b - bhat
xhat.shape, bhat.shape, r.shape, A.T @ r
What is the shape of
xhat?
What is the shape of
bhat?
What mathematical object is stored in
xhat?
What mathematical object is stored in
bhat?
Which variable stores the closest reachable output?
Does
A.T @ r being close to zero imply that
r is close to the zero vector?
Solution .
The matrix
\(A\) has two columns and three rows. Therefore
xhat has shape
(2,), while
bhat and
r both have shape
(3,).
The vector
xhat stores the least-squares coefficient vector
\begin{equation*}
\widehat{\mathbf{x}}
=
\begin{bmatrix}
4/3\\
7/3
\end{bmatrix}.
\end{equation*}
The vector
bhat stores the fitted vector
\begin{equation*}
\widehat{\mathbf{b}}
=
A\widehat{\mathbf{x}}
=
\begin{bmatrix}
4/3\\
7/3\\
11/3
\end{bmatrix}.
\end{equation*}
The vector
r stores the residual
\begin{equation*}
\mathbf{r}
=
\mathbf{b}-\widehat{\mathbf{b}}
=
\begin{bmatrix}
-1/3\\
-1/3\\
1/3
\end{bmatrix}.
\end{equation*}
The closest reachable output is
bhat, not
xhat.
computes the dot products of the residual with the columns of
\(A\text{.}\) It checks that
\begin{equation*}
\mathbf{r}
\perp
\operatorname{col}(A).
\end{equation*}
This does not mean that the residual is zero. Here,
\begin{equation*}
\|\mathbf{r}\|^2
=
\frac13,
\end{equation*}
so the residual is nonzero.
Activity 4.5.2 . Roundoff and the wrong residual check (U4-LO4, U4-LO7).
array([2.2e-16, -1.1e-16])
What geometric condition is checked by the first output?
Why are the entries tiny rather than exactly zero?
Does the first output say that the residual vector is zero?
What does the second output show?
np.allclose(A.T @ r, np.zeros(A.shape[1]))
a more appropriate numerical check than
A.T @ r == 0?
to check residual orthogonality. What is wrong with this expression?
Solution .
The first output checks whether the residual is orthogonal to every column of
\(A\text{.}\) In exact arithmetic,
\begin{equation*}
A^T\mathbf{r}
=
\mathbf{0}.
\end{equation*}
Floating-point arithmetic may produce tiny roundoff errors instead of exact zeros.
The first output does not say that
\(\mathbf{r}=\mathbf{0}\text{.}\) The second output shows that
\begin{equation*}
\|\mathbf{r}\|
\approx
0.57735
=
\frac{1}{\sqrt{3}},
\end{equation*}
so the residual is nonzero.
The function
np.allclose checks whether the entries are numerically close to zero within a tolerance. An entry-by-entry equality check is usually too strict for floating-point output.
The expression
A @ r is not the correct check. Here
\(A\) has shape
\(3\times2\text{,}\) while
\(\mathbf{r}\) has length
\(3\text{,}\) so the product is not even defined. More importantly, residual orthogonality requires dot products with the columns of
\(A\text{,}\) which are collected by
\begin{equation*}
A^T\mathbf{r}.
\end{equation*}
Activity 4.5.3 . Reading a regression design matrix (U4-LO4, U4-LO7).
t = np.array([0.0, 1.0, 2.0])
y = np.array([1.0, 2.0, 2.0])
X = np.column_stack([np.ones_like(t), t])
c = np.linalg.lstsq(X, y, rcond=None)[0]
yhat = X @ c
r = y - yhat
What do the two columns of
\(X\) represent?
Why does
c have two entries?
Why does
yhat have three entries?
Which variable is the fitted vector?
Which variable is the coefficient vector?
Which vector lies in
\(\operatorname{col}(X)\text{?}\)
In the
\(t\) -
\(y\) data plot, are the residual segments vertical or perpendicular to the fitted line?
What quantity is computed by
Solution .
The first column of
\(X\) is the constant feature. The second column contains the input values
\(t\text{.}\)
The vector
c has two entries because the model has two coefficients: the intercept and slope. The vector
yhat has three entries because the model makes one prediction at each of the three input values.
The coefficient vector is
\begin{equation*}
\widehat{\mathbf{c}}
=
\begin{bmatrix}
7/6\\
1/2
\end{bmatrix}.
\end{equation*}
\begin{equation*}
\widehat{\mathbf{y}}
=
X\widehat{\mathbf{c}}
=
\begin{bmatrix}
7/6\\
5/3\\
13/6
\end{bmatrix}.
\end{equation*}
Thus
c is the coefficient vector, while
yhat is the fitted vector. The vector
yhat lies in
\(\operatorname{col}(X)\text{.}\)
The right panel of β
Two views of the same line fitΒ 4.3.9 β shows the projection in
\(\mathbb{R}^3\text{.}\) In the
\(t\) -
\(y\) data plot, the residual segments are vertical. They are not generally perpendicular to the fitted line.
computes the sum of squared residuals. Its value is
\begin{equation*}
\frac16
\end{equation*}
Activity 4.5.4 . Reading QR least-squares code (U4-LO6, U4-LO7).
Q, R = np.linalg.qr(X, mode="reduced")
c_qr = np.linalg.solve(R, Q.T @ y)
np.allclose(Q.T @ Q, np.eye(Q.shape[1]))
np.allclose(Q @ R, X)
c_qr
What should the shapes of
Q and
R be?
What does the first
np.allclose line check?
What does the second
np.allclose line check?
What mathematical system is solved by
np.linalg.solve(R, Q.T @ y)
Why is a system with
\(R\) easier to solve than the original least-squares problem?
What should
c_qr be close to?
Does obtaining
c_qr mean that
\begin{equation*}
X\mathbf{c}=\mathbf{y}
\end{equation*}
Why might NumPy return signs in
\(Q\) and
\(R\) that differ from a hand GramβSchmidt calculation?
Solution .
The design matrix
\(X\) has shape
\(3\times2\text{.}\) Reduced QR therefore gives
Q.shape == (3, 2)
R.shape == (2, 2)
\begin{equation*}
Q^TQ=I_2,
\end{equation*}
so the columns of
\(Q\) are orthonormal. The second check verifies
\begin{equation*}
QR=X.
\end{equation*}
np.linalg.solve(R, Q.T @ y)
\begin{equation*}
R\widehat{\mathbf{c}}
=
Q^T\mathbf{y}.
\end{equation*}
The matrix
\(R\) is square and upper triangular, so this is a standard linear system rather than a rectangular least-squares problem.
The result should be close to
\begin{equation*}
\widehat{\mathbf{c}}
=
\begin{bmatrix}
7/6\\
1/2
\end{bmatrix}.
\end{equation*}
This does not mean that
\(X\mathbf{c}=\mathbf{y}\) has an exact solution. The fitted vector has a nonzero residual.
A QR factorization allows simultaneous sign changes in a column of
\(Q\) and the corresponding row of
\(R\text{.}\) Therefore numerical software may return signs different from a hand calculation while still satisfying
\begin{equation*}
Q^TQ=I
\qquad\text{and}\qquad
QR=X.
\end{equation*}