Section 4.7 Unit 4 highlights
Subsection Mathematical quick reference
Orthogonal coordinates.
Orthogonal and orthonormal lists. Distinct vectors in an orthogonal list have dot product zero. An orthonormal list also has every norm equal to \(1\text{.}\) A list of nonzero pairwise orthogonal vectors is independent.
Coordinates in an orthogonal basis. If \(\mathbf f_1,\ldots,\mathbf f_k\) is an orthogonal basis of \(U\) and \(\mathbf x\in U\text{,}\) then
\begin{equation*}
\mathbf x=\sum_{i=1}^k c_i\mathbf f_i,\qquad c_i=\frac{\mathbf x\cdot\mathbf f_i}{\mathbf f_i\cdot\mathbf f_i}.
\end{equation*}
For an orthonormal basis \(\mathbf q_i\text{,}\) this simplifies to \(c_i=\mathbf x\cdot\mathbf q_i\text{.}\)
Orthonormal columns. For \(Q\in\mathbb R^{m\times k}\text{,}\) \(Q^TQ=I_k\) exactly when its columns are orthonormal. If \(Q\) is square, it is an orthogonal matrix and \(Q^{-1}=Q^T\text{.}\)
Orthogonal complements and projection.
Orthogonal complement.
\begin{equation*}
U^\perp=\{\mathbf r:\mathbf r\cdot\mathbf u=0\text{ for every }\mathbf u\in U\}.
\end{equation*}
For \(A\in\mathbb R^{m\times n}\text{,}\) \(\operatorname{col}(A)^\perp=\operatorname{null}(A^T)\) and \(\operatorname{row}(A)^\perp=\operatorname{null}(A)\text{.}\)
Orthogonal decomposition. For a subspace \(U\subseteq\mathbb R^m\text{,}\) every \(\mathbf b\) has a unique decomposition
\begin{equation*}
\mathbf b=\mathbf p+\mathbf r,\qquad \mathbf p\in U,\quad \mathbf r\in U^\perp.
\end{equation*}
Here \(\mathbf p=\operatorname{proj}_U\mathbf b\text{,}\) and \(\dim U+\dim U^\perp=m\text{.}\)
Projection formulas. For a nonzero \(\mathbf v\text{,}\) an orthogonal basis \(\mathbf f_1,\ldots,\mathbf f_k\) of \(U\text{,}\) or a matrix \(Q\) with orthonormal columns spanning \(U\text{,}\) respectively,
\begin{equation*}
\operatorname{proj}_{\operatorname{span}\{\mathbf v\}}\mathbf b=\frac{\mathbf b\cdot\mathbf v}{\mathbf v\cdot\mathbf v}\mathbf v,\qquad \operatorname{proj}_U\mathbf b=\sum_{i=1}^k\frac{\mathbf b\cdot\mathbf f_i}{\mathbf f_i\cdot\mathbf f_i}\mathbf f_i=QQ^T\mathbf b.
\end{equation*}
Best approximation. If \(\mathbf p=\operatorname{proj}_U\mathbf b\text{,}\) then for every \(\mathbf u\in U\text{,}\)
\begin{equation*}
\|\mathbf b-\mathbf u\|^2=\|\mathbf b-\mathbf p\|^2+\|\mathbf p-\mathbf u\|^2.
\end{equation*}
Thus \(\mathbf p\) is the unique closest vector in \(U\text{,}\) and \(\operatorname{dist}(\mathbf b,U)=\|\mathbf b-\mathbf p\|\text{.}\)
Gram–Schmidt and QR factorization.
Gram–Schmidt. For independent columns \(\mathbf a_1,\ldots,\mathbf a_n\in\mathbb R^m\text{,}\) process \(j=1,\ldots,n\text{:}\)
\begin{equation*}
r_{ij}=\mathbf q_i^T\mathbf a_j\ (i\lt j),\qquad \mathbf v_j=\mathbf a_j-\sum_{i\lt j}r_{ij}\mathbf q_i,
\end{equation*}
\begin{equation*}
r_{jj}=\|\mathbf v_j\|,\qquad \mathbf q_j=\frac{\mathbf v_j}{r_{jj}}.
\end{equation*}
Independence guarantees \(r_{jj}>0\text{.}\) Each initial span is preserved.
Reduced QR factorization. For \(A\in\mathbb R^{m\times n}\) with independent columns,
\begin{equation*}
A=QR,\qquad Q\in\mathbb R^{m\times n},\quad Q^TQ=I_n,\quad R\in\mathbb R^{n\times n}.
\end{equation*}
The matrix \(R\) is upper triangular, with the coefficients \(r_{ij}\) above and on its diagonal, and is invertible.
Least squares and normal equations.
Least-squares problem. Given \(A\in\mathbb R^{m\times n}\) and \(\mathbf b\in\mathbb R^m\text{,}\) minimize \(\|A\mathbf x-\mathbf b\|^2\text{.}\) For a minimizer \(\widehat{\mathbf x}\text{,}\) distinguish the coefficient vector, fitted vector, and residual:
\begin{equation*}
\widehat{\mathbf x}\in\mathbb R^n,\qquad \mathbf p=A\widehat{\mathbf x}\in\mathbb R^m,\qquad \mathbf r=\mathbf b-\mathbf p\in\mathbb R^m.
\end{equation*}
Equivalent tests.
\begin{equation*}
A\widehat{\mathbf x}=\operatorname{proj}_{\operatorname{col}(A)}\mathbf b\quad\Longleftrightarrow\quad A^T\mathbf r=\mathbf0\quad\Longleftrightarrow\quad A^TA\widehat{\mathbf x}=A^T\mathbf b.
\end{equation*}
Least-squares solutions always exist. The fitted vector and residual are unique; the coefficients are unique exactly when the columns of \(A\) are independent.
Full column rank. In this case \(A^TA\) is invertible, so
\begin{equation*}
\widehat{\mathbf x}=(A^TA)^{-1}A^T\mathbf b,\qquad \mathbf p=A(A^TA)^{-1}A^T\mathbf b.
\end{equation*}
With \(A=QR\text{,}\) solve the upper-triangular system \(R\widehat{\mathbf x}=Q^T\mathbf b\) instead.
Design matrix. For prescribed functions \(\phi_j\) and sample points \(t_i\text{,}\) set \(A_{ij}=\phi_j(t_i)\text{.}\) Fitting the linear combination \(\sum_j c_j\phi_j\) to values \(b_i\) gives \(A^TA\widehat{\mathbf c}=A^T\mathbf b\text{.}\)
Subsection Common mistakes
Confusing orthogonal with orthonormal; omitting the squared norm in a projection coefficient; assuming \(QQ^T=I_m\) for a rectangular \(Q\text{;}\) normalizing a zero Gram–Schmidt residual; confusing coefficients with fitted values; assuming \(A^T\mathbf r=\mathbf0\) means \(\mathbf r=\mathbf0\text{;}\) or using \((A^TA)^{-1}\) without checking column independence.
Subsection Connections
Unit 1 supplies dot products and linear combinations. Unit 2 supplies subspaces, rank, and null spaces. Unit 3 uses the projection of a gradient onto a unit direction for directional derivatives. Unit 5 derives the normal equations using gradients. Unit 7 extends projection and Gram–Schmidt to polynomial and inner product spaces.
