Skip to main content

MATH 345: Linear Algebra and Optimization

Section 5.9 Unit 5 highlights

Subsection Mathematical quick reference

Eigenvalues and eigenspaces.

Eigenpair. For a square matrix \(A\text{,}\) an eigenvalue \(\lambda\) and eigenvector \(\mathbf v\neq\mathbf0\) satisfy
\begin{equation*} A\mathbf v=\lambda\mathbf v. \end{equation*}
The eigenspace is \(E_\lambda=\operatorname{null}(A-\lambda I)\text{;}\) it includes the zero vector, although an eigenvector cannot be zero.
Calculation. Solve \(\det(A-\lambda I)=0\text{,}\) then solve \((A-\lambda I)\mathbf v=\mathbf0\) for each eigenvalue. Eigenvectors for distinct eigenvalues are independent.
Multiplicities. Algebraic multiplicity is multiplicity as a root of the characteristic polynomial; geometric multiplicity is \(\dim E_\lambda\text{,}\) and is at least \(1\) and at most the algebraic multiplicity.

Diagonalization.

Eigenvector basis. A real \(n\times n\) matrix has a real diagonalization exactly when it has a basis of real eigenvectors. With eigenvectors as columns of \(P\) and their eigenvalues in the same order on \(D\text{,}\)
\begin{equation*} AP=PD,\qquad A=PDP^{-1},\qquad A^k=PD^kP^{-1}\quad(k=0,1,2,\ldots). \end{equation*}
In particular, \(n\) distinct real eigenvalues suffice. Repeated eigenvalues require checking eigenspace dimensions.
Spectral theorem. A real symmetric matrix has an orthonormal eigenvector basis:
\begin{equation*} A=QDQ^T,\qquad Q^TQ=I. \end{equation*}
Find orthonormal bases in the eigenspaces and assemble \(Q\text{;}\) eigenvectors for distinct eigenvalues are automatically orthogonal.

Quadratic forms and definiteness.

Quadratic form. \(q_A(\mathbf x)=\mathbf x^TA\mathbf x\text{.}\) For symmetric \(A=QDQ^T\text{,}\) use principal coordinates \(\mathbf y=Q^T\mathbf x\text{:}\)
\begin{equation*} q_A(\mathbf x)=\mathbf y^TD\mathbf y=\sum_{i=1}^n\lambda_i y_i^2. \end{equation*}
Sign tests for symmetric matrices. Positive definite means every eigenvalue is positive, equivalently \(q_A(\mathbf x)>0\) for all \(\mathbf x\neq\mathbf0\text{.}\) Negative definite means every eigenvalue is negative, equivalently \(q_A(\mathbf x)\lt0\) for all nonzero \(\mathbf x\text{.}\) Indefinite means at least one positive and one negative eigenvalue, so the quadratic form takes both signs.
Two-by-two shortcut. For \(A=\begin{bmatrix}a\amp b\\b\amp d\end{bmatrix}\) and \(\Delta=ad-b^2\text{:}\) positive definite if \(\Delta>0\) and \(a>0\text{;}\) negative definite if \(\Delta>0\) and \(a\lt0\text{;}\) indefinite if \(\Delta\lt0\text{.}\)

Hessians and the second derivative test.

Quadratic local approximation. If \(f\) has continuous second partial derivatives near \(\mathbf a\text{,}\) then
\begin{equation*} f(\mathbf a+\mathbf h)\approx f(\mathbf a)+\nabla f(\mathbf a)^T\mathbf h+\frac12\mathbf h^TH_f(\mathbf a)\mathbf h. \end{equation*}
Classify an interior critical point. First check \(\nabla f(\mathbf a)=\mathbf0\text{.}\) For a function with continuous second partials near that point, all positive Hessian eigenvalues give a strict local minimum; all negative give a strict local maximum; a positive and a negative eigenvalue give a saddle point. All remaining sign patterns are inconclusive.
Two-variable test. At a critical point \((a,b)\text{,}\) set
\begin{equation*} D=f_{xx}(a,b)f_{yy}(a,b)-f_{xy}(a,b)^2. \end{equation*}
If \(D>0\text{,}\) use the sign of \(f_{xx}(a,b)\text{:}\) positive gives a strict local minimum, negative a strict local maximum. If \(D\lt0\text{,}\) there is a saddle point. If \(D=0\text{,}\) the test is inconclusive.

Least squares: projection, gradient, and curvature.

Squared-error function. For \(L(\mathbf x)=\|A\mathbf x-\mathbf b\|^2\text{,}\)
\begin{equation*} \nabla L(\mathbf x)=2A^T(A\mathbf x-\mathbf b),\qquad H_L=2A^TA,\qquad \mathbf z^TH_L\mathbf z=2\|A\mathbf z\|^2\geq0. \end{equation*}
Minimizer conditions. With \(\mathbf r=\mathbf b-A\widehat{\mathbf x}\text{,}\)
\begin{equation*} \nabla L(\widehat{\mathbf x})=\mathbf0\quad\Longleftrightarrow\quad A^T\mathbf r=\mathbf0\quad\Longleftrightarrow\quad A^TA\widehat{\mathbf x}=A^T\mathbf b. \end{equation*}
Every solution of these equations is a global minimizer, since
\begin{equation*} L(\widehat{\mathbf x}+\mathbf z)=L(\widehat{\mathbf x})+\|A\mathbf z\|^2. \end{equation*}
Uniqueness. The minimizer is unique exactly when \(\operatorname{null}(A)=\{\mathbf0\}\text{.}\) Otherwise the complete minimizer set is \(\widehat{\mathbf x}+\operatorname{null}(A)\text{.}\) Singularity of the Hessian does not prevent global minimality here.
Fixed functions, variable coefficients. If \(q(t)=\sum_j c_j\phi_j(t)\) and the functions \(\phi_j\) are fixed, set \(A_{ij}=\phi_j(t_i)\text{.}\) Minimizing \(\sum_i(q(t_i)-b_i)^2\) is the same least-squares problem in \(\mathbf c\text{,}\) even when the \(\phi_j\) are nonlinear functions of \(t\text{.}\)

Subsection Common mistakes

Using the zero vector as an eigenvector; assuming every square matrix is diagonalizable; ordering the eigenvalues differently from the columns of \(P\text{;}\) replacing \(P^{-1}\) by \(P^T\) without orthonormal columns; judging definiteness only from diagonal entries; applying the Hessian test away from a critical point; or treating a zero eigenvalue as conclusive. Mixed positive and negative eigenvalues still imply a saddle even if other eigenvalues are zero.
Confusing a local extremum with a global one; assuming residual orthogonality means a zero residual; assuming least-squares coefficients are unique without a rank check; or overlooking linearity in the coefficients when the fixed functions are nonlinear.

Subsection Connections

Unit 1 supplies matrix maps, products, and transposes. Unit 2 supplies eigenspace calculations through null spaces and explains nonuniqueness through rank. Unit 3 supplies gradients, Hessians, and critical points. Unit 4 supplies projection and residual orthogonality. Unit 6 relates constrained quadratic-form extrema to eigenvectors, principal directions, and singular values.