. The opening block reviews gradient descent and a final-layer update from Unit 3. The Unit 5 material interprets Hessian eigenvalues, checks least squares through residual and gradient conditions, distinguishes global minimality from coefficient uniqueness, and fits fixed nonlinear features.
For a quick reference on entrywise functions, design matrices, eigenvalue and least-squares commands, transposes, and numerical checks such as βclose to zeroβ, see the programming appendix sections B.4, B.8, B.10, and B.9.
The matrix being studied is the Hessian matrix. Its eigenvalues have opposite signs, so the quadratic form \(\mathbf{h}^T H\mathbf{h}\) takes both positive and negative values. If \(H\) is the Hessian at a critical point, the second derivative test classifies the critical point as a saddle point. The phrase βat a critical pointβ is important because the Hessian test classifies local behavior after the first-order term has disappeared.
The expression A.T @ r checks whether the residual is orthogonal to the columns of \(A\text{.}\) It can be close to zero even when r is not close to zero. Unit 4 used this same residual-orthogonality condition to derive the normal equations. Unit 5 obtains the same equations by setting the gradient of \(\|A\mathbf{x}-\mathbf{b}\|^2\) equal to zero.
The entries of evals are the eigenvalues \(1\) and \(3\text{.}\) The corresponding columns of Q are orthonormal eigenvectors. NumPy may return either sign for either eigenvector, so the signs of the displayed columns may change without changing the mathematics.
Because \(A\) is symmetric, its orthonormal eigenvectors form the columns of an orthogonal matrix \(Q\text{,}\) and \(Q^TAQ\) is the diagonal matrix of eigenvalues. The expression x @ (A @ x) computes the quadratic form \(q_A(\mathbf x)=\mathbf x^TA\mathbf x\text{,}\) which equals \(2\) for the displayed vector. Both eigenvalues are positive, so \(q_A\) is positive definite.