Unit 3 Beyond linear maps: local linearization and first-order optimization
Unit 1 and Unit 2 focused on linear maps \(\mathbf x\mapsto A\mathbf x\text{.}\) Most functions are not linear. This unit studies general scalar-valued and vector-valued functions, then uses derivatives and Jacobian matrices to recover linear maps that describe nonlinear functions near one point. For scalar-valued functions, the same local model also identifies directional change, critical points, and descent steps.
\begin{equation*}
F(\mathbf{a}+\mathbf{h})\approx F(\mathbf{a})+J_F(\mathbf{a})\mathbf{h}.
\end{equation*}
Big questions. If a function is not linear, what linear map best describes it near one point? How can that local model guide a step that decreases a scalar-valued function?
Learning outcomes. By the end of this unit, students should be able to:
-
U3-LO1. Classify, evaluate, and compose scalar-valued and vector-valued functions; write component functions and parametric descriptions of curves; and describe domains, ranges, graphs, level curves, and images of sets.
-
U3-LO2. Compute limits, check continuity, differentiate scalar-valued and vector-valued functions of one variable, and compute partial derivatives of multivariable functions.
-
U3-LO3. Compute and interpret partial derivatives, gradients, Hessians, and Jacobian matrices.
-
U3-LO4. Construct tangent planes and local linear approximations for scalar-valued functions.
-
U3-LO5. Use a Jacobian matrix as a local linear map for vector-valued functions: \(F(\mathbf{a}+\mathbf{h})\approx F(\mathbf{a})+J_F(\mathbf{a})\mathbf{h}\text{.}\)
-
U3-LO6. Apply the chain rule in component form and as matrix multiplication of Jacobian matrices.
-
U3-LO7. Interpret simple composed nonlinear systems, such as a tiny neural-network forward pass, by identifying affine pieces, nonlinear pieces, and the local linear approximation.
-
U3-LO8. Compute and interpret directional derivatives, and use the gradient to identify directions of greatest increase and decrease.
-
U3-LO9. Identify interior critical points as candidates for local extrema, take gradient-descent steps, and diagnose the effect of the learning rate from iterates, code, or loss tables.
Toolbox skills. Write component functions, compute partial derivatives and higher partial derivatives, compute gradients and Hessians, compute Jacobian matrices, write tangent plane equations, evaluate local linear approximations, check Jacobian matrix shapes, multiply Jacobian matrices in the chain rule, compare actual changes with linear predictions, find critical points, write critical-point equations, take gradient-descent steps, read a gradient-descent update in code, and compare learning rates from a loss table.
