Section 3.9 Unit 3 highlights
Subsection Mathematical quick reference
Functions, domains, and local language.
Function types. A scalar-valued function has the form \(f:D\subseteq\mathbb R^n\to\mathbb R\text{.}\) A vector-valued function \(F:D\subseteq\mathbb R^n\to\mathbb R^m\) has scalar component functions \(F_1,\ldots,F_m\text{.}\) The domain consists of allowed inputs; the range consists of attained outputs.
Graph and level set. The graph of \(f\) is \(\{(\mathbf x,f(\mathbf x)):\mathbf x\in D\}\text{;}\) the level set at \(c\) is \(\{\mathbf x\in D:f(\mathbf x)=c\}\text{.}\)
Open ball. \(B_r(\mathbf a)=\{\mathbf x:\|\mathbf x-\mathbf a\|\lt r\}\text{,}\) with \(r>0\text{.}\) A point is interior to \(D\) if some ball about it lies in \(D\text{;}\) it is a boundary point if every ball meets both \(D\) and its complement. A set is closed if it contains all its boundary points, and bounded if it lies in some ball.
Limit. At an accumulation point of the domain, \(F(\mathbf x)\to\mathbf L\) as \(\mathbf x\to\mathbf a\) means that for every \(\epsilon>0\) there is \(\delta>0\) such that
\begin{equation*}
\mathbf x\in D,\quad 0\lt\|\mathbf x-\mathbf a\|\lt\delta\quad\Longrightarrow\quad\|F(\mathbf x)-\mathbf L\|\lt\epsilon.
\end{equation*}
Limits of vector-valued functions are computed componentwise. Two paths giving different limits disprove existence; agreement along a few paths does not prove existence.
Continuity. At an accumulation point \(\mathbf a\in D\text{,}\) continuity means \(\lim_{\mathbf x\to\mathbf a}F(\mathbf x)=F(\mathbf a)\text{,}\) with inputs restricted to \(D\text{.}\) Continuity is automatic at isolated domain points.
Derivatives and their dimensions.
One-parameter derivative. For \(\mathbf r(t)=(r_1(t),\ldots,r_m(t))\text{,}\)
\begin{equation*}
\mathbf r'(t)=(r_1'(t),\ldots,r_m'(t)).
\end{equation*}
When \(\mathbf r'(a)\neq\mathbf0\text{,}\) the tangent line is \(\mathbf r(a)+s\mathbf r'(a)\text{.}\)
Partial derivative. Differentiate in one coordinate while fixing the others:
\begin{equation*}
f_{x_j}(\mathbf a)=\lim_{h\to0}\frac{f(\mathbf a+h\mathbf e_j)-f(\mathbf a)}h.
\end{equation*}
Gradient and Hessian. For a scalar-valued function of \(n\) variables,
\begin{equation*}
\nabla f=\begin{bmatrix}f_{x_1}\\\vdots\\f_{x_n}\end{bmatrix}\in\mathbb R^n,\qquad H_f=[f_{x_i x_j}]\in\mathbb R^{n\times n}.
\end{equation*}
Continuous second partial derivatives ensure equality of mixed partials and symmetry of the Hessian.
Jacobian matrix. For \(F:\mathbb R^n\to\mathbb R^m\text{,}\)
\begin{equation*}
J_F=\left[\frac{\partial F_i}{\partial x_j}\right]\in\mathbb R^{m\times n},\qquad J_f=(\nabla f)^T.
\end{equation*}
Row \(i\) differentiates component \(F_i\text{;}\) column \(j\) describes change in input direction \(\mathbf e_j\text{.}\)
Differentiability and local approximation.
Differentiability. At an interior point \(\mathbf a\text{,}\) the derivative is a linear map whose error is small relative to the input change:
\begin{equation*}
\lim_{\mathbf h\to\mathbf0}\frac{\|F(\mathbf a+\mathbf h)-F(\mathbf a)-J_F(\mathbf a)\mathbf h\|}{\|\mathbf h\|}=0.
\end{equation*}
Continuous first partial derivatives near \(\mathbf a\) are sufficient for differentiability. Existence of partial derivatives alone is not sufficient.
Local prediction.
\begin{equation*}
F(\mathbf a+\mathbf h)\approx F(\mathbf a)+J_F(\mathbf a)\mathbf h,\qquad f(\mathbf a+\mathbf h)\approx f(\mathbf a)+\nabla f(\mathbf a)\cdot\mathbf h.
\end{equation*}
Tangent plane. For a differentiable \(z=f(x,y)\) at \((a,b)\text{,}\)
\begin{equation*}
z=f(a,b)+f_x(a,b)(x-a)+f_y(a,b)(y-b).
\end{equation*}
A normal vector is \((f_x(a,b),f_y(a,b),-1)\text{.}\)
First-order zero change. If \(\mathbf h\in\operatorname{null}(J_F(\mathbf a))\text{,}\) the predicted first-order output change is zero; the actual nonlinear change need not be zero.
Chain rule.
Composition. If \(G:\mathbb R^n\to\mathbb R^p\) is differentiable at \(\mathbf a\) and \(F:\mathbb R^p\to\mathbb R^m\) is differentiable at \(G(\mathbf a)\text{,}\) then
\begin{equation*}
J_{F\circ G}(\mathbf a)=J_F(G(\mathbf a))J_G(\mathbf a).
\end{equation*}
The dimensions are \((m\times p)(p\times n)=m\times n\text{;}\) the right factor acts first.
Scalar function along a curve.
\begin{equation*}
\frac{d}{dt}f(\mathbf r(t))=\nabla f(\mathbf r(t))\cdot\mathbf r'(t).
\end{equation*}
Directional change and first-order optimization.
Directional derivative. For differentiable \(f\) and a unit direction \(\mathbf u\text{,}\)
\begin{equation*}
D_{\mathbf u}f(\mathbf a)=\nabla f(\mathbf a)\cdot\mathbf u.
\end{equation*}
Normalize a proposed nonzero direction before using the formula.
Steepest directions. If \(\nabla f(\mathbf a)\neq\mathbf0\text{,}\) the greatest and least directional derivatives are \(\|\nabla f(\mathbf a)\|\) and \(-\|\nabla f(\mathbf a)\|\text{,}\) attained in directions \(\pm\nabla f(\mathbf a)/\|\nabla f(\mathbf a)\|\text{.}\) If the gradient is zero, every directional derivative is zero.
Local and global extrema. A local minimum compares values in a neighborhood; a global minimum compares all values on the domain. Reverse the inequalities for maxima.
Critical point. In this book, a differentiable point \(\mathbf a\) with \(\nabla f(\mathbf a)=\mathbf0\text{.}\) An interior differentiable local extremum must be critical. The converse fails; boundary points and points where differentiability fails require separate checks.
Gradient descent. Starting from \(\mathbf x_0\text{,}\) with step size \(\alpha>0\text{,}\)
\begin{equation*}
\mathbf x_{k+1}=\mathbf x_k-\alpha\nabla f(\mathbf x_k).
\end{equation*}
For a differentiable function with a nonzero gradient, a sufficiently small step in the negative-gradient direction decreases the function. Arbitrary step sizes need not decrease it or yield convergence.
Subsection Common mistakes
Confusing domain with range; using a few paths to prove a multivariable limit; assuming partial derivatives guarantee differentiability; treating a local approximation as exact; confusing \(\mathbf h\) with \(\mathbf a+\mathbf h\text{;}\) reversing chain-rule factors or evaluating them at the wrong point; using a nonunit direction; treating every critical point as an extremum; ignoring boundaries; or assuming every gradient-descent step decreases the function.
