. The main sequence practices symbolic Jacobian matrices, evaluating a Jacobian matrix at a point, local linear prediction, square-grid visualizations, checking Jacobian matrix shapes in the chain rule, and short nonlinear model-block computations.
For a quick reference on array shapes, matrix products, numerical checks, and SymPy symbolic commands, see the programming appendix sections B.2, B.4, B.9, and B.11.
The component functions are \(F_1(x,y)=xy\) and \(F_2(x,y)=x^2+y\text{.}\) The Jacobian matrix is \(\begin{bmatrix}y&x\\2x&1\end{bmatrix}\text{,}\) so it is \(2\times 2\text{.}\) Its first column measures how the two output components change with respect to the first input variable \(x\text{.}\)
Activity3.7.2.Actual output versus local prediction (U3-LO5).
def F(v):
x, y = v
return np.array([x**2 + y, x - y**2])
a = np.array([1.0, 2.0])
h = np.array([0.1, -0.2])
J_at_a = np.array([[2.0, 1.0],
[1.0, -4.0]])
actual = F(a + h)
linear = F(a) + J_at_a @ h
actual, linear, actual - linear
The true value is actual, which is [3.01, -2.14]. The local linear prediction is linear, which is [3.0, -2.1]. The difference actual - linear is [0.01, -0.04], the local approximation error for this input change.
The input dimension is \(3\) and the output dimension is \(2\text{.}\) The slice J_at_a[:, 0] selects the first column of the Jacobian matrix, [2.0, 0.0]. For the input change [0.1, 0.0, 0.0], the predicted output change is [0.2, 0.0].
The map whose Jacobian matrix is Jg is applied first. Since Jg is \(3\times 2\text{,}\) it takes two input directions to three intermediate directions; Jf_at_g then takes those three directions to four output directions. The product has shape \(4\times 2\text{,}\) and the displayed output is a \(4\times 2\) array whose entries are all \(3\text{.}\)
The line J_at_x = W2 @ D @ W1 computes the Jacobian matrix of the block at x by the chain rule for Jacobian matrices. The shape of J_at_x is \(2\times 2\text{:}\) the input has two coordinates and the output has two coordinates.
The learning rate is alpha. The vector grad_f(x) must have the same shape as x, since the update subtracts one vector from another. If the minus sign were replaced by a plus sign, the update would move in the direction of steepest increase instead of steepest decrease. The loop does not prove that a minimum has been found. It only describes the repeated update; convergence depends on the function, starting point, learning rate, and stopping rule.
The learning rate \(\alpha=0.05\) is making slow but steady progress. The learning rate \(\alpha=0.20\) is making faster useful progress. The learning rate \(\alpha=1.05\) appears unstable because the loss is increasing. A decreasing loss table does not prove that the global minimum has been found. It only shows what happened for the displayed iterates.
Activity3.7.8.Reading a next-token loss (U3-LO8, U3-LO9).
A language model produces scores for possible next tokens, converts those scores into probabilities, and uses a loss to measure the probability assigned to the observed next token. A simplified training objective for a sequence \(t_1,\ldots,t_L\) can be written
The parameters are collected in \(\boldsymbol{\theta}\text{.}\) The objective \(\mathcal L(\boldsymbol{\theta})\) is the loss. The learning rate is \(\alpha\text{.}\) The gradient \(\nabla\mathcal L(\boldsymbol{\theta}_k)\) gives the local direction of steepest increase, so the negative gradient is used for descent. This is an optimization problem because training means adjusting parameters to reduce a loss.
with a fixed matrix \(A\) and a single trained coefficient vector \(\mathbf c\text{.}\) It is a general differentiable loss optimized by gradient descent.
If \(\mathbf g\in\mathbb R^m\) and \(\mathbf h\in\mathbb R^d\text{,}\) then this outer product has shape \(m\times d\text{.}\) Every column is a scalar multiple of \(\mathbf g\text{,}\) so its rank is at most one. The update changes \(W\) in the negative-gradient direction for this training example.
If \(\mathbf g\) has length \(m\) and \(\mathbf h\) has length \(d\text{,}\) then G has shape \(m\times d\text{.}\) The command np.outer(g, h) forms the outer product even when the two vectors are stored as one-dimensional NumPy arrays. The second line represents
\begin{equation*}
W_{\mathrm{new}}
=
W-\alpha\nabla_W\mathcal L.
\end{equation*}