Preface Preface
This course is about linear algebra, together with the parts of multivariable calculus most relevant to optimization problems: given a quantity depending on multiple parameters, finding parameters that maximize or minimize that quantity. These tools are fundamental to data science, where one often builds models of data by choosing parameters that make a model best approximate a given data set. The course has one recurring theme: linear algebra studies linear maps; differential calculus studies nonlinear functions by replacing them locally with linear maps; optimization uses both to solve data problems.
The book is organized into seven units. Unit 1 introduces vectors, matrices, and linear maps. Unit 2 studies the anatomy of a linear map: systems, geometry, redundancy, rank, and forgotten directions. Unit 3 develops local linearization, directional derivatives, critical points, and gradient descent. Unit 4 develops orthogonality, projection, and least squares. Unit 5 develops eigenvalues, diagonalization, quadratic forms, Hessians, and least-squares applications. Unit 6 develops constrained optimization and Lagrange multipliers, then applies constrained quadratic forms to principal component analysis (PCA) and the singular value decomposition (SVD). Unit 7 uses abstract vector spaces and polynomial approximation to synthesize projection, least squares, Taylor approximation, and Gram-Schmidt.
Applications. The application threads are intentionally recurrent. Vectors, distance, and dot products introduce embeddings, nearest neighbors, and cosine similarity in Unit 1; dot products return through projection and least squares in Unit 4 and through eigenvector geometry and quadratic forms in Unit 5. Weighted linear combinations preview attention in Unit 1 and return through score-then-combine comparisons in Unit 4. Linear maps \(\mathbf{x}\mapsto A\mathbf{x}\) begin in Unit 1 and return through Units 2 and 3. Systems \(Ax=b\text{,}\) null spaces, rank, lines and planes, and local linearization connect solution sets, forgotten directions, tangent planes, and sensitivity across Units 2 and 3. Projection and least squares drive regression in Unit 4 and return in Units 5 and 7. Gradients and gradient descent organize first-order optimization in Unit 3; eigenvalues, quadratic forms, and Hessians organize second-order analysis in Unit 5; Lagrange multipliers, PCA, and SVD organize constraints and principal directions in Unit 6; polynomial inner products organize Taylor approximation and regression in Unit 7. The Transformer Wrap-Up Worksheet revisits token embeddings, attention scores, weighted averages, affine layers, nonlinear blocks, training loss, gradients, and low-rank updates. The course is not a course on large language models. The goal is to learn mathematical tools that are useful across many areas: data science, optimization, modeling, numerical computation, statistics, engineering, and machine learning.
How to use learning outcomes. Each unit begins with learning outcomes labeled U1-LO1, U1-LO2, and so on. Each outcome names a skill students should be able to use in a new problem. Toolbox skills are smaller calculations that support these outcomes. Activities, exercises, and the lab index identify the relevant outcomes using the same labels. For example, Learning outcomes. U4-LO4 identifies work on the Unit 4 least-squares residual outcome. Optional material is identified separately.
Outcome recurrence. Dot products and cosine similarity first appear in U1-LO1 and U1-LO2, then return in Units 4 and 5 through projection, least squares, eigenvector geometry, and quadratic forms. Matrix-vector products and composition first appear in U1-LO4 and U1-LO6, then return in Units 2 and 3 through linear maps, matrix actions, and local linearization. Reachable outputs and null directions first appear in U2-LO3 and U2-LO6, then return in Units 4 and 6 through least squares, forgotten directions, rank, and the four fundamental subspaces. Gradients and gradient descent first appear in Unit 3. Eigenvalues, quadratic forms, and Hessian classifications appear in Unit 5. Projection and least squares first appear in U4-LO4 and U4-LO5, then return in Units 5, 6, and 7 through regression, normal equations, PCA reconstruction, and polynomial approximation. Symmetric quadratic forms from U5-LO4 return in U6-LO4 and U6-LO5 through constrained extrema, maximum stretch, and PCA. SVD first appears in U6-LO6 and U6-LO7, connecting orthonormal input and output directions with matrix stretch, rank, the four fundamental subspaces, and PCA scores. Polynomial inner products and best approximation first appear in U7-LO3 and U7-LO5 and close the final synthesis through Taylor approximation, continuous least squares, and sampled least squares.
Labs. The notebooks are linked supplements to the textbook. The lab sequence includes Lab U1 on vectors, similarity, attention, and matrix actions; Lab U2 on auditing a linear map; Lab U3 on Jacobian matrices and local linearization; Lab U4 on regression as projection; Lab U5 on Hessian eigenvalues, least squares, and fixed nonlinear features, with a short review of Unit 3 gradient descent; Lab U6 on constraints, PCA, and SVD; and Lab U7 on polynomial approximation. The Transformer Wrap-Up Lab is an extension attached to the Transformer Wrap-Up Worksheet appendix. A detailed lab index appears at the end of the book.
