Let \(A=\begin{bmatrix}\mathbf{a}_1 \amp \mathbf{a}_2 \amp \cdots \amp \mathbf{a}_n\end{bmatrix}\) be an \(m\times n\) matrix and let \(\mathbf{x}=(x_1,\ldots,x_n)\) be a column vector in \(\mathbb R^n\text{.}\) The matrix-vector product is
Write \(A=[\mathbf{a}_1\ \cdots\ \mathbf{a}_n]\text{,}\) where \(\mathbf{a}_j=(a_{1j},\ldots,a_{mj})\text{,}\) and let \(\mathbf{x}=(x_1,\ldots,x_n)\text{.}\) Since
We prove the first property. Write \(A=[\mathbf{a}_1\ \cdots\ \mathbf{a}_n]\text{,}\)\(\mathbf{x}=(x_1,\ldots,x_n)\text{,}\) and \(\mathbf{y}=(y_1,\ldots,y_n)\text{.}\) By the definition above,
Before forming \(A\mathbf{x}\text{,}\) check that the number of columns of \(A\) equals the number of entries of \(\mathbf{x}\text{.}\) If \(A\) is \(m\times n\) and \(\mathbf{x}\in\mathbb R^n\text{,}\) then \(A\mathbf{x}\in\mathbb R^m\text{:}\) its \(m\) entries come from the \(m\) row dot products above.
A matrix-vector product has two complementary readings. In the row view, each row measures the input by taking a dot product with \(\mathbf{x}\text{.}\) In the column view,
The same product \(A\mathbf{x}\) can represent different operations, depending on what the rows and columns of \(A\) mean. The next activities show three common patterns: scoring, selecting, and measuring changes.
Activity1.3.11.Document ranking in matrix code (U1-LO4, U1-LO8).
In ActivityΒ 1.1.24, we ranked documents by cosine similarity one score at a time. The following code repeats that ranking using a data matrix whose rows are the document vectors.
The rows of X are \(\mathbf{D}_1\text{,}\)\(\mathbf{D}_2\text{,}\) and \(\mathbf{D}_3\text{.}\) The product X @ q computes the three dot products with the query.
np.argsort(scores) returns the indices that put the scores in increasing order, rather than the sorted score values. NumPy indices start at 0. For scores [1, 0.5, 0], it returns [2, 1, 0]: the smallest score is at index 2, the next at index 1, and the largest at index 0.
[::-1] reverses that index array. A slice has the form [start:stop:step]; leaving the endpoints blank and using step -1 traverses the whole array backward. Thus [2, 1, 0] becomes [0, 1, 2], which lists the indices from largest score to smallest.
order = np.argsort(scores)[::-1]
ranking = doc_names[order]
Here order is [0, 1, 2], so ranking contains doc_names[0], doc_names[1], and doc_names[2], in that order. Thus the document names are arranged from highest score to lowest. If the scores changed, the indices could select the names in a different order.
The line s = K @ q computes the three dot products. The token βbirdβ receives the largest weight. The line out = alpha @ V forms the weighted average of the rows of V. The entries of alpha add to \(1\) because the scores were divided by their sum.
If \(A=\begin{bmatrix}\mathbf{a}_1 \amp \cdots \amp \mathbf{a}_n\end{bmatrix}\text{,}\) then \(A\mathbf{e}_j=\mathbf{a}_j\text{.}\) Applying a matrix to a standard basis vector picks out a column.
The coordinates of \(\mathbf{e}_j\) are all zero except for a \(1\) in position \(j\text{.}\) Thus the column formula for the matrix-vector product gives
Indeed, the columns of the identity matrix from DefinitionΒ 1.2.17 are \(\mathbf{e}_1,\ldots,\mathbf{e}_n\text{.}\) Therefore, writing \(\mathbf{x}=(x_1,\ldots,x_n)\text{,}\)
SubsectionGeometric matrix actions in \(\mathbb R^2\text{:}\) building intuition
In \(\mathbb R^2\text{,}\) a matrix \(A=\begin{bmatrix}\mathbf{a}_1 \amp \mathbf{a}_2\end{bmatrix}\) sends \(\mathbf{e}_1\) to \(\mathbf{a}_1\) and \(\mathbf{e}_2\) to \(\mathbf{a}_2\text{.}\) Every vector \(\mathbf{x}=(x_1,x_2)\) can be written as \(x_1\mathbf{e}_1+x_2\mathbf{e}_2\text{,}\) since
The second equality uses linearity, and the last uses FactΒ 1.3.15. Thus a \(2\times 2\) matrix is geometrically determined by where it sends the two coordinate directions.
To understand a \(2\times 2\) matrix \(A\text{,}\) apply it to the four corners of the unit square. The image is usually a parallelogram. The two edges leaving the origin are the columns of \(A\text{.}\) We will refer to this as the unit-square visualization.
Three panels appear from left to right: the original unit square, the two transformed basis vectors, and the image parallelogram. In the left panel, a red horizontal arrow marks \(\mathbf e_1\) and a green vertical arrow marks \(\mathbf e_2\text{.}\) In the middle panel, a red arrow marks \(A\mathbf e_1\) and a green arrow marks \(A\mathbf e_2\text{;}\) in the right panel, those two arrows form adjacent edges of the sheared parallelogram. Blue dots label the corner \((1,1)\) in the left panel and its image \((2,1)\) in the right panel.
Figure1.3.17.A \(2\times 2\) matrix sends the unit square to the parallelogram determined by its two columns. This is the geometric version of columns contribute.
The left panel shows the unit square with a red horizontal basis arrow and a green vertical basis arrow. The middle panel shows the columns of \(A=\begin{bmatrix}2\amp0\\0\amp1\end{bmatrix}\text{:}\) the red arrow reaches \((2,0)\) and the green arrow reaches \((0,1)\text{.}\) The right panel shows the image rectangle with corners \((0,0)\text{,}\)\((2,0)\text{,}\)\((2,1)\text{,}\) and \((0,1)\text{.}\) The width doubles and the height stays the same.
Figure1.3.18.A horizontal stretch by a factor of \(2\) sends the unit square to a rectangle twice as wide. The first column doubles the horizontal basis vector, while the second column leaves the vertical basis vector unchanged.
Let \(A=\begin{bmatrix}1\amp1\\0\amp1\end{bmatrix}\text{.}\) Compute \(A\mathbf{e}_1\text{,}\)\(A\mathbf{e}_2\text{,}\) and \(A\mathbf{x}\text{,}\) where \(\mathbf{x}=(1,1)\text{.}\)
Activity1.3.22.Reflection across the \(y\)-axis (U1-LO4).
Let \(A=\begin{bmatrix}-1\amp0\\0\amp1\end{bmatrix}\text{.}\) Compute \(A\mathbf{e}_1\text{,}\)\(A\mathbf{e}_2\text{,}\)\(A\mathbf{x}\text{,}\) and \(A\mathbf{y}\text{,}\) where \(\mathbf{x}=(1,1)\) and \(\mathbf{y}=(0,1)\text{.}\)
The left panel shows the unit square with a red horizontal basis arrow and a green vertical basis arrow. The middle panel shows the columns of \(A=\begin{bmatrix}-1\amp0\\0\amp1\end{bmatrix}\text{:}\) the red arrow points left to \((-1,0)\) and the green arrow points up to \((0,1)\text{.}\) The right panel shows the reflected square with corners \((0,0)\text{,}\)\((-1,0)\text{,}\)\((-1,1)\text{,}\) and \((0,1)\text{.}\) Its labeled arrows preserve the correspondence with the original basis vectors.
Figure1.3.23.Reflection across the \(y\)-axis sends the unit square to its left-hand mirror image. The first column reverses the horizontal direction, while the second column leaves the vertical direction unchanged.
Thus the two coordinates are exchanged. If \(x=y\text{,}\) exchanging the coordinates does not move the point, so every point on the line \(y=x\) stays fixed.
The left panel shows the unit square with a red arrow to \(\mathbf e_1=(1,0)\) and a green arrow to \(\mathbf e_2=(0,1)\text{.}\) The middle panel shows the columns of \(A=\begin{bmatrix}0\amp1\\1\amp0\end{bmatrix}\text{:}\) the red arrow \(A\mathbf e_1\) points up and the green arrow \(A\mathbf e_2\) points right. The right panel fills the same unit square, with the red edge now vertical and the green edge horizontal. The dashed line \(y=x\) passes through the two fixed corners, \((0,0)\) and \((1,1)\text{.}\)
Figure1.3.25.The unit square occupies the same region after reflection across \(y=x\text{,}\) even though its horizontal and vertical coordinate directions have exchanged, so an unlabeled image of the square can appear unchanged. The labeled red and green edges show the exchange. The corners \((0,0)\) and \((1,1)\) on the dashed mirror line stay fixed.
The domain of \(F\) is its set of allowed inputs, \(\mathbb R^n\text{,}\) and its co-domain is the specified set containing its outputs, \(\mathbb R^m\text{.}\) A matrix gives one important source of such functions by the rule \(\mathbf{x}\mapsto A\mathbf{x}\text{.}\)
Note1.3.28.Linear maps preserve linear combinations.
If \(T:\mathbb R^n\to\mathbb R^m\) is linear, then for any vectors \(\mathbf{v}_1,\ldots,\mathbf{v}_k\in\mathbb R^n\) and scalars \(c_1,\ldots,c_k\in\mathbb R\text{,}\)
Thus a linear combination of inputs is sent to the linear combination of their outputs with the same coefficients. This follows by repeatedly applying the two linearity rules.
Every vector \(\mathbf{x}=(x_1,\ldots,x_n)\in\mathbb R^n\) can be written as a linear combination of the standard basis vectors. Indeed, \(x_j\mathbf{e}_j\) has entry \(x_j\) in position \(j\) and zeros elsewhere, so adding these vectors coordinate by coordinate gives
Since a linear map preserves linear combinations, its values on \(\mathbf{e}_1,\ldots,\mathbf{e}_n\) determine its value on every input. This observation leads to the following theorem.
Every linear map \(T:\mathbb R^n\to\mathbb R^m\) is given by multiplication by a unique matrix. The columns of that matrix are \(T(\mathbf{e}_1),\ldots,T(\mathbf{e}_n)\text{.}\)
Using the standard-basis expansion above, let \(A=[\,T(\mathbf{e}_1)\ \cdots\ T(\mathbf{e}_n)\,]\text{.}\) By linearity and the column view of matrix-vector multiplication,
The second equality uses linearity, and the third uses the column view of matrix-vector multiplication. This matrix is unique because its \(j\)-th column must be \(A\mathbf{e}_j=T(\mathbf{e}_j)\text{.}\)
Activity1.3.31.Images of basis vectors determine the matrix (U1-LO5).
Suppose a linear map \(T:\mathbb R^2\to\mathbb R^2\) satisfies \(T(\mathbf{e}_1)=(2,1)\) and \(T(\mathbf{e}_2)=(-1,3)\text{.}\) Find the matrix \(A\) such that \(T(\mathbf{x})=A\mathbf{x}\text{.}\)
An affine map is linear only when \(\mathbf{b}=\mathbf{0}\text{.}\) A quick test is the zero vector: every linear map sends \(\mathbf{0}\) to \(\mathbf{0}\text{,}\) but
Three panels share the same coordinate scale. The left panel shows the unit square with red and green standard basis arrows and the corner \((1,1)\) marked in blue. The middle panel shows the linear image under \(A=\begin{bmatrix}2\amp0\\0\amp1\end{bmatrix}\text{:}\) a rectangle with corners \((0,0)\text{,}\)\((2,0)\text{,}\)\((2,1)\text{,}\) and \((0,1)\text{.}\) The right panel shows its translation by \(\mathbf b=(1,1)\text{,}\) with corners \((1,1)\text{,}\)\((3,1)\text{,}\)\((3,2)\text{,}\) and \((1,2)\text{.}\) A purple arrow from the origin to \((1,1)\) marks the translation vector and the image of zero. The red and green edge arrows now start at \(\mathbf b\text{,}\) and the blue corner is labeled \((3,2)\text{.}\)
Figure1.3.34.The horizontal stretch from FigureΒ 1.3.18 followed by translation by \(\mathbf b=(1,1)\text{.}\) The translation shifts every point one unit right and one unit up, preserving the stretched rectangleβs shape. In particular, \(F(\mathbf 0)=\mathbf b\ne\mathbf 0\text{,}\) so this affine map is not linear.
The symbol \(\mapsto\) means βmaps toβ: the input vector \(\mathbf{x}\) is sent to the output vector \(W\mathbf{x}+\mathbf{b}\text{.}\) The matrix \(W\) mixes the input coordinates. The vector \(\mathbf{b}\) shifts the result. Since