Skip to main content

MATH 345: Linear Algebra and Optimization

Section 1.3 Matrix-vector product and linear maps

Subsection Matrix-vector product

A vector can be written as a row or as a column. When a matrix acts on a vector in this course, we usually write the vector as a column:
\begin{equation*} \mathbf{x}=\begin{bmatrix}x_1\\x_2\\\vdots\\x_n\end{bmatrix}. \end{equation*}
From now on, when we write \(\mathbf{x}=(x_1,\ldots,x_n)\text{,}\) \(\mathbf{x}\) is interpreted as a column vector.
If \(A\) is an \(m\times n\) matrix, then \(A\) has \(n\) columns. We write
\begin{equation*} A=\begin{bmatrix}\mathbf{a}_1 \amp \mathbf{a}_2 \amp \cdots \amp \mathbf{a}_n\end{bmatrix} \end{equation*}
where each \(\mathbf{a}_j\) is a column vector in \(\mathbb R^m\text{.}\)

Definition 1.3.1. Matrix-vector product.

Let \(A=\begin{bmatrix}\mathbf{a}_1 \amp \mathbf{a}_2 \amp \cdots \amp \mathbf{a}_n\end{bmatrix}\) be an \(m\times n\) matrix and let \(\mathbf{x}=(x_1,\ldots,x_n)\) be a column vector in \(\mathbb R^n\text{.}\) The matrix-vector product is
\begin{equation*} A\mathbf{x}=x_1\mathbf{a}_1+x_2\mathbf{a}_2+\cdots+x_n\mathbf{a}_n. \end{equation*}

Activity 1.3.2. Computing a matrix-vector product (U1-LO4).

Compute \(A\mathbf{x}\) where
\begin{equation*} A=\begin{bmatrix}-1 \amp 4 \amp -5\\ 3 \amp 1 \amp -2\end{bmatrix}, \qquad \mathbf{x}=\begin{bmatrix}2\\-3\\4\end{bmatrix}. \end{equation*}
Solution.
Using columns,
\begin{align*} A\mathbf{x} \amp=2\begin{bmatrix}-1\\3\end{bmatrix}-3\begin{bmatrix}4\\1\end{bmatrix}+4\begin{bmatrix}-5\\-2\end{bmatrix}\\ \amp=\begin{bmatrix}2(-1)-3(4)+4(-5)\\2(3)-3(1)+4(-2)\end{bmatrix}\\ \amp=\begin{bmatrix}-34\\-5\end{bmatrix}. \end{align*}

Why is this true?.

Write \(A=[\mathbf{a}_1\ \cdots\ \mathbf{a}_n]\text{,}\) where \(\mathbf{a}_j=(a_{1j},\ldots,a_{mj})\text{,}\) and let \(\mathbf{x}=(x_1,\ldots,x_n)\text{.}\) Since
\begin{equation*} A\mathbf{x}=x_1\mathbf{a}_1+\cdots+x_n\mathbf{a}_n, \end{equation*}
its \(i\)-th entry is
\begin{equation*} x_1a_{i1}+\cdots+x_na_{in}=\operatorname{row}_i(A)\cdot\mathbf{x}. \end{equation*}
Stacking these entries gives the stated formula.

Activity 1.3.4. Computing the same product using rows (U1-LO4).

We revisit ActivityΒ 1.3.2 using row dot products. Compute \(A\mathbf{x}\) where
\begin{equation*} A=\begin{bmatrix}-1 \amp 4 \amp -5\\ 3 \amp 1 \amp -2\end{bmatrix}, \qquad \mathbf{x}=\begin{bmatrix}2\\-3\\4\end{bmatrix}. \end{equation*}
Solution.
Take the dot product of each row of \(A\) with \(\mathbf{x}\text{:}\)
\begin{align*} A\mathbf{x} \amp=\begin{bmatrix}(-1,4,-5)\cdot(2,-3,4)\\(3,1,-2)\cdot(2,-3,4)\end{bmatrix}\\ \amp=\begin{bmatrix}(-1)(2)+4(-3)+(-5)(4)\\3(2)+1(-3)+(-2)(4)\end{bmatrix}\\ \amp=\begin{bmatrix}-34\\-5\end{bmatrix}. \end{align*}
Each dot product gives one component of the same output vector obtained using columns.

Why is this true?.

We prove the first property. Write \(A=[\mathbf{a}_1\ \cdots\ \mathbf{a}_n]\text{,}\) \(\mathbf{x}=(x_1,\ldots,x_n)\text{,}\) and \(\mathbf{y}=(y_1,\ldots,y_n)\text{.}\) By the definition above,
\begin{align*} A(\mathbf{x}+\mathbf{y}) \amp= (x_1+y_1)\mathbf{a}_1+\cdots+(x_n+y_n)\mathbf{a}_n\\ \amp= A\mathbf{x}+A\mathbf{y}. \end{align*}

Note 1.3.6. Shape habit for matrix-vector products.

Before forming \(A\mathbf{x}\text{,}\) check that the number of columns of \(A\) equals the number of entries of \(\mathbf{x}\text{.}\) If \(A\) is \(m\times n\) and \(\mathbf{x}\in\mathbb R^n\text{,}\) then \(A\mathbf{x}\in\mathbb R^m\text{:}\) its \(m\) entries come from the \(m\) row dot products above.

Note 1.3.7. Rows measure; columns contribute.

A matrix-vector product has two complementary readings. In the row view, each row measures the input by taking a dot product with \(\mathbf{x}\text{.}\) In the column view,
\begin{equation*} A\mathbf{x}=x_1\mathbf{a}_1+\cdots+x_n\mathbf{a}_n, \end{equation*}
so the input coordinates tell how much each column contributes to the output.

Subsubsection Examples of matrix-vector products

The same product \(A\mathbf{x}\) can represent different operations, depending on what the rows and columns of \(A\) mean. The next activities show three common patterns: scoring, selecting, and measuring changes.
Activity 1.3.8. Scoring objects by features (U1-LO4).
Let the rows of \(X\) represent three objects with two features:
\begin{equation*} X=\begin{bmatrix}3\amp 1\\1\amp 4\\2\amp 2\end{bmatrix}, \qquad \mathbf{w}=(2,-1). \end{equation*}
Compute \(X\mathbf{w}\text{.}\) Interpret the result.
Solution.
We have
\begin{equation*} X\mathbf{w}= \begin{bmatrix}3\amp 1\\1\amp 4\\2\amp 2\end{bmatrix} \begin{bmatrix}2\\-1\end{bmatrix} = \begin{bmatrix}5\\-2\\2\end{bmatrix}. \end{equation*}
Each output is a score:
\begin{equation*} 2(\text{first feature})-(\text{second feature}). \end{equation*}
Activity 1.3.9. A selector matrix (U1-LO4, U1-LO5).
Let
\begin{equation*} S=\begin{bmatrix}1\amp0\amp0\amp0\\0\amp0\amp1\amp0\end{bmatrix}, \qquad \mathbf{x}=\begin{bmatrix}5\\6\\7\\8\end{bmatrix}. \end{equation*}
Compute \(S\mathbf{x}\text{.}\) What does \(S\) do?
Solution.
We have
\begin{equation*} \begin{aligned} S\mathbf{x} \amp= \begin{bmatrix} 1\cdot 5+0\cdot 6+0\cdot 7+0\cdot 8\\ 0\cdot 5+0\cdot 6+1\cdot 7+0\cdot 8 \end{bmatrix}\\ \amp= \begin{bmatrix}5\\7\end{bmatrix}. \end{aligned} \end{equation*}
The matrix \(S\) selects the first and third entries of \(\mathbf{x}\text{.}\)
Activity 1.3.10. A difference matrix (U1-LO4, U1-LO5).
Let
\begin{equation*} D=\begin{bmatrix}-1\amp 1\amp 0\\0\amp -1\amp 1\end{bmatrix}, \qquad \mathbf{x}=\begin{bmatrix}2\\5\\9\end{bmatrix}. \end{equation*}
Compute \(D\mathbf{x}\text{.}\) What does \(D\) measure?
Solution.
We have
\begin{equation*} \begin{aligned} D\mathbf{x} \amp= \begin{bmatrix} -1\cdot 2+1\cdot 5+0\cdot 9\\ 0\cdot 2+(-1)\cdot 5+1\cdot 9 \end{bmatrix}\\ \amp= \begin{bmatrix}3\\4\end{bmatrix}. \end{aligned} \end{equation*}
The entries are consecutive differences:
\begin{equation*} 5-2=3,\qquad 9-5=4. \end{equation*}
Thus \(D\) measures changes in a short time series.
Activity 1.3.11. Document ranking in matrix code (U1-LO4, U1-LO8).
In ActivityΒ 1.1.24, we ranked documents by cosine similarity one score at a time. The following code repeats that ranking using a data matrix whose rows are the document vectors.
import numpy as np

doc_names = np.array(["D1", "D2", "D3"])

X = np.array([
    [3., 0., 3.],
    [1., 1., 0.],
    [0., 2., 0.],
])

q = np.array([1., 0., 1.])

scores = (X @ q) / (
    np.linalg.norm(X, axis=1) * np.linalg.norm(q)
)
ranking = doc_names[np.argsort(scores)[::-1]]

scores, ranking
Output:
(array([1. , 0.5, 0. ]),
 array(['D1', 'D2', 'D3']))
  1. What does the second row of X represent?
  2. Which expression computes the three dot products \(\mathbf{D}_i\cdot\mathbf{q}\text{?}\)
  3. What does np.argsort(scores)[::-1] do?
Solution.
The rows of X are \(\mathbf{D}_1\text{,}\) \(\mathbf{D}_2\text{,}\) and \(\mathbf{D}_3\text{.}\) The product X @ q computes the three dot products with the query.
To understand np.argsort(scores)[::-1], read it in two steps:
  1. np.argsort(scores) returns the indices that put the scores in increasing order, rather than the sorted score values. NumPy indices start at 0. For scores [1, 0.5, 0], it returns [2, 1, 0]: the smallest score is at index 2, the next at index 1, and the largest at index 0.
  2. [::-1] reverses that index array. A slice has the form [start:stop:step]; leaving the endpoints blank and using step -1 traverses the whole array backward. Thus [2, 1, 0] becomes [0, 1, 2], which lists the indices from largest score to smallest.
The assignment to ranking uses these indices to select entries from doc_names. We could write the same calculation in two lines:
order = np.argsort(scores)[::-1]
ranking = doc_names[order]
Here order is [0, 1, 2], so ranking contains doc_names[0], doc_names[1], and doc_names[2], in that order. Thus the document names are arranged from highest score to lowest. If the scores changed, the indices could select the names in a different order.
The displayed output array(['D1', 'D2', 'D3']) describes a one-dimensional NumPy array containing three text strings.
The scores agree with ActivityΒ 1.1.24: \(1,\frac12,0\text{.}\)
Example 1.3.12. The attention activity in matrix form.
In ActivityΒ 1.1.29, we computed the scores
\begin{equation*} \mathbf{q}\cdot\mathbf{k}_{\mathrm{small}},\qquad \mathbf{q}\cdot\mathbf{k}_{\mathrm{red}},\qquad \mathbf{q}\cdot\mathbf{k}_{\mathrm{bird}} \end{equation*}
one at a time. Matrix-vector multiplication computes the same scores at once.
Put the key vectors as the rows of
\begin{equation*} K=\begin{bmatrix} \mathbf{k}_{\mathrm{small}}^T\\ \mathbf{k}_{\mathrm{red}}^T\\ \mathbf{k}_{\mathrm{bird}}^T \end{bmatrix}. \end{equation*}
Then
\begin{equation*} K\mathbf{q}= \begin{bmatrix} \mathbf{k}_{\mathrm{small}}\cdot\mathbf{q}\\ \mathbf{k}_{\mathrm{red}}\cdot\mathbf{q}\\ \mathbf{k}_{\mathrm{bird}}\cdot\mathbf{q} \end{bmatrix}. \end{equation*}
For the numbers in ActivityΒ 1.1.29,
\begin{equation*} K=\begin{bmatrix}1\amp0\\0\amp1\\1\amp1\end{bmatrix}, \qquad \mathbf{q}=\begin{bmatrix}1\\1\end{bmatrix}, \end{equation*}
so
\begin{equation*} K\mathbf{q}=\begin{bmatrix}1\\1\\2\end{bmatrix}. \end{equation*}
Activity 1.3.13. The same attention calculation in code (U1-LO4, U1-LO8).
The following code performs the calculation from the token attention activity.
import numpy as np

tokens = np.array(["small", "red", "bird"])

q = np.array([1., 1.])

K = np.array([
    [1., 0.],
    [0., 1.],
    [1., 1.],
])

V = np.array([
    [4., 0.],
    [0., 4.],
    [4., 4.],
])

s = K @ q
alpha = s / s.sum()
out = alpha @ V

s, alpha, out
Output:
(array([1., 1., 2.]),
 array([0.25, 0.25, 0.5 ]),
 array([3., 3.]))
  1. Which line computes the token scores?
  2. Which token receives the largest weight?
  3. Which line forms the weighted average?
  4. Why do the entries of alpha add to \(1\text{?}\)
Solution.
The line s = K @ q computes the three dot products. The token β€œbird” receives the largest weight. The line out = alpha @ V forms the weighted average of the rows of V. The entries of alpha add to \(1\) because the scores were divided by their sum.
Definition 1.3.14. Standard basis.
The \(j\)-th standard basis vector \(\mathbf{e}_j\) is the column vector with a \(1\) in position \(j\) and zeros elsewhere.
The term β€œbasis” will be explained later in the course. For now, use the concrete description of \(\mathbf{e}_j\) above.
Why is this true?.
The coordinates of \(\mathbf{e}_j\) are all zero except for a \(1\) in position \(j\text{.}\) Thus the column formula for the matrix-vector product gives
\begin{equation*} A\mathbf{e}_j=\sum_{k=1}^{n}(\mathbf{e}_j)_k\mathbf{a}_k=\mathbf{a}_j. \end{equation*}
Here \((\mathbf{e}_j)_k\) denotes the \(k\)-th coordinate of \(\mathbf{e}_j\text{,}\) which equals \(1\) when \(k=j\) and \(0\) otherwise.
Example 1.3.16. The identity matrix leaves vectors unchanged.
For every column vector \(\mathbf{x}\in\mathbb R^n\text{,}\) \(I_n\mathbf{x}=\mathbf{x}\text{.}\)
Indeed, the columns of the identity matrix from DefinitionΒ 1.2.17 are \(\mathbf{e}_1,\ldots,\mathbf{e}_n\text{.}\) Therefore, writing \(\mathbf{x}=(x_1,\ldots,x_n)\text{,}\)
\begin{equation*} I_n\mathbf{x}=x_1\mathbf{e}_1+\cdots+x_n\mathbf{e}_n=\begin{bmatrix}x_1\\\vdots\\x_n\end{bmatrix}=\mathbf{x}. \end{equation*}

Subsection Geometric matrix actions in \(\mathbb R^2\text{:}\) building intuition

In \(\mathbb R^2\text{,}\) a matrix \(A=\begin{bmatrix}\mathbf{a}_1 \amp \mathbf{a}_2\end{bmatrix}\) sends \(\mathbf{e}_1\) to \(\mathbf{a}_1\) and \(\mathbf{e}_2\) to \(\mathbf{a}_2\text{.}\) Every vector \(\mathbf{x}=(x_1,x_2)\) can be written as \(x_1\mathbf{e}_1+x_2\mathbf{e}_2\text{,}\) since
\begin{equation*} x_1\mathbf{e}_1+x_2\mathbf{e}_2=x_1\begin{bmatrix}1\\0\end{bmatrix}+x_2\begin{bmatrix}0\\1\end{bmatrix}=\begin{bmatrix}x_1\\x_2\end{bmatrix}=\mathbf{x}. \end{equation*}
Therefore,
\begin{align*} A\mathbf{x} \amp=A(x_1\mathbf{e}_1+x_2\mathbf{e}_2)\\ \amp=x_1A\mathbf{e}_1+x_2A\mathbf{e}_2\\ \amp=x_1\mathbf{a}_1+x_2\mathbf{a}_2. \end{align*}
The second equality uses linearity, and the last uses FactΒ 1.3.15. Thus a \(2\times 2\) matrix is geometrically determined by where it sends the two coordinate directions.
To understand a \(2\times 2\) matrix \(A\text{,}\) apply it to the four corners of the unit square. The image is usually a parallelogram. The two edges leaving the origin are the columns of \(A\text{.}\) We will refer to this as the unit-square visualization.
The unit-square visualization for a shear matrix.
Three panels appear from left to right: the original unit square, the two transformed basis vectors, and the image parallelogram. In the left panel, a red horizontal arrow marks \(\mathbf e_1\) and a green vertical arrow marks \(\mathbf e_2\text{.}\) In the middle panel, a red arrow marks \(A\mathbf e_1\) and a green arrow marks \(A\mathbf e_2\text{;}\) in the right panel, those two arrows form adjacent edges of the sheared parallelogram. Blue dots label the corner \((1,1)\) in the left panel and its image \((2,1)\) in the right panel.
Figure 1.3.17. A \(2\times 2\) matrix sends the unit square to the parallelogram determined by its two columns. This is the geometric version of columns contribute.
Three-panel unit-square visualization of a horizontal stretch by a factor of two.
The left panel shows the unit square with a red horizontal basis arrow and a green vertical basis arrow. The middle panel shows the columns of \(A=\begin{bmatrix}2\amp0\\0\amp1\end{bmatrix}\text{:}\) the red arrow reaches \((2,0)\) and the green arrow reaches \((0,1)\text{.}\) The right panel shows the image rectangle with corners \((0,0)\text{,}\) \((2,0)\text{,}\) \((2,1)\text{,}\) and \((0,1)\text{.}\) The width doubles and the height stays the same.
Figure 1.3.18. A horizontal stretch by a factor of \(2\) sends the unit square to a rectangle twice as wide. The first column doubles the horizontal basis vector, while the second column leaves the vertical basis vector unchanged.

Activity 1.3.20. Horizontal shear (U1-LO4).

Let \(A=\begin{bmatrix}1\amp1\\0\amp1\end{bmatrix}\text{.}\) Compute \(A\mathbf{e}_1\text{,}\) \(A\mathbf{e}_2\text{,}\) and \(A\mathbf{x}\text{,}\) where \(\mathbf{x}=(1,1)\text{.}\)
Solution.
We have
\begin{equation*} A\mathbf{e}_1=\begin{bmatrix}1\\0\end{bmatrix},\qquad A\mathbf{e}_2=\begin{bmatrix}1\\1\end{bmatrix},\qquad A\mathbf{x}=\begin{bmatrix}2\\1\end{bmatrix}. \end{equation*}
The formula is \(A(x,y)=(x+y,y)\text{.}\) Points with \(y=0\) stay fixed, and points with \(y=1\) shift right by \(1\text{.}\)

Note 1.3.21.

For the shear matrix \(A=\begin{bmatrix}1\amp1\\0\amp1\end{bmatrix}\text{,}\) the row view gives
\begin{equation*} A\begin{bmatrix}x\\y\end{bmatrix} = \begin{bmatrix} (1,1)\cdot(x,y)\\ (0,1)\cdot(x,y) \end{bmatrix} = \begin{bmatrix}x+y\\y\end{bmatrix}. \end{equation*}
The first row measures horizontal plus vertical, while the second row measures vertical only. The column view gives
\begin{equation*} A\begin{bmatrix}x\\y\end{bmatrix} = x\begin{bmatrix}1\\0\end{bmatrix} + y\begin{bmatrix}1\\1\end{bmatrix}, \end{equation*}
so the \(y\)-coordinate contributes both upward and rightward motion.

Activity 1.3.22. Reflection across the \(y\)-axis (U1-LO4).

Let \(A=\begin{bmatrix}-1\amp0\\0\amp1\end{bmatrix}\text{.}\) Compute \(A\mathbf{e}_1\text{,}\) \(A\mathbf{e}_2\text{,}\) \(A\mathbf{x}\text{,}\) and \(A\mathbf{y}\text{,}\) where \(\mathbf{x}=(1,1)\) and \(\mathbf{y}=(0,1)\text{.}\)
Solution.
We get
\begin{align*} A\mathbf{e}_1 \amp= \begin{bmatrix}-1\\0\end{bmatrix},\qquad A\mathbf{e}_2=\begin{bmatrix}0\\1\end{bmatrix},\\ A\mathbf{x} \amp= \begin{bmatrix}-1\\1\end{bmatrix},\qquad A\mathbf{y}=\begin{bmatrix}0\\1\end{bmatrix}. \end{align*}
The first coordinate flips sign and the second coordinate is unchanged.
Three-panel unit-square visualization of reflection across the y-axis.
The left panel shows the unit square with a red horizontal basis arrow and a green vertical basis arrow. The middle panel shows the columns of \(A=\begin{bmatrix}-1\amp0\\0\amp1\end{bmatrix}\text{:}\) the red arrow points left to \((-1,0)\) and the green arrow points up to \((0,1)\text{.}\) The right panel shows the reflected square with corners \((0,0)\text{,}\) \((-1,0)\text{,}\) \((-1,1)\text{,}\) and \((0,1)\text{.}\) Its labeled arrows preserve the correspondence with the original basis vectors.
Figure 1.3.23. Reflection across the \(y\)-axis sends the unit square to its left-hand mirror image. The first column reverses the horizontal direction, while the second column leaves the vertical direction unchanged.

Activity 1.3.24. Reflection across the line \(y=x\) (U1-LO4, U1-LO5).

Let
\begin{equation*} A=\begin{bmatrix}0\amp1\\1\amp0\end{bmatrix}. \end{equation*}
  1. Compute \(A\mathbf{e}_1\) and \(A\mathbf{e}_2\text{.}\)
  2. Compute \(A\begin{bmatrix}x\\y\end{bmatrix}\text{.}\)
  3. What happens to a point on the line \(y=x\text{?}\)
Solution.
We have
\begin{equation*} A\mathbf{e}_1=\begin{bmatrix}0\\1\end{bmatrix}=\mathbf{e}_2, \qquad A\mathbf{e}_2=\begin{bmatrix}1\\0\end{bmatrix}=\mathbf{e}_1. \end{equation*}
More generally, write the input in terms of the standard basis:
\begin{equation*} \begin{aligned} A\begin{bmatrix}x\\y\end{bmatrix} &=A(x\mathbf{e}_1+y\mathbf{e}_2) =xA\mathbf{e}_1+yA\mathbf{e}_2\\ &=x\mathbf{e}_2+y\mathbf{e}_1 =\begin{bmatrix}y\\x\end{bmatrix}. \end{aligned} \end{equation*}
Thus the two coordinates are exchanged. If \(x=y\text{,}\) exchanging the coordinates does not move the point, so every point on the line \(y=x\) stays fixed.
Three-panel unit-square visualization of reflection across y equals x, with exchanged basis arrows.
The left panel shows the unit square with a red arrow to \(\mathbf e_1=(1,0)\) and a green arrow to \(\mathbf e_2=(0,1)\text{.}\) The middle panel shows the columns of \(A=\begin{bmatrix}0\amp1\\1\amp0\end{bmatrix}\text{:}\) the red arrow \(A\mathbf e_1\) points up and the green arrow \(A\mathbf e_2\) points right. The right panel fills the same unit square, with the red edge now vertical and the green edge horizontal. The dashed line \(y=x\) passes through the two fixed corners, \((0,0)\) and \((1,1)\text{.}\)
Figure 1.3.25. The unit square occupies the same region after reflection across \(y=x\text{,}\) even though its horizontal and vertical coordinate directions have exchanged, so an unlabeled image of the square can appear unchanged. The labeled red and green edges show the exchange. The corners \((0,0)\) and \((1,1)\) on the dashed mirror line stay fixed.

Subsection Linear and affine maps

A function assigns one output to each input. In this course, many functions have vectors as inputs and vectors as outputs:
\begin{equation*} F:\mathbb R^n\to\mathbb R^m. \end{equation*}
The domain of \(F\) is its set of allowed inputs, \(\mathbb R^n\text{,}\) and its co-domain is the specified set containing its outputs, \(\mathbb R^m\text{.}\) A matrix gives one important source of such functions by the rule \(\mathbf{x}\mapsto A\mathbf{x}\text{.}\)

Definition 1.3.26. Matrix map.

Given an \(m\times n\) matrix \(A\text{,}\) the matrix map induced by \(A\) is the function \(T_A:\mathbb R^n\to\mathbb R^m\) defined by
\begin{equation*} T_A(\mathbf{x})=A\mathbf{x}. \end{equation*}

Definition 1.3.27. Linear map.

A function \(T:\mathbb R^n\to\mathbb R^m\) is linear if
\begin{align*} T(\mathbf{u}+\mathbf{v}) \amp= T(\mathbf{u})+T(\mathbf{v}),\\ T(c\mathbf{v}) \amp= cT(\mathbf{v}) \end{align*}
for all vectors \(\mathbf{u},\mathbf{v}\) and all scalars \(c\text{.}\)

Note 1.3.28. Linear maps preserve linear combinations.

If \(T:\mathbb R^n\to\mathbb R^m\) is linear, then for any vectors \(\mathbf{v}_1,\ldots,\mathbf{v}_k\in\mathbb R^n\) and scalars \(c_1,\ldots,c_k\in\mathbb R\text{,}\)
\begin{equation*} T(c_1\mathbf{v}_1+\cdots+c_k\mathbf{v}_k)=c_1T(\mathbf{v}_1)+\cdots+c_kT(\mathbf{v}_k). \end{equation*}
Thus a linear combination of inputs is sent to the linear combination of their outputs with the same coefficients. This follows by repeatedly applying the two linearity rules.

Why is this true?.

Matrix-vector multiplication distributes over vector addition and scalar multiplication:
\begin{equation*} A(\mathbf{u}+\mathbf{v})=A\mathbf{u}+A\mathbf{v},\qquad A(c\mathbf{v})=cA\mathbf{v}. \end{equation*}
These are exactly the two linearity rules.
Every vector \(\mathbf{x}=(x_1,\ldots,x_n)\in\mathbb R^n\) can be written as a linear combination of the standard basis vectors. Indeed, \(x_j\mathbf{e}_j\) has entry \(x_j\) in position \(j\) and zeros elsewhere, so adding these vectors coordinate by coordinate gives
\begin{equation*} x_1\mathbf{e}_1+\cdots+x_n\mathbf{e}_n=\begin{bmatrix}x_1\\\vdots\\x_n\end{bmatrix}=\mathbf{x}. \end{equation*}
Since a linear map preserves linear combinations, its values on \(\mathbf{e}_1,\ldots,\mathbf{e}_n\) determine its value on every input. This observation leads to the following theorem.

Why is this true?.

Using the standard-basis expansion above, let \(A=[\,T(\mathbf{e}_1)\ \cdots\ T(\mathbf{e}_n)\,]\text{.}\) By linearity and the column view of matrix-vector multiplication,
\begin{align*} T(\mathbf{x}) \amp=T(x_1\mathbf{e}_1+\cdots+x_n\mathbf{e}_n)\\ \amp=x_1T(\mathbf{e}_1)+\cdots+x_nT(\mathbf{e}_n)\\ \amp=\begin{bmatrix}T(\mathbf{e}_1)\amp\cdots\amp T(\mathbf{e}_n)\end{bmatrix}\begin{bmatrix}x_1\\\vdots\\x_n\end{bmatrix}\\ \amp=A\mathbf{x}. \end{align*}
The second equality uses linearity, and the third uses the column view of matrix-vector multiplication. This matrix is unique because its \(j\)-th column must be \(A\mathbf{e}_j=T(\mathbf{e}_j)\text{.}\)

Activity 1.3.31. Images of basis vectors determine the matrix (U1-LO5).

Suppose a linear map \(T:\mathbb R^2\to\mathbb R^2\) satisfies \(T(\mathbf{e}_1)=(2,1)\) and \(T(\mathbf{e}_2)=(-1,3)\text{.}\) Find the matrix \(A\) such that \(T(\mathbf{x})=A\mathbf{x}\text{.}\)
Solution.
The columns are the images of the standard basis vectors, so
\begin{equation*} A=\begin{bmatrix}2 \amp -1\\1 \amp 3\end{bmatrix}. \end{equation*}

Warning 1.3.32.

We use β€œlinear map” and β€œlinear transformation” interchangeably; β€œtransformation” is often used when emphasizing geometry.

Definition 1.3.33. Affine map.

Let \(A\) be an \(m\times n\) real matrix and let \(\mathbf{b}\in\mathbb R^m\) be a fixed column vector. The function
\begin{equation*} F:\mathbb R^n\to\mathbb R^m,\qquad F(\mathbf{x})=A\mathbf{x}+\mathbf{b}, \end{equation*}
where \(\mathbf{x}\in\mathbb R^n\) is a column vector, is called affine. It is built from a linear map followed by a translation.
An affine map is linear only when \(\mathbf{b}=\mathbf{0}\text{.}\) A quick test is the zero vector: every linear map sends \(\mathbf{0}\) to \(\mathbf{0}\text{,}\) but
\begin{equation*} A\mathbf{0}+\mathbf{b}=\mathbf{b}. \end{equation*}
If \(\mathbf{b}\ne\mathbf{0}\text{,}\) the map is not linear.
A unit square is stretched horizontally and then translated one unit right and one unit up.
Three panels share the same coordinate scale. The left panel shows the unit square with red and green standard basis arrows and the corner \((1,1)\) marked in blue. The middle panel shows the linear image under \(A=\begin{bmatrix}2\amp0\\0\amp1\end{bmatrix}\text{:}\) a rectangle with corners \((0,0)\text{,}\) \((2,0)\text{,}\) \((2,1)\text{,}\) and \((0,1)\text{.}\) The right panel shows its translation by \(\mathbf b=(1,1)\text{,}\) with corners \((1,1)\text{,}\) \((3,1)\text{,}\) \((3,2)\text{,}\) and \((1,2)\text{.}\) A purple arrow from the origin to \((1,1)\) marks the translation vector and the image of zero. The red and green edge arrows now start at \(\mathbf b\text{,}\) and the blue corner is labeled \((3,2)\text{.}\)
Figure 1.3.34. The horizontal stretch from FigureΒ 1.3.18 followed by translation by \(\mathbf b=(1,1)\text{.}\) The translation shifts every point one unit right and one unit up, preserving the stretched rectangle’s shape. In particular, \(F(\mathbf 0)=\mathbf b\ne\mathbf 0\text{,}\) so this affine map is not linear.

Example 1.3.35. An affine layer.

A layer in a neural network can take the form
\begin{equation*} \mathbf{x}\mapsto W\mathbf{x}+\mathbf{b}. \end{equation*}
The symbol \(\mapsto\) means β€œmaps to”: the input vector \(\mathbf{x}\) is sent to the output vector \(W\mathbf{x}+\mathbf{b}\text{.}\) The matrix \(W\) mixes the input coordinates. The vector \(\mathbf{b}\) shifts the result. Since
\begin{equation*} W\mathbf{0}+\mathbf{b}=\mathbf{b}, \end{equation*}
this map is affine. It is linear exactly when \(\mathbf{b}=\mathbf{0}\text{.}\)