Skip to main content

MATH 345: Linear Algebra and Optimization

Section 5.2 Diagonalization

Eigenvalues describe individual special directions. Diagonalization asks whether there are enough independent eigenvectors to form a basis. In eigenvector coordinates, the matrix acts by scaling each coordinate separately.

Definition 5.2.1.

An \(n\times n\) matrix \(A\) is called diagonalizable if there exists an invertible \(n \times n\) matrix \(P\) so that \(P^{-1}AP\) is a diagonal matrix (recall Definitionย 2.5.36). The matrix \(P\) is called a diagonalizing matrix for \(A\text{.}\)

Activity 5.2.2. Diagonalization and a matrix power (U5-LO2).

\begin{equation*} A=\begin{bmatrix} 3 \amp 1 \\ -2 \amp 0 \end{bmatrix}\text{.} \end{equation*}

(a)

Compute \(P^{-1}AP\) where
\begin{equation*} P=\begin{bmatrix} -1/2 \amp -1 \\ 1 \amp 1 \end{bmatrix}\text{.} \end{equation*}
You may use the fact that
\begin{equation*} P^{-1}=\begin{bmatrix} 2 \amp 2 \\ -2 \amp -1 \end{bmatrix}\text{.} \end{equation*}
What does this computation say about the matrix \(A\text{?}\)
Solution.
We start by calculating that
\begin{align*} AP \amp = \begin{bmatrix} (3)(-1/2) + (1)(1) \amp (3)(-1) + (1)(1) \\ (-2)(-1/2) + (0)(1) \amp (-2)(-1) + (0)(1) \end{bmatrix}\\ \amp = \begin{bmatrix} -1/2 \amp -2 \\ 1 \amp 2 \end{bmatrix} \end{align*}
and thus that
\begin{align*} P^{-1}AP \amp = \begin{bmatrix} 2 \amp 2 \\ -2 \amp -1 \end{bmatrix} \begin{bmatrix} -1/2 \amp -2 \\ 1 \amp 2 \end{bmatrix}\\ \amp = \begin{bmatrix} (2)(-1/2) + (2)(1) \amp (2)(-2) + (2)(2) \\ (-2)(-1/2) + (-1)(1) \amp (-2)(-2) + (-1)(2) \end{bmatrix}\\ \amp = \begin{bmatrix} 1 \amp 0 \\ 0 \amp 2 \end{bmatrix} \end{align*}
Since \(P^{-1}AP\) is a diagonal matrix, the matrix \(A\) is diagonalizable.

(b)

Use the previous computation to calculate \(A^5\text{.}\)
Solution.
Letโ€™s start by computing smaller powers of \(A\text{.}\) We find that
\begin{align*} A^2 \amp = A \cdot A = (PDP^{-1})(PDP^{-1})\\ \amp = PD(P^{-1}P)DP^{-1} = PD^2P^{-1} \end{align*}
and
\begin{align*} A^3 \amp = A \cdot A^2 = (PDP^{-1})(PD^2P^{-1})\\ \amp = PD(P^{-1}P)D^2P^{-1} = PD^3P^{-1} \end{align*}
Continuing this calculation, noticing the pattern, we conclude that \(A^5 = PD^5P^{-1}\text{.}\)
Multiplying diagonal matrices is easy, i.e.,
\begin{equation*} \begin{bmatrix} a_1 \amp 0 \\ 0 \amp b_1 \end{bmatrix} \begin{bmatrix} a_2 \amp 0 \\ 0 \amp b_2 \end{bmatrix} \cdots \begin{bmatrix} a_m \amp 0 \\ 0 \amp b_m \end{bmatrix} = \begin{bmatrix} a_1 \cdots a_m \amp 0 \\ 0 \amp b_1 \cdots b_m \end{bmatrix}\text{.} \end{equation*}
Thus
\begin{equation*} D^5 = \begin{bmatrix} 1 \amp 0 \\ 0 \amp 2 \end{bmatrix}^5 = \begin{bmatrix} 1^5 \amp 0 \\ 0 \amp 2^5 \end{bmatrix} = \begin{bmatrix} 1 \amp 0 \\ 0 \amp 32 \end{bmatrix}\text{.} \end{equation*}
We use this computation to compute \(A^5\text{,}\) i.e., we calculate that
\begin{equation*} PD^5 = \begin{bmatrix} -1/2 \amp -1 \\ 1 \amp 1 \end{bmatrix} \begin{bmatrix} 1 \amp 0 \\ 0 \amp 32 \end{bmatrix} = \begin{bmatrix} -1/2 \amp -32 \\ 1 \amp 32 \end{bmatrix} \end{equation*}
and thus that
\begin{align*} PD^5P^{-1} \amp = \begin{bmatrix} -1/2 \amp -32 \\ 1 \amp 32 \end{bmatrix} \begin{bmatrix} 2 \amp 2 \\ -2 \amp -1 \end{bmatrix}\\ \amp = \begin{bmatrix} 63 \amp 31 \\ -62 \amp -30 \end{bmatrix}\text{.} \end{align*}
The previous activity shows why diagonalizing a matrix may be useful. But given a matrix \(A\text{,}\) how do we potentially find \(P\) such that \(P^{-1}AP\) is a diagonal matrix?

Activity 5.2.3. Interpreting the columns in a diagonalization (U5-LO1, U5-LO2).

Suppose \(P^{-1}AP = D\) for some diagonal matrix \(D = \diag(\lambda_1,\dots,\lambda_n)\text{.}\) Let \(\mathbf{v}_1, \dots, \mathbf{v}_n\) be the columns of \(P\text{.}\) Note that the equation \(P^{-1} A P = D\) holds if and only if \(AP = PD\text{.}\)

(a)

What happens when I take \(PD\text{?}\)
Solution.
\begin{align*} PD \amp = [\mathbf{v}_1 \mathbf{v}_2 \cdots \mathbf{v}_n] \begin{bmatrix} \lambda_1 \amp 0 \amp \cdots \amp 0 \\ 0 \amp \lambda_2 \amp \cdots \amp 0 \\ \vdots \amp \vdots \amp \ddots \amp \vdots \\ 0 \amp 0 \amp \cdots \amp \lambda_n \end{bmatrix} \end{align*}
Matrix multiplication gives us:
\begin{equation*} PD = [\lambda_1\mathbf{v}_1 \; \lambda_2\mathbf{v}_2 \; \cdots \; \lambda_n\mathbf{v}_n] \end{equation*}
So \(PD\) is the matrix whose columns are the vectors \(\mathbf{v}_i\) scaled by their corresponding \(\lambda_i\) values.

(b)

Calculate \(AP\) and \(PD\) in terms of the vectors \(\mathbf{v}_1,\dots,\mathbf{v}_n\text{,}\) the scalars \(\lambda_1,\dots,\lambda_n\) and the matrix \(A\text{.}\)
Solution.
When we take \(AP\text{,}\) we compute that
\begin{equation*} AP = A \begin{bmatrix} \mathbf{v}_1 \amp \mathbf{v}_2 \amp \cdots \amp \mathbf{v}_n \end{bmatrix} = \begin{bmatrix} A\mathbf{v}_1 \amp A\mathbf{v}_2 \amp \cdots \amp A\mathbf{v}_n \end{bmatrix}\text{.} \end{equation*}
When we calculate \(PD\text{,}\) we compute that
\begin{align*} PD \amp = \begin{bmatrix} \mathbf{v}_1 \amp \mathbf{v}_2 \amp \cdots \amp \mathbf{v}_n \end{bmatrix} \begin{bmatrix} \lambda_1 \amp 0 \amp \cdots \amp 0 \\ 0 \amp \lambda_2 \amp \cdots \amp 0 \\ \vdots \amp \vdots \amp \ddots \amp \vdots \\ 0 \amp 0 \amp \cdots \amp \lambda_n \end{bmatrix}\\ \amp = \begin{bmatrix} \lambda_1 \mathbf{v}_1 \amp \cdots \amp \lambda_n \mathbf{v}_n \end{bmatrix}\text{.} \end{align*}

(c)

So what conditions on \(P\) and \(D\) are necessary and sufficient for the equation \(AP = PD\) to hold?
Solution.
For the equation \(AP = PD\) to hold, we need \(A \mathbf{v}_i = \lambda_i \mathbf{v}_i\) for each \(i\text{.}\) Thus we need each column of \(P\) to be an eigenvector of \(A\text{,}\) whose eigenvalue is the corresponding entry of the diagonal matrix \(D\text{.}\)
To see this differently, suppose \(P^{-1}AP\) is a diagonal matrix with diagonal entries \(\lambda_1,\dots,\lambda_n\text{.}\) Then for each standard basis vector (recall Definitionย 1.3.14) \(P^{-1}AP \mathbf{e}_i = \lambda_i \mathbf{e}_i\text{.}\) Multiplying this equation by \(P\) on the left-hand side and setting \(\mathbf{v}_i = P \mathbf{e}_i\text{,}\) we obtain
\begin{equation*} A \mathbf{v}_i = P(P^{-1}AP \mathbf{e}_i) = P(\lambda_i \mathbf{e}_i) = \lambda_i \mathbf{v}_i\text{.} \end{equation*}
Thus \(\mathbf{v}_i\) is an eigenvector of the matrix \(A\) associated with the eigenvalue \(\lambda_i\text{.}\)
If \(\R^n\) has a basis of eigenvectors \(\mathbf{v}_1,\ldots,\mathbf{v}_n\) with corresponding eigenvalues \(\lambda_1,\ldots,\lambda_n\text{,}\) then every vector \(\mathbf{v} \in \R^n\) can be written as a linear combination of these eigenvectors, i.e., there exists constants \(a_1,\dots,a_n\) such that
\begin{equation*} \mathbf{v} = a_1 \mathbf{v}_1 + a_2 \mathbf{v}_2 + \cdots + a_n \mathbf{v}_n\text{.} \end{equation*}
Thus
\begin{align*} T_A (\mathbf{v}) \amp = A \mathbf{v}\\ \amp = A(a_1 \mathbf{v}_1 + a_2 \mathbf{v}_2 + \cdots + a_n \mathbf{v}_n)\\ \amp = a_1 A \mathbf{v}_1 + a_2 A\mathbf{v}_2 + \cdots + a_n A \mathbf{v}_n\\ \amp = a_1 \lambda_1 \mathbf{v}_1 + a_2 \lambda_2 \mathbf{v}_2 + \cdots + a_n \lambda_n \mathbf{v}_n\\ \amp = \lambda_1 (a_1 \mathbf{v}_1) + \lambda_2 (a_2 \mathbf{v}_2) + \cdots + \lambda_n (a_n \mathbf{v}_n) \end{align*}
In other words, \(T_A\) โ€˜stretchesโ€™ the components of \(\mathbf{v}\) along each eigenvector by the corresponding eigenvalue.

Activity 5.2.6. Checking a proposed eigenvector pair (U5-LO1).

Consider the matrix
\begin{equation*} M = \begin{bmatrix} 5/3 \amp -2/3 \\ -1/3 \amp 4/3 \end{bmatrix}\text{.} \end{equation*}
Verify that
\begin{equation*} \mathbf{u} = \begin{bmatrix} 1 \\ 1 \end{bmatrix} \quad\text{and}\quad \mathbf{v} = \begin{bmatrix} 1 \\ -1/2 \end{bmatrix} \end{equation*}
are eigenvectors of \(M\) corresponding to different eigenvalues. What are the eigenvalues?
Solution.
We compute that
\begin{align*} M\mathbf{u} \amp = \begin{bmatrix} 5/3 \amp -2/3 \\ -1/3 \amp 4/3 \end{bmatrix}\begin{bmatrix} 1 \\ 1 \end{bmatrix}\\ \amp = \begin{bmatrix} 5/3 - 2/3 \\ -1/3 + 4/3 \end{bmatrix}\\ \amp = \begin{bmatrix} 1 \\ 1 \end{bmatrix}\\ \amp = 1 \cdot \mathbf{u} \end{align*}
and
\begin{align*} M\mathbf{v} \amp = \begin{bmatrix} 5/3 \amp -2/3 \\ -1/3 \amp 4/3 \end{bmatrix}\begin{bmatrix} 1 \\ -1/2 \end{bmatrix}\\ \amp = \begin{bmatrix} 5/3 - 2/3 \cdot (-1/2) \\ -1/3 + 4/3 \cdot (-1/2) \end{bmatrix}\\ \amp = \begin{bmatrix} 5/3 + 1/3 \\ -1/3 - 2/3 \end{bmatrix}\\ \amp = \begin{bmatrix} 2 \\ -1 \end{bmatrix}\\ \amp = 2 \cdot \begin{bmatrix} 1 \\ -1/2 \end{bmatrix}\\ \amp = 2 \cdot \mathbf{v}\text{.} \end{align*}
Therefore, \(\mathbf{u}\) is an eigenvector with eigenvalue \(\lambda_1 = 1\) and \(\mathbf{v}\) is an eigenvector with eigenvalue \(\lambda_2 = 2\text{.}\)
Figureย 5.2.7 illustrates the effect of the matrix \(M\) from Activityย 5.2.6 on a grid in \(\R^2\text{.}\) In particular, it highlights an important geometric interpretation of eigenvectors and eigenvalues; when we apply the linear map defined by the matrix \(M\) to any vector in \(\R^2\text{,}\) the matrix map stretches eigenvector directions by their eigenvalues.
Before-and-after grid diagram for a matrix with two highlighted eigenvector directions.
The left panel shows a square coordinate grid with a red line in the direction \(\mathbf{u}=(1,1)\) and a blue line in the direction \(\mathbf{v}=(1,-1/2)\text{.}\) The right panel shows the image of the grid under \(M\text{.}\) The red direction is unchanged, while the blue direction is stretched to twice its original length.
Figure 5.2.7. Visualization of the action of the matrix \(M\) from Activityย 5.2.6 on a grid. The red line along the eigenvector \(\mathbf{u}\) is unchanged, because \(\mathbf{u}\) has eigenvalue \(1\text{.}\) The blue line along the eigenvector \(\mathbf{v}\) is stretched by a factor of \(2\text{,}\) because \(\mathbf{v}\) has eigenvalue \(2\text{.}\) Adapted from: Stanfordโ€™s MATH 51 textbook.

Remark 5.2.8. Non-diagonalizable matrices.

Consider the rotation matrix
\begin{equation*} A_\theta = \begin{bmatrix} \cos\theta \amp -\sin\theta \\ \sin\theta \amp \cos\theta \end{bmatrix} \end{equation*}
where \(\theta\) is not a multiple of \(180ยฐ\text{.}\) For any nonzero vector \(\mathbf{v} \in \R^2\text{,}\) the vector \(A_\theta\mathbf{v}\) is obtained by rotating \(\mathbf{v}\) counterclockwise by an angle \(\theta\text{,}\) so \(A_\theta\mathbf{v}\) is never on the line spanned by \(\mathbf{v}\text{.}\) Thus \(A_\theta\mathbf{v}=\lambda\mathbf{v}\) cannot hold for any real scalar \(\lambda\) and any nonzero \(\mathbf{v}\text{.}\)
Thus the matrix \(A_\theta\) has no real eigenvalues or real eigenvectors. It is therefore not diagonalizable over \(\R\text{.}\) This makes geometric sense: we cannot find special directions in \(\R^2\) so that rotating vectors in the plane is obtained by stretching in those directions.
Remarkย 5.2.8 illustrates why not every matrix is diagonalizable. A matrix is diagonalizable precisely when it has a basis of eigenvectors, but this is not always possible.

Activity 5.2.9. Testing a matrix for diagonalizability (U5-LO2).

Let
\begin{equation*} A=\begin{bmatrix} 3 \amp 0 \\ -1 \amp -4 \end{bmatrix}\text{.} \end{equation*}
Is \(A\) diagonalizable?
Solution.
We need to check if \(A\) has a basis of eigenvectors. In Activityย 5.1.3, we showed that \(A\) has eigenvectors
\begin{equation*} \mathbf{v}_1 = \begin{bmatrix} -7 \\ 1 \end{bmatrix} \quad\text{and}\quad \mathbf{v}_2 = \begin{bmatrix} 0 \\ 1 \end{bmatrix}\text{.} \end{equation*}
The two vectors are linearly independent as they are not parallel, and thus form a basis for \(\R^2\text{.}\) Therefore, \(A\) is diagonalizable.

Activity 5.2.10. A second diagonalizability test (U5-LO2).

Let
\begin{equation*} A = \begin{bmatrix} 4 \amp 1 \\ 0 \amp 4 \end{bmatrix}\text{.} \end{equation*}
Is \(A\) diagonalizable?
Solution.
We start by calculating the characteristic polynomial of \(A\text{,}\) that
\begin{equation*} \det(\lambda I - A) = \det\begin{bmatrix} \lambda-4 \amp -1 \\ 0 \amp \lambda-4 \end{bmatrix} = (\lambda-4)^2\text{.} \end{equation*}
Thus \(\lambda = 4\) is the only eigenvalue of \(A\text{.}\)
To find the eigenvectors of \(A\text{,}\) we solve the equation \((4I - A)\mathbf{v} = \mathbf{0}\text{,}\) which we expand as
\begin{equation*} \begin{bmatrix} 0 \amp 1 \\ 0 \amp 0 \end{bmatrix}\begin{bmatrix} v_1 \\ v_2 \end{bmatrix} = \begin{bmatrix} 0 \\ 0 \end{bmatrix}\text{.} \end{equation*}
The equation has only a single basic solution, i.e.,
\begin{equation*} \begin{bmatrix} 1 \\ 0 \end{bmatrix}\text{.} \end{equation*}
Recalling Factย 5.1.11, every eigenvector of \(A\) is a scalar multiple of this vector. So \(A\) cannot have a basis of eigenvectors, and so \(A\) is not diagonalizable.