. The main sequence practices vector applications, document similarity, token attention, matrix-vector products, geometric matrix actions, applying one matrix to many points, composition, affine maps, and short code outputs. The final part contains review questions and a short forward look to square-grid visualizations.
The value is ["D1", "D2", "D3"]. The documents are being ranked by cosine similarity with the query. This is a direction-based ranking because cosine similarity compares direction after normalization.
The vector s contains dot-product scores. The vector alpha contains weights because the scores were divided by their sum. The vector out is a weighted average of the value vectors stored in the rows of V.
The columns of X are vertices of the unit square. The columns of Y are their images under \(A\text{.}\) This is a horizontal shear because \(A\) sends \((x,y)\) to \((x+y,y)\text{.}\)
No. A linear map must send \(\mathbf{0}\) to \(\mathbf{0}\text{.}\) But \(f(\mathbf{0})=W\mathbf{0}+\mathbf{b}=\mathbf{b}\text{,}\) and \(\mathbf{b}\) is not the zero vector. The rule is affine, but not linear.