Matrices as transformations
See a matrix as a machine that moves every point of the plane, learn why its columns say where the axes land, and build one layer of a neural network.
FreeAbout 15 min
A matrix moves the whole plane
In the last lesson, a vector was a point or an arrow. A matrix is a machine that moves points: put a vector in, get a new vector out. Do that to every point of the plane at once and the whole plane turns, stretches, slants or flattens.
Definition (Matrix)
An matrix is a grid of real numbers with rows and columns. The entry in row and column is . A matrix is written
Press the presets and watch the grid, the arrows and , and the letter F. Each preset is a different matrix acting on the plane.
Notice what never happens: grid lines never bend, and the origin never moves. That's the mark of the transformations matrices make.
Multiplying a matrix by a vector
Definition (Matrix–vector product)
For and , the product is the vector in whose -th entry is the dot product of row of with :
The vector needs as many entries as has columns, and the answer has as many entries as has rows.
Example (By hand)
Take and . Row 1 gives , and row 2 gives , so
The matrix moves the point to .
Exercise, level: Core
The columns are where the axes land
The standard basis vectors of are and . There's a second way to read a matrix–vector product, and it explains the whole picture.
Theorem (Column picture)
Let and be the columns of a matrix . Then , , and for every ,
Proof
By the definition, . Split it by which each term contains:
With (so , ) this gives , and with it gives . The same argument works for any number of columns.
End of proof.
So the first column is where lands and the second column is where lands, and every other point follows. The point is "3 steps along , then 2 steps along ", so it lands 3 steps along the new and 2 along the new . For the matrix above:
That's the same answer as the rows gave. Here is that matrix. Drag the tips of and , or type new entries, and watch the columns change with them.
Exercise, level: Core
Linear maps
Why do grid lines never bend? Because of two rules that every matrix obeys.
Definition (Linear map)
A function is linear if, for all vectors and every number ,
Theorem (Matrices and linear maps are the same thing)
- For every matrix , the map is linear.
- Every linear map is for exactly one matrix : the one whose columns are and .
Proof
Part 1. Each entry of is a dot product of a row with , and the dot product spreads over addition and lets numbers pull out (lesson 1). So each entry of is the matching entry of plus that of , and each entry of is times that of .
Part 2. Every can be written . Using the two rules,
By the column picture, that's for the matrix with columns and . Any matrix that does the same must send to and to , so it has the same columns: it's the same matrix.
End of proof.
Two consequences you saw in the widget: a linear map keeps the origin fixed, because ; and it keeps lines straight, because the points of a line land on the points , another line (or a single point when ).
The presets are these matrices. Read each column as where or goes:
| Preset | Matrix | What it does |
|---|---|---|
| Rotate | turns the plane counterclockwise | |
| Scale | stretches by 2 and by 1.5 | |
| Shear | slides each point sideways by its height | |
| Reflect | mirrors the plane in the -axis | |
| Project | flattens the plane onto the -axis |
The widget also shows , the determinant. Its absolute value is the factor by which the matrix scales areas, and it's negative when the matrix flips the plane over. We state that here without proof; determinants get a full module in Linear Algebra for ML.
Exercise, level: Derivation
One transformation after another
Apply first, then : the point goes to , then to . The result is again a linear map, so by the theorem it's a single matrix.
Definition (Matrix product)
The product is the matrix whose -th column is times the -th column of . It's defined when has as many columns as has rows, and then for every .
Why that column rule: the combined map sends to , and is the -th column of . Its matrix has those images as columns. Note the order: in , the matrix on the right acts first.
Example (Order matters)
Let (turn ) and (shear). Turning twice is , a half turn: every point goes to its opposite. But the two orders of and give different matrices:
Shearing then turning isn't the same as turning then shearing. Matrix products usually depend on the order.
Exercise, level: Derivation
By hand: one layer of a neural network
A layer of a neural network takes a vector and computes
The weights are a matrix, so the first part is a linear map: it turns, stretches and shears space. The bias shifts everything over. Then , applied to each entry, replaces negative numbers with : it folds the parts of space with negative coordinates flat onto the axes.
Example (One layer)
Take , and the input .
The ReLU matters more than it looks. Without it, two layers in a row would be : a single matrix and a single shift, no more powerful than one layer. The fold between layers is what lets a deep network bend space into shapes no single matrix can make.
In code
NumPy writes the matrix–vector product with @, just like the dot product. A[:, 0] is the first column:
import numpy as np
A = np.array([[2.0, -1.0],
[1.0, 1.0]])
x = np.array([3.0, 2.0])
print(A @ x) # rows dotted with x
print(3 * A[:, 0] + 2 * A[:, 1]) # 3 of column 1 plus 2 of column 2Both lines print [4. 5.]. Put many points in the columns of one array and a single @ moves them all. Here are 200 points on the unit circle, moved by the same matrix:
import numpy as np
A = np.array([[2.0, -1.0], [1.0, 1.0]])
t = np.linspace(0, 2 * np.pi, 200)
circle = np.stack([np.cos(t), np.sin(t)]) # 2 rows, 200 columns: one point per column
moved = A @ circle
show(moved[0], moved[1], kind="scatter", title="The unit circle after A")The circle comes out as a tilted ellipse: a linear map keeps lines straight, but it can stretch a circle. Finally, the layer from the last step:
import numpy as np
W = np.array([[1.0, -1.0], [2.0, 1.0]])
b = np.array([0.0, -1.0])
x = np.array([1.0, 2.0])
print(np.maximum(0, W @ x + b)) # ReLU(Wx + b)It prints [0. 3.], as we found by hand. Change x and run it again.
Where it's used in ML
Tip
Every dense layer of a neural network is a matrix. A layer that reads a image as a vector of pixels and outputs numbers has a weight matrix and a bias with entries. Training (lesson 4) adjusts all of those numbers. Language models are built from the same pieces, with matrices of many millions of entries.
Exercise, level: Core
Exercise, level: Exam
Want to go further, with the four fundamental subspaces, determinants and more? Matrices are the second module of Linear Algebra for ML.
or press Enter