矩阵变化
Type: Build
Languages: Python, Julia
Prerequisites: Phase 1, Lessons 01-02 (Linear Algebra Intuition, Vectors & Matrices Operations)
Time: ~75 minutes
学习目标
- 构建旋转,扩展,切割和反射矩阵,并将它们应用于2D和3D点
- 通过矩阵乘法编写多个转换,并验证顺序是否重要
- 从特征方程计算2x2矩阵的自值和自向量
- 解释为什么自值决定PCA方向,RNN稳定性和光谱集群行为
问题
你读到PCA,看"求出对差矩阵的自向量".你读到模型稳定性,看"检查所有自值是否小于1."你读到数据增强,看"应用随机旋转".
矩阵不仅仅是数字的网格.它们是空间机器.一个旋转矩阵旋转点.一个扩展矩阵伸缩它们.一个切割矩阵倾斜它们.一个神经网络对数据的每一个转变都是这些操作之一或它们的组成.这个课程使这些操作具体.
概念
转变为矩阵
任何2D的线性转换都可以写成2x2矩阵.矩阵告诉你底向量 [1, 0] 和 [0, 1] 最终到底是哪里.其他的一切都会随之而来.
graph LR
subgraph Before["Standard Basis"]
e1["e1 = [1, 0] (along x)"]
e2["e2 = [0, 1] (along y)"]
end
subgraph Transform["Matrix M"]
M["M = columns are new basis vectors"]
end
subgraph After["After Transformation M"]
e1p["e1' = new x-basis"]
e2p["e2' = new y-basis"]
end
e1 --> M --> e1p
e2 --> M --> e2p转动
通过角度旋转,保持距离和角落完整. 它沿着圆弧移动每个点.
graph LR
subgraph Before["Before Rotation"]
A["A(2, 1)"]
B["B(0, 2)"]
end
subgraph Rot["Rotate 45 degrees"]
R["R(θ) = [[cos θ, -sin θ], [sin θ, cos θ]]"]
end
subgraph After["After Rotation"]
Ap["A'(0.71, 2.12)"]
Bp["B'(-1.41, 1.41)"]
end
A --> R --> Ap
B --> R --> Bp在3D中,你旋转一个轴,每个轴都有自己的旋转矩阵:
Rz(theta) = | cos -sin 0 | Rotate around z-axis
| sin cos 0 | (x-y plane spins, z stays)
| 0 0 1 |
Rx(theta) = | 1 0 0 | Rotate around x-axis
| 0 cos -sin | (y-z plane spins, x stays)
| 0 sin cos |
Ry(theta) = | cos 0 sin | Rotate around y-axis
| 0 1 0 | (x-z plane spins, y stays)
| -sin 0 cos |规模化
尺度延伸或压缩在每个轴线上独立.
graph LR
subgraph Before["Before Scaling"]
A["A(2, 1)"]
B["B(0, 2)"]
end
subgraph Scale["Scale sx=2, sy=0.5"]
S["S = [[2, 0], [0, 0.5]]"]
end
subgraph After["After Scaling"]
Ap["A'(4, 0.5)"]
Bp["B'(0, 1)"]
end
A --> S --> Ap
B --> S --> Bp切割
切削曲一个轴,同时保持另一个固定. 它将矩形变成平行图.
graph LR
subgraph Before["Before Shear"]
A["A(1, 0)"]
B["B(0, 1)"]
end
subgraph Shear["Shear in x, k=1"]
Sh["Shx = [[1, k], [0, 1]]"]
end
subgraph After["After Shear"]
Ap["A(1, 0) unchanged"]
Bp["B'(1, 1) shifted"]
end
A --> Sh --> Ap
B --> Sh --> Bp切割矩阵:
Shx = [[1, k], [0, 1]]转变 x 乘以 k * yShy = [[1, 0], [k, 1]]转变为 y 乘以 k * x
思考
反映反映在轴或线的点.
graph LR
subgraph Before["Before Reflection"]
A["A(2, 1)"]
end
subgraph Reflect["Reflect across y-axis"]
R["[[-1, 0], [0, 1]]"]
end
subgraph After["After Reflection"]
Ap["A'(-2, 1)"]
end
A --> R --> Ap反映矩阵:
- 反射在 y 轴上:
[[-1, 0], [0, 1]] - 通过x轴反射:
[[1, 0], [0, -1]]
组成:链接转换
应用转换A然后B是相同的乘以它们的矩阵:result = B @ A @ point顺序是重要的. 旋转然后尺度给出不同的结果,
graph LR
subgraph Path1["Rotate 90 then Scale (2, 0.5)"]
P1["(1, 0)"] -->|"Rotate 90"| P2["(0, 1)"] -->|"Scale"| P3["(0, 0.5)"]
end组成:S @ R = [[0, -2], [0.5, 0]]
graph LR
subgraph Path2["Scale (2, 0.5) then Rotate 90"]
Q1["(1, 0)"] -->|"Scale"| Q2["(2, 0)"] -->|"Rotate 90"| Q3["(0, 2)"]
end组成:R @ S = [[0, -0.5], [2, 0]]
矩阵乘法不是交换式的.
自身值和自身向量
矩阵碰到它们时,大多数向量都改变方向.自向量是特殊的:矩阵只会缩小它们,从来没有旋转它们.缩小因素是自值.
A @ v = lambda * v
v is the eigenvector (direction that survives)
lambda is the eigenvalue (how much it stretches)
Example: A = | 2 1 |
| 1 2 |
Eigenvector [1, 1] with eigenvalue 3:
A @ [1,1] = [3, 3] = 3 * [1, 1] (same direction, scaled by 3)
Eigenvector [1, -1] with eigenvalue 1:
A @ [1,-1] = [1, -1] = 1 * [1, -1] (same direction, unchanged)矩阵延伸空间3x沿 [1, 1]并保持[1, -1]不变.其他方向都是这两个混合.
自身组成
如果矩阵具有 n 线性独立的自向量,则可以分解:
A = V @ D @ V^(-1)
V = matrix whose columns are eigenvectors
D = diagonal matrix of eigenvalues
V^(-1) = inverse of V
This says: rotate into eigenvector coordinates, scale along each axis, rotate back.为什么自有价值很重要
PCA.变量矩阵的自向量是主要组件.自值值告诉你每个组件捕获多少变量.按自值排序,保持顶部k,你有维度减少.
Stability.在复发网络和动态系统中,大小 > 1 的自值导致输出爆炸.大小 < 1 导致它们消失.这是一个句子中所述的消失/爆炸梯度问题.
Spectral methods.图形神经网络使用邻近矩阵的自值.谱系集群使用拉普拉西亚的自值.自向量揭示图形的结构.
定量量缩小因素的定量
转换矩阵的定量符告诉你它在面积 (2D) 或体积 (3D) 范围内是多少.
det = 1: area preserved (rotation)
det = 2: area doubled
det = 0: space crushed to lower dimension (singular)
det = -1: area preserved but orientation flipped (reflection)
| det(Rotation) | = 1 (always)
| det(Scale sx, sy) | = sx * sy
| det(Shear) | = 1 (area preserved)
| det(Reflection) | = -1 (orientation flipped)建立它
步骤1:从零开始的转换矩阵 (Python)
pythonimport math
def rotation_2d(theta):
c, s = math.cos(theta), math.sin(theta)
return [[c, -s], [s, c]]
def scaling_2d(sx, sy):
return [[sx, 0], [0, sy]]
def shearing_2d(kx, ky):
return [[1, kx], [ky, 1]]
def reflection_x():
return [[1, 0], [0, -1]]
def reflection_y():
return [[-1, 0], [0, 1]]
def mat_vec_mul(matrix, vector):
return [
sum(matrix[i][j] * vector[j] for j in range(len(vector)))
for i in range(len(matrix))
]
def mat_mul(a, b):
rows_a, cols_b = len(a), len(b[0])
cols_a = len(a[0])
return [
[sum(a[i][k] * b[k][j] for k in range(cols_a)) for j in range(cols_b)]
for i in range(rows_a)
]
point = [1.0, 0.0]
angle = math.pi / 4
rotated = mat_vec_mul(rotation_2d(angle), point)
print(f"Rotate (1,0) by 45 deg: ({rotated[0]:.4f}, {rotated[1]:.4f})")
scaled = mat_vec_mul(scaling_2d(2, 3), [1.0, 1.0])
print(f"Scale (1,1) by (2,3): ({scaled[0]:.1f}, {scaled[1]:.1f})")
sheared = mat_vec_mul(shearing_2d(1, 0), [1.0, 1.0])
print(f"Shear (1,1) kx=1: ({sheared[0]:.1f}, {sheared[1]:.1f})")
reflected = mat_vec_mul(reflection_y(), [2.0, 1.0])
print(f"Reflect (2,1) across y: ({reflected[0]:.1f}, {reflected[1]:.1f})")转换的组成
pythonR = rotation_2d(math.pi / 2)
S = scaling_2d(2, 0.5)
rotate_then_scale = mat_mul(S, R)
scale_then_rotate = mat_mul(R, S)
point = [1.0, 0.0]
result1 = mat_vec_mul(rotate_then_scale, point)
result2 = mat_vec_mul(scale_then_rotate, point)
print(f"Rotate 90 then scale: ({result1[0]:.2f}, {result1[1]:.2f})")
print(f"Scale then rotate 90: ({result2[0]:.2f}, {result2[1]:.2f})")
print(f"Same? {result1 == result2}")步骤3:自动值从零开始 (2x2)
对于2x2矩阵[[a, b], [c, d]]个性化方程的自值解法:lambda^2 - (a+d)*lambda + (ad - bc) = 0现在,我们要去.
pythondef eigenvalues_2x2(matrix):
a, b = matrix[0]
c, d = matrix[1]
trace = a + d
det = a * d - b * c
discriminant = trace ** 2 - 4 * det
if discriminant < 0:
real = trace / 2
imag = (-discriminant) ** 0.5 / 2
return (complex(real, imag), complex(real, -imag))
sqrt_disc = discriminant ** 0.5
return ((trace + sqrt_disc) / 2, (trace - sqrt_disc) / 2)
def eigenvector_2x2(matrix, eigenvalue):
a, b = matrix[0]
c, d = matrix[1]
if abs(b) > 1e-10:
v = [b, eigenvalue - a]
elif abs(c) > 1e-10:
v = [eigenvalue - d, c]
else:
if abs(a - eigenvalue) < 1e-10:
v = [1, 0]
else:
v = [0, 1]
mag = (v[0] ** 2 + v[1] ** 2) ** 0.5
return [v[0] / mag, v[1] / mag]
A = [[2, 1], [1, 2]]
vals = eigenvalues_2x2(A)
print(f"Matrix: {A}")
print(f"Eigenvalues: {vals[0]:.4f}, {vals[1]:.4f}")
for val in vals:
vec = eigenvector_2x2(A, val)
result = mat_vec_mul(A, vec)
scaled = [val * vec[0], val * vec[1]]
print(f" lambda={val:.1f}, v={[round(x,4) for x in vec]}")
print(f" A@v = {[round(x,4) for x in result]}")
print(f" l*v = {[round(x,4) for x in scaled]}")步骤4:作为体积缩放因素的决定因素
pythondef det_2x2(matrix):
return matrix[0][0] * matrix[1][1] - matrix[0][1] * matrix[1][0]
print(f"det(rotation 45) = {det_2x2(rotation_2d(math.pi/4)):.4f}")
print(f"det(scale 2,3) = {det_2x2(scaling_2d(2, 3)):.1f}")
print(f"det(shear kx=1) = {det_2x2(shearing_2d(1, 0)):.1f}")
print(f"det(reflect y) = {det_2x2(reflection_y()):.1f}")
singular = [[1, 2], [2, 4]]
print(f"det(singular) = {det_2x2(singular):.1f}")
print("Singular: columns are proportional, space collapses to a line.")用它
通过优化程序来处理所有这些.
pythonimport numpy as np
theta = np.pi / 4
R = np.array([[np.cos(theta), -np.sin(theta)],
[np.sin(theta), np.cos(theta)]])
point = np.array([1.0, 0.0])
print(f"Rotate (1,0) by 45 deg: {R @ point}")
S = np.diag([2.0, 3.0])
composed = S @ R
print(f"Scale(2,3) after Rotate(45): {composed @ point}")
A = np.array([[2, 1], [1, 2]], dtype=float)
eigenvalues, eigenvectors = np.linalg.eig(A)
print(f"\nEigenvalues: {eigenvalues}")
print(f"Eigenvectors (columns):\n{eigenvectors}")
for i in range(len(eigenvalues)):
v = eigenvectors[:, i]
lam = eigenvalues[i]
print(f" A @ v{i} = {A @ v}, lambda * v{i} = {lam * v}")
print(f"\ndet(R) = {np.linalg.det(R):.4f}")
print(f"det(S) = {np.linalg.det(S):.1f}")
B = np.array([[3, 1], [0, 2]], dtype=float)
vals, vecs = np.linalg.eig(B)
D = np.diag(vals)
V = vecs
reconstructed = V @ D @ np.linalg.inv(V)
print(f"\nEigendecomposition A = V @ D @ V^-1:")
print(f"Original:\n{B}")
print(f"Reconstructed:\n{reconstructed}")通过 NumPy 进行3D旋转
pythondef rotation_3d_z(theta):
c, s = np.cos(theta), np.sin(theta)
return np.array([[c, -s, 0], [s, c, 0], [0, 0, 1]])
def rotation_3d_x(theta):
c, s = np.cos(theta), np.sin(theta)
return np.array([[1, 0, 0], [0, c, -s], [0, s, c]])
point_3d = np.array([1.0, 0.0, 0.0])
rotated_z = rotation_3d_z(np.pi / 2) @ point_3d
rotated_x = rotation_3d_x(np.pi / 2) @ point_3d
print(f"\n3D point: {point_3d}")
print(f"Rotate 90 around z: {np.round(rotated_z, 4)}")
print(f"Rotate 90 around x: {np.round(rotated_x, 4)}")运送它
这一课构建了PCA (第二阶段) 和神经网络权重分析的几何基础.在这里构建的自值/自向量代码是相同的算法,它支持生产ML系统中的维度减少,光谱集群和稳定分析.
运动
- 按一个单位方形 (角在 [0,0], [1,0], [1,1], [0,1]) 进行旋转,扩展和切割. 打印每个角的转换角. 检查旋转是否保持角之间的距离.
- 通过使用特征方程手动找到矩阵的自值值.然后用从头开始的函数和NumPy来验证.
- 创建一个三变的组合 (旋转30度,以 [1.5,0.8] 缩小,切割 kx=0.3) 并将其应用于圆形排列的8个点. 打印前后坐标.计算复合矩阵的确定量,并验证它等于单个确定量的产量.
关键词
| Term | What people say | What it actually means |
|---|---|---|
| Rotation matrix | "Spins things" | An orthogonal matrix that moves points along circular arcs while preserving distances and angles. Determinant is always 1. |
| Scaling matrix | "Makes things bigger" | A diagonal matrix that stretches or compresses independently along each axis. Determinant is the product of scale factors. |
| Shearing matrix | "Slants things" | A matrix that shifts one coordinate proportionally to another, turning rectangles into parallelograms. Determinant is 1. |
| Reflection | "Mirrors things" | A matrix that flips space across an axis or plane. Determinant is -1. |
| Composition | "Do two things" | Multiplying transformation matrices to chain operations. Order matters: B @ A means apply A first, then B. |
| Eigenvector | "Special direction" | A direction that the matrix only scales, never rotates. The transformation's fingerprint. |
| Eigenvalue | "How much it stretches" | The scalar factor by which the matrix scales its eigenvector. Can be negative (flip) or complex (rotation). |
| Eigendecomposition | "Break the matrix apart" | Writing a matrix as V @ D @ V^(-1), separating it into its fundamental scaling directions and magnitudes. |
| Determinant | "A single number from a matrix" | The factor by which the transformation scales area (2D) or volume (3D). Zero means the transformation is irreversible. |
| Characteristic equation | "Where eigenvalues come from" | det(A - lambda * I) = 0. The polynomial whose roots are the eigenvalues. |
进一步阅读
- 3Blue1Brown: Linear Transformations视觉直觉来了解矩阵如何重塑空间
- 3Blue1Brown: Eigenvectors and Eigenvalues它们是对自向量对几何的最佳视觉解释.
- MIT 18.06 Lecture 21: Eigenvalues and Eigenvectors吉尔伯特·斯特朗的经典治疗方法
This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.
Browse the complete course catalog or open this lesson on GitHub.