矢量,矩阵和运算
Type: Build
Languages: Python, Julia
Prerequisites: Phase 1, Lesson 01 (Linear Algebra Intuition)
Time: ~60 minutes
学习目标
- 构建一个矩阵类,以元素理性操作,矩阵乘法,转换,定数和反
- 区分元素式乘法与矩阵乘法,并解释每一个乘法是什么时候适用的
- 实现单一密集神经网络层 (
relu(W @ x + b)) 仅使用从零开始的矩阵类 - 解释广播规则以及神经网络框架中偏见加算的运作方式
问题
你想建立一个神经网络,你读出代码,看看这个:
output = activation(weights @ input + bias)这@它们是矩阵乘法.weights它们是矩阵.input如果您不知道这些操作是什么,这个线是魔术.如果您知道,这是一个层的整个前进通过在三个操作.
你的模型处理的每一个图像都是像素值的矩阵.每个嵌入的字符都是矢量.每个神经网络的每个层都是矩阵转换.你不能在矩阵操作中流利地构建人工智能系统,就像你不能在理解变量的情况下编写代码一样.
这一课从零开始就能让你变得流利.
概念
矢量:排列列的数目列表
矢量是指向和大小的数量列表.在人工智能中,矢量代表数据点,特征或参数.
v = [3, 4] -- a 2D vector
w = [1, 0, -2] -- a 3D vector两个维向量[3, 4]它们的长度 (大小) 是5 (三角形3-4-5).
矩阵:数字网
矩阵是一个二维格格. 列和列.一个m x n矩阵有m列和n列.
A = | 1 2 3 | -- 2x3 matrix (2 rows, 3 columns)
| 4 5 6 |在神经网络中,重量矩阵将输入向量转化为输出向量.一个有784个输入和128个输出的层使用 128x784的重量矩阵.
形状为什么重要
矩阵乘法有一个严格的规则:(m x n) @ (n x p) = (m x p)内部的尺寸必须相匹配.
(128 x 784) @ (784 x 1) = (128 x 1)
weights input output
Inner dimensions: 784 = 784 -- valid如果在 PyTorch 中出现了形状不匹配错误,
运营地图
| Operation | What it does | Neural network use |
|---|---|---|
| Addition | Element-wise combine | Adding bias to output |
| Scalar multiply | Scale every element | Learning rate * gradients |
| Matrix multiply | Transform vectors | Layer forward pass |
| Transpose | Flip rows and columns | Backpropagation |
| Determinant | Single number summary | Checking invertibility |
| Inverse | Undo a transformation | Solving linear systems |
| Identity | Do-nothing matrix | Initialization, residual connections |
元素智能与矩阵乘法
这种区别会让初学者不断地陷入困境.
按元素的角度,乘以相匹配的位置.
| 1 2 | | 5 6 | | 5 12 |
| 3 4 | * | 7 8 | = | 21 32 |矩阵乘法:列和列的点产量.内面尺寸必须匹配.
| 1 2 | | 5 6 | | 1*5+2*7 1*6+2*8 | | 19 22 |
| 3 4 | @ | 7 8 | = | 3*5+4*7 3*6+4*8 | = | 43 50 |不同的操作,不同的结果,不同的规则.
广播
输出矩阵中添加偏差向量时,形状不匹配.
| 1 2 3 | + [10, 20, 30]
| 4 5 6 |
Broadcasting stretches the vector across rows:
| 1 2 3 | | 10 20 30 | | 11 22 33 |
| 4 5 6 | + | 10 20 30 | = | 14 25 36 |任何现代框架都会自动做到这一点.
建立它
步骤1:向量类
pythonclass Vector:
def __init__(self, data):
self.data = list(data)
self.size = len(self.data)
def __repr__(self):
return f"Vector({self.data})"
def __add__(self, other):
return Vector([a + b for a, b in zip(self.data, other.data)])
def __sub__(self, other):
return Vector([a - b for a, b in zip(self.data, other.data)])
def __mul__(self, scalar):
return Vector([x * scalar for x in self.data])
def dot(self, other):
return sum(a * b for a, b in zip(self.data, other.data))
def magnitude(self):
return sum(x ** 2 for x in self.data) ** 0.5步骤2: 核心操作的矩阵类
pythonclass Matrix:
def __init__(self, data):
self.data = [list(row) for row in data]
self.rows = len(self.data)
self.cols = len(self.data[0])
self.shape = (self.rows, self.cols)
def __repr__(self):
rows_str = "\n ".join(str(row) for row in self.data)
return f"Matrix({self.shape}):\n {rows_str}"
def __add__(self, other):
return Matrix([
[self.data[i][j] + other.data[i][j] for j in range(self.cols)]
for i in range(self.rows)
])
def __sub__(self, other):
return Matrix([
[self.data[i][j] - other.data[i][j] for j in range(self.cols)]
for i in range(self.rows)
])
def scalar_multiply(self, scalar):
return Matrix([
[self.data[i][j] * scalar for j in range(self.cols)]
for i in range(self.rows)
])
def element_wise_multiply(self, other):
return Matrix([
[self.data[i][j] * other.data[i][j] for j in range(self.cols)]
for i in range(self.rows)
])
def matmul(self, other):
return Matrix([
[
sum(self.data[i][k] * other.data[k][j] for k in range(self.cols))
for j in range(other.cols)
]
for i in range(self.rows)
])
def transpose(self):
return Matrix([
[self.data[j][i] for j in range(self.rows)]
for i in range(self.cols)
])
def determinant(self):
if self.shape == (1, 1):
return self.data[0][0]
if self.shape == (2, 2):
return self.data[0][0] * self.data[1][1] - self.data[0][1] * self.data[1][0]
det = 0
for j in range(self.cols):
minor = Matrix([
[self.data[i][k] for k in range(self.cols) if k != j]
for i in range(1, self.rows)
])
det += ((-1) ** j) * self.data[0][j] * minor.determinant()
return det
def inverse_2x2(self):
det = self.determinant()
if det == 0:
raise ValueError("Matrix is singular, no inverse exists")
return Matrix([
[self.data[1][1] / det, -self.data[0][1] / det],
[-self.data[1][0] / det, self.data[0][0] / det]
])
@staticmethod
def identity(n):
return Matrix([
[1 if i == j else 0 for j in range(n)]
for i in range(n)
])步骤3: 让它发挥作用
pythonA = Matrix([[1, 2], [3, 4]])
B = Matrix([[5, 6], [7, 8]])
print("A + B =", (A + B).data)
print("A @ B =", A.matmul(B).data)
print("A^T =", A.transpose().data)
print("det(A) =", A.determinant())
print("A^-1 =", A.inverse_2x2().data)
I = Matrix.identity(2)
print("A @ A^-1 =", A.matmul(A.inverse_2x2()).data)步骤4:连接到神经网络
pythonimport random
inputs = Matrix([[0.5], [0.8], [0.2]])
weights = Matrix([
[random.uniform(-1, 1) for _ in range(3)]
for _ in range(2)
])
bias = Matrix([[0.1], [0.1]])
def relu_matrix(m):
return Matrix([[max(0, val) for val in row] for row in m.data])
pre_activation = weights.matmul(inputs) + bias
output = relu_matrix(pre_activation)
print(f"Input shape: {inputs.shape}")
print(f"Weight shape: {weights.shape}")
print(f"Output shape: {output.shape}")
print(f"Output: {output.data}")这是一个单密层:output = relu(W @ x + b)每个神经网络的密集层都会做这么做.
用它
平在更少的线上和更快的规模上完成了以上的一切.
pythonimport numpy as np
A = np.array([[1, 2], [3, 4]])
B = np.array([[5, 6], [7, 8]])
print("A + B =\n", A + B)
print("A * B (element-wise) =\n", A * B)
print("A @ B (matrix multiply) =\n", A @ B)
print("A^T =\n", A.T)
print("det(A) =", np.linalg.det(A))
print("A^-1 =\n", np.linalg.inv(A))
print("I =\n", np.eye(2))
inputs = np.random.randn(3, 1)
weights = np.random.randn(2, 3)
bias = np.array([[0.1], [0.1]])
output = np.maximum(0, weights @ inputs + bias)
print(f"\nNeural network layer: {weights.shape} @ {inputs.shape} = {output.shape}")
print(f"Output:\n{output}")其他@在Python调用中操作员__matmul__通过C和Fortan编写的优化BLAS程序实现了它.
在NumPy中播放:
pythonmatrix = np.array([[1, 2, 3], [4, 5, 6]])
bias = np.array([10, 20, 30])
print(matrix + bias)通过 NumPy 实现了自动播放1D偏差在两个行中.
运送它
通过几何直觉来教导矩阵操作.outputs/prompt-matrix-operations.md现在,我们要去.
在这个阶段建立的矩阵类是我们在3阶段的10课时建立的微神经网络框架的基础.
运动
- Verify the inverse.乘以
A @ A.inverse_2x2()现在我们可以通过两个不同的2x2矩阵来测试,然后确认我们得到了身份矩阵.
- Implement 3x3 inverse.通过结方法,将矩阵类扩展到计算3x3矩阵的逆数.
np.linalg.inv现在,我们要去.
- Build a two-layer network.通过使用您的矩阵类 (没有NumPy),创建一个两个层神经网络:输入 (3) ->隐藏 (4) ->输出 (2). 启动随机权重,运行前进传递,并验证所有形状是正确的.
关键词
| Term | What people say | What it actually means |
|---|---|---|
| Vector | "An arrow" | An ordered list of numbers. In AI: a point in high-dimensional space. |
| Matrix | "A table of numbers" | A linear transformation. It maps vectors from one space to another. |
| Matrix multiply | "Just multiply the numbers" | Dot products between every row of the first matrix and every column of the second. Order matters. |
| Transpose | "Flip it" | Swap rows and columns. Turns an m x n matrix into n x m. Critical in backpropagation. |
| Determinant | "Some number from the matrix" | Measures how much the matrix scales area (2D) or volume (3D). Zero means the transformation crushes a dimension. |
| Inverse | "Undo the matrix" | The matrix that reverses the transformation. Only exists when the determinant is not zero. |
| Identity matrix | "The boring matrix" | The matrix equivalent of multiplying by 1. Used in residual connections (ResNets). |
| Broadcasting | "Magic shape fixing" | Stretching a smaller array to match a larger one by repeating along missing dimensions. |
| Element-wise | "Regular multiplication" | Multiply matching positions. Both arrays must have the same shape (or be broadcastable). |
进一步阅读
- 3Blue1Brown: Essence of Linear Algebra- 视觉直觉对于每一个操作
- NumPy documentation on broadcasting- 准确的规则
- Stanford CS229 Linear Algebra Review- 简要参考 ML 特定线性代数
This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.
Browse the complete course catalog or open this lesson on GitHub.