调解神经网络
Type: Build
Languages: Python, PyTorch
Prerequisites: Phase 03 Lessons 01-10 (especially backpropagation, loss functions, optimizers)
Time: ~90 minutes
学习目标
- 使用系统性调试策略诊断常见的神经网络故障 (NaN损失,平损曲线,过度适应,振荡)
- 应用"超适合一批"技术来验证模型架构和训练循环是否正确
- 检查梯度大小,激活分布和重量规范,以确定消失/爆炸梯度问题
- 建立一个调试检查清单,涵盖数据管道,模型架构,损失函数,优化器和学习速度问题
问题
传统软件在破产时会崩. 零指标会产生例外. 编译时类型不匹配失败. 一次一次错误会产生明显错误输出.
网络不会给你提供这种奢.
破产的神经网络运行到完成, 打印损失值, 损失可能会减少. 预测可能看起来是可行的. 但模型是默默地错误的 - - 学习快捷方式,记住噪音,或者接近无用的本地最低点. 谷歌研究人员估计,在ML调试时间中的60-70%,用于"沉默"的错误,
工作模型与破产模型的区别通常是一个错误的线条:一个缺失的线条.zero_grad()们的学习速度是10倍. 们的神经网络训练的食谱 (2019) 开始了这样:"最常见的神经网络错误是不会崩的错误.
这课教你如何找到这些虫子.
概念
错误的思维方式
忘记打印和试验调试. 网络调试需要系统方法,因为反循环很慢 (每次训练跑的分钟到小时),症状很模糊 (坏损失可能意味着20种不同的东西).
黄金规则:start simple, add complexity one piece at a time, and verify each piece independently.
flowchart TD
A["Loss not decreasing"] --> B{"Check learning rate"}
B -->|"Too high"| C["Loss oscillates or explodes"]
B -->|"Too low"| D["Loss barely moves"]
B -->|"Reasonable"| E{"Check gradients"}
E -->|"All zeros"| F["Dead ReLUs or vanishing gradients"]
E -->|"NaN/Inf"| G["Exploding gradients"]
E -->|"Normal"| H{"Check data pipeline"}
H -->|"Labels shuffled"| I["Random-chance accuracy"]
H -->|"Preprocessing bug"| J["Model learns noise"]
H -->|"Data is fine"| K{"Check architecture"}
K -->|"Too small"| L["Underfitting"]
K -->|"Too deep"| M["Optimization difficulty"]症状1: 损失不减少
训练循环是流行的,时代过去了,损失保持平稳或动.
Wrong learning rate.对于亚当来说,从1e-3开始.对于SGD来说,从1e-1或1e-2开始. 在得出其他错误之前,总是尝试3个学习率,每个学习率跨越10倍 (例如,1e-2,1e-3,1e-4)
Dead ReLUs.如果一个ReLU神经元接收了大量的负输入,它输出0并且其梯度是0.它永远不会再次激活.如果足够的神经元死亡,网络无法学习.检查:印出每一个ReLU层后的精确0的激活率.如果50%以上死亡,则切换到LeakyReLU或降低学习速度.
Vanishing gradients.在具有sigmoid或tanh激活的深层网络中,渐变随着向后传播而呈指数缩小.到达第一层时,它们是0.第一层停止学习.
Exploding gradients.降梯率增长指数.在RNN和非常深度网络中常见.损失跳到NaN.torch.nn.utils.clip_grad_norm_),降低学习率,或增加正常化.
症状2:损失减少,但模式不好
损失下降了,训练精度达到99%,但测试精度是55%.
Overfitting.模型记忆训练数据而不是学习模式.训练和验证损失之间的差距随着时间的推移增加.
Data leakage.测试数据泄露到训练中.精度可疑高.常见原因:在分开之前混动,预处理数据集的统计数据,在分开之间复制样本. 修复:分开第一,预处理第二,检查复制.
Label errors.大多数真实数据集中的5-10%的标签是错误的 (Northcutt等同,2021年"测试集中的普遍标签错误").模型学习噪音.修复:使用自信学习来找到和修复错误标签的例子,或使用损失缩小忽略高损失样本.
症状3: NaN或Inf在损失中
损失值将成为nan或inf训练已经结束了.
Learning rate too high.更新速度超越了重量爆炸.
log(0) or log(negative).跨缩损失计算器log(p)如果你的模型输出精确的0或负概率,记录会爆炸.[eps, 1-eps]在哪里eps=1e-7现在,我们要去.
Division by zero.批量正常化按标准偏差分开. 一批具有恒定值的批量具有 std=0. 修正:将epsilon添加到分母中 (PyTorch默认执行,但可能不会实现).
Numerical overflow.输入了大量激活exp()解决问题:在指数化之前减去最大值 (日积-exp技巧).
技术1:渐进检查
根据分析的差异,你必须将分析的差异 (从后方向) 进行比较.
参数的数值梯度w其他:
grad_numerical = (loss(w + eps) - loss(w - eps)) / (2 * eps)协议指标 (相对差异):
rel_diff = |grad_analytical - grad_numerical| / max(|grad_analytical|, |grad_numerical|, 1e-8)如果rel_diff < 1e-5答案是正确的.rel_diff > 1e-3几乎肯定是个虫子.
flowchart LR
A["Parameter w"] --> B["w + eps"]
A --> C["w - eps"]
B --> D["Forward pass"]
C --> E["Forward pass"]
D --> F["loss+"]
E --> G["loss-"]
F --> H["(loss+ - loss-) / 2eps"]
G --> H
H --> I["Compare to backprop gradient"]技术2:激活统计
训练期间,监测每层激活后的平均和标准偏差.健康网络保持在 0 附近的平均和 1 附近的 STD (正常化后) 或至少有界限的激活.
| Health indicator | Mean | Std | Diagnosis |
|---|---|---|---|
| Healthy | ~0 | ~1 | Network is learning normally |
| Saturated | >>0 or <<0 | ~0 | Activations stuck at extreme values |
| Dead | 0 | 0 | Neurons are dead (all zeros) |
| Exploding | >>10 | >>10 | Activations growing without bound |
技术3:渐进流量可视化
图表每层的平均梯度大小.在健康的网络中,梯度大小应该在各层之间大致相同.如果早期层的梯度比后层小1000倍,则有消失的梯度.
graph LR
subgraph "Healthy Gradient Flow"
L1["Layer 1<br/>grad: 0.05"] --- L2["Layer 2<br/>grad: 0.04"] --- L3["Layer 3<br/>grad: 0.06"] --- L4["Layer 4<br/>grad: 0.05"]
endgraph LR
subgraph "Vanishing Gradient Flow"
V1["Layer 1<br/>grad: 0.0001"] --- V2["Layer 2<br/>grad: 0.003"] --- V3["Layer 3<br/>grad: 0.02"] --- V4["Layer 4<br/>grad: 0.08"]
end技术4:超级适应一批试验
对于深度学习来说,这是最重要的调试技术.
运行一个小批量 (8-32个样本). 训练100次以上. 损失应该达到接近零,训练精度应该达到100%. 如果没有,你的模型或训练循环有基本的错误 - - 不要继续进行完整的训练.
这项测试发现:
- 破损函数
- 破碎的后行
- 建筑物太小,无法代表数据
- 没有连接到模型参数的优化器
- 错误地调整数据和标签
这需要30秒的时间才能运行,
技术5:学习率查询器
莱斯利·史密斯 (2017) 提出在一个时代内将学习率从非常小 (1e-7) 扫到非常大 (10) 扫描,同时记录损失. 剧情损失与学习率.最佳学习率大约是10倍小于损失开始减速速度的速度.
graph TD
subgraph "LR Finder Plot"
direction LR
A["1e-7: loss=2.3"] --> B["1e-5: loss=2.3"]
B --> C["1e-3: loss=1.8"]
C --> D["1e-2: loss=0.9 -- steepest"]
D --> E["1e-1: loss=0.5"]
E --> F["1.0: loss=NaN -- too high"]
end在本例中最好的LR: ~1e-3 (最点前一个大小顺序).
常见的皮托尔奇虫
这些是PyTorch社区最多时间浪费的虫子:
| Bug | Symptom | Fix |
|---|---|---|
Forgetting optimizer.zero_grad() | Gradients accumulate across batches, loss oscillates | Add optimizer.zero_grad() before loss.backward() |
Forgetting model.eval() at test time | Dropout and batch norm behave differently, test accuracy varies between runs | Add model.eval() and torch.no_grad() |
| Wrong tensor shapes | Silent broadcasting produces wrong results, no error | Print shapes after every operation during debugging |
| CPU/GPU mismatch | RuntimeError: expected CUDA tensor | Use .to(device) on model AND data |
| Not detaching tensors | Computation graph grows forever, OOM | Use .detach() or with torch.no_grad() |
| In-place operations breaking autograd | RuntimeError: modified by in-place operation | Replace x += 1 with x = x + 1 |
| Data not normalized | Loss stuck at random-chance level | Normalize inputs to mean=0, std=1 |
| Labels as wrong dtype | Cross-entropy expects Long, got Float | Cast labels: labels.long() |
导师调试表
| Symptom | Likely cause | First thing to try |
|---|---|---|
| Loss stuck at -log(1/num_classes) | Model predicting uniform distribution | Check data pipeline, verify labels match inputs |
| Loss NaN after a few steps | Learning rate too high | Reduce LR by 10x |
| Loss NaN immediately | log(0) or division by zero | Add epsilon to log/division operations |
| Loss oscillating wildly | LR too high or batch size too small | Reduce LR, increase batch size |
| Loss decreasing then plateaus | LR too high for fine-tuning phase | Add LR schedule (cosine or step decay) |
| Training acc high, test acc low | Overfitting | Add dropout, weight decay, more data |
| Training acc = test acc = chance | Model not learning anything | Run overfit-one-batch test |
| Training acc = test acc but both low | Underfitting | Bigger model, more layers, more features |
| Gradients all zero | Dead ReLUs or detached computation graph | Switch to LeakyReLU, check .requires_grad |
| Out of memory during training | Batch too large or graph not freed | Reduce batch size, use torch.no_grad() for eval |
建立它
检测工具包监测激活,梯度和损失曲线.
步骤1:网络调试器类
入 PyTorch 模型,以记录每层的激活和梯度统计.
pythonimport torch
import torch.nn as nn
import math
class NetworkDebugger:
def __init__(self, model):
self.model = model
self.activation_stats = {}
self.gradient_stats = {}
self.loss_history = []
self.lr_losses = []
self.hooks = []
self._register_hooks()
def _register_hooks(self):
for name, module in self.model.named_modules():
if isinstance(module, (nn.Linear, nn.Conv2d, nn.ReLU, nn.LeakyReLU)):
hook = module.register_forward_hook(self._make_activation_hook(name))
self.hooks.append(hook)
hook = module.register_full_backward_hook(self._make_gradient_hook(name))
self.hooks.append(hook)
def _make_activation_hook(self, name):
def hook(module, input, output):
with torch.no_grad():
out = output.detach().float()
self.activation_stats[name] = {
"mean": out.mean().item(),
"std": out.std().item(),
"fraction_zero": (out == 0).float().mean().item(),
"min": out.min().item(),
"max": out.max().item(),
}
return hook
def _make_gradient_hook(self, name):
def hook(module, grad_input, grad_output):
if grad_output[0] is not None:
with torch.no_grad():
grad = grad_output[0].detach().float()
self.gradient_stats[name] = {
"mean": grad.mean().item(),
"std": grad.std().item(),
"abs_mean": grad.abs().mean().item(),
"max": grad.abs().max().item(),
}
return hook
def record_loss(self, loss_value):
self.loss_history.append(loss_value)
def check_loss_health(self):
if len(self.loss_history) < 2:
return "NOT_ENOUGH_DATA"
recent = self.loss_history[-10:]
if any(math.isnan(v) or math.isinf(v) for v in recent):
return "NAN_OR_INF"
if len(self.loss_history) >= 20:
first_half = sum(self.loss_history[:10]) / 10
second_half = sum(self.loss_history[-10:]) / 10
if second_half >= first_half * 0.99:
return "NOT_DECREASING"
if len(recent) >= 5:
diffs = [recent[i+1] - recent[i] for i in range(len(recent)-1)]
if max(diffs) - min(diffs) > 2 * abs(sum(diffs) / len(diffs)):
return "OSCILLATING"
return "HEALTHY"
def check_activations(self):
issues = []
for name, stats in self.activation_stats.items():
if stats["fraction_zero"] > 0.5:
issues.append(f"DEAD_NEURONS: {name} has {stats['fraction_zero']:.0%} zero activations")
if abs(stats["mean"]) > 10:
issues.append(f"EXPLODING_ACTIVATIONS: {name} mean={stats['mean']:.2f}")
if stats["std"] < 1e-6:
issues.append(f"COLLAPSED_ACTIVATIONS: {name} std={stats['std']:.2e}")
return issues if issues else ["HEALTHY"]
def check_gradients(self):
issues = []
grad_magnitudes = []
for name, stats in self.gradient_stats.items():
grad_magnitudes.append((name, stats["abs_mean"]))
if stats["abs_mean"] < 1e-7:
issues.append(f"VANISHING_GRADIENT: {name} abs_mean={stats['abs_mean']:.2e}")
if stats["abs_mean"] > 100:
issues.append(f"EXPLODING_GRADIENT: {name} abs_mean={stats['abs_mean']:.2e}")
if len(grad_magnitudes) >= 2:
first_mag = grad_magnitudes[0][1]
last_mag = grad_magnitudes[-1][1]
if last_mag > 0 and first_mag / last_mag > 100:
issues.append(f"GRADIENT_RATIO: first/last = {first_mag/last_mag:.0f}x (vanishing)")
return issues if issues else ["HEALTHY"]
def print_report(self):
print("\n=== NETWORK DEBUGGER REPORT ===")
print(f"\nLoss health: {self.check_loss_health()}")
if self.loss_history:
print(f" Last 5 losses: {[f'{v:.4f}' for v in self.loss_history[-5:]]}")
print("\nActivation diagnostics:")
for item in self.check_activations():
print(f" {item}")
print("\nGradient diagnostics:")
for item in self.check_gradients():
print(f" {item}")
print("\nPer-layer activation stats:")
for name, stats in self.activation_stats.items():
print(f" {name}: mean={stats['mean']:.4f} std={stats['std']:.4f} zero={stats['fraction_zero']:.1%}")
print("\nPer-layer gradient stats:")
for name, stats in self.gradient_stats.items():
print(f" {name}: abs_mean={stats['abs_mean']:.2e} max={stats['max']:.2e}")
def remove_hooks(self):
for hook in self.hooks:
hook.remove()
self.hooks.clear()第二步:一批过度适应的测试
pythondef overfit_one_batch(model, x_batch, y_batch, criterion, lr=0.01, steps=200):
optimizer = torch.optim.Adam(model.parameters(), lr=lr)
model.train()
print("\n=== OVERFIT ONE BATCH TEST ===")
print(f"Batch size: {x_batch.shape[0]}, Steps: {steps}")
for step in range(steps):
optimizer.zero_grad()
output = model(x_batch)
loss = criterion(output, y_batch)
loss.backward()
optimizer.step()
if step % 50 == 0 or step == steps - 1:
with torch.no_grad():
preds = (output > 0).float() if output.shape[-1] == 1 else output.argmax(dim=1)
targets = y_batch if y_batch.dim() == 1 else y_batch.squeeze()
acc = (preds.squeeze() == targets).float().mean().item()
print(f" Step {step:3d} | Loss: {loss.item():.6f} | Accuracy: {acc:.1%}")
final_loss = loss.item()
if final_loss > 0.1:
print(f"\n FAIL: Loss did not converge ({final_loss:.4f}). Model or training loop is broken.")
return False
print(f"\n PASS: Loss converged to {final_loss:.6f}")
return True步骤3:学习率查询器
pythondef find_learning_rate(model, x_data, y_data, criterion, start_lr=1e-7, end_lr=10, steps=100):
import copy
original_state = copy.deepcopy(model.state_dict())
optimizer = torch.optim.SGD(model.parameters(), lr=start_lr)
lr_mult = (end_lr / start_lr) ** (1 / steps)
model.train()
results = []
best_loss = float("inf")
current_lr = start_lr
print("\n=== LEARNING RATE FINDER ===")
for step in range(steps):
optimizer.zero_grad()
output = model(x_data)
loss = criterion(output, y_data)
if math.isnan(loss.item()) or loss.item() > best_loss * 10:
break
best_loss = min(best_loss, loss.item())
results.append((current_lr, loss.item()))
loss.backward()
optimizer.step()
current_lr *= lr_mult
for param_group in optimizer.param_groups:
param_group["lr"] = current_lr
model.load_state_dict(original_state)
if len(results) < 10:
print(" Could not complete LR sweep -- loss diverged too quickly")
return results
min_loss_idx = min(range(len(results)), key=lambda i: results[i][1])
suggested_lr = results[max(0, min_loss_idx - 10)][0]
print(f" Swept {len(results)} steps from {start_lr:.0e} to {results[-1][0]:.0e}")
print(f" Minimum loss {results[min_loss_idx][1]:.4f} at lr={results[min_loss_idx][0]:.2e}")
print(f" Suggested learning rate: {suggested_lr:.2e}")
return results步骤4: 测量度
pythondef _flat_to_multi_index(flat_idx, shape):
multi_idx = []
remaining = flat_idx
for dim in reversed(shape):
multi_idx.insert(0, remaining % dim)
remaining //= dim
return tuple(multi_idx)
def gradient_check(model, x, y, criterion, eps=1e-4):
model.train()
x_double = x.double()
y_double = y.double()
model_double = model.double()
print("\n=== GRADIENT CHECK ===")
overall_max_diff = 0
checked = 0
for name, param in model_double.named_parameters():
if not param.requires_grad:
continue
layer_max_diff = 0
model_double.zero_grad()
output = model_double(x_double)
loss = criterion(output, y_double)
loss.backward()
analytical_grad = param.grad.clone()
num_checks = min(5, param.numel())
for i in range(num_checks):
idx = _flat_to_multi_index(i, param.shape)
original = param.data[idx].item()
param.data[idx] = original + eps
with torch.no_grad():
loss_plus = criterion(model_double(x_double), y_double).item()
param.data[idx] = original - eps
with torch.no_grad():
loss_minus = criterion(model_double(x_double), y_double).item()
param.data[idx] = original
numerical = (loss_plus - loss_minus) / (2 * eps)
analytical = analytical_grad[idx].item()
denom = max(abs(numerical), abs(analytical), 1e-8)
rel_diff = abs(numerical - analytical) / denom
layer_max_diff = max(layer_max_diff, rel_diff)
checked += 1
overall_max_diff = max(overall_max_diff, layer_max_diff)
status = "OK" if layer_max_diff < 1e-5 else "MISMATCH"
print(f" {name}: max_rel_diff={layer_max_diff:.2e} [{status}]")
model.float()
print(f"\n Checked {checked} parameters")
if overall_max_diff < 1e-5:
print(" PASS: Gradients match (rel_diff < 1e-5)")
elif overall_max_diff < 1e-3:
print(" WARN: Small differences (1e-5 < rel_diff < 1e-3)")
else:
print(" FAIL: Gradient mismatch detected (rel_diff > 1e-3)")
return overall_max_diff步骤5:故意破解网络
现在将工具包应用到破产的网络,
pythondef demo_broken_networks():
torch.manual_seed(42)
x = torch.randn(64, 10)
y = (x[:, 0] > 0).long()
print("\n" + "=" * 60)
print("BUG 1: Learning rate too high (lr=10)")
print("=" * 60)
model1 = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 2))
debugger1 = NetworkDebugger(model1)
optimizer1 = torch.optim.SGD(model1.parameters(), lr=10.0)
criterion = nn.CrossEntropyLoss()
for step in range(20):
optimizer1.zero_grad()
out = model1(x)
loss = criterion(out, y)
debugger1.record_loss(loss.item())
loss.backward()
optimizer1.step()
debugger1.print_report()
debugger1.remove_hooks()
print("\n" + "=" * 60)
print("BUG 2: Dead ReLUs from bad initialization")
print("=" * 60)
model2 = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 32), nn.ReLU(), nn.Linear(32, 2))
with torch.no_grad():
for m in model2.modules():
if isinstance(m, nn.Linear):
m.weight.fill_(-1.0)
m.bias.fill_(-5.0)
debugger2 = NetworkDebugger(model2)
optimizer2 = torch.optim.Adam(model2.parameters(), lr=1e-3)
for step in range(50):
optimizer2.zero_grad()
out = model2(x)
loss = criterion(out, y)
debugger2.record_loss(loss.item())
loss.backward()
optimizer2.step()
debugger2.print_report()
debugger2.remove_hooks()
print("\n" + "=" * 60)
print("BUG 3: Missing zero_grad (gradients accumulate)")
print("=" * 60)
model3 = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 2))
debugger3 = NetworkDebugger(model3)
optimizer3 = torch.optim.SGD(model3.parameters(), lr=0.01)
for step in range(50):
out = model3(x)
loss = criterion(out, y)
debugger3.record_loss(loss.item())
loss.backward()
optimizer3.step()
debugger3.print_report()
debugger3.remove_hooks()
print("\n" + "=" * 60)
print("HEALTHY NETWORK: Correct setup for comparison")
print("=" * 60)
model_good = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 2))
debugger_good = NetworkDebugger(model_good)
optimizer_good = torch.optim.Adam(model_good.parameters(), lr=1e-3)
for step in range(50):
optimizer_good.zero_grad()
out = model_good(x)
loss = criterion(out, y)
debugger_good.record_loss(loss.item())
loss.backward()
optimizer_good.step()
debugger_good.print_report()
debugger_good.remove_hooks()
print("\n" + "=" * 60)
print("OVERFIT-ONE-BATCH TEST (healthy model)")
print("=" * 60)
model_test = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 2))
overfit_one_batch(model_test, x[:8], y[:8], criterion)
print("\n" + "=" * 60)
print("LEARNING RATE FINDER")
print("=" * 60)
model_lr = nn.Sequential(nn.Linear(10, 32), nn.ReLU(), nn.Linear(32, 2))
find_learning_rate(model_lr, x, y, criterion)
print("\n" + "=" * 60)
print("GRADIENT CHECK")
print("=" * 60)
model_grad = nn.Sequential(nn.Linear(10, 8), nn.ReLU(), nn.Linear(8, 2))
gradient_check(model_grad, x[:4], y[:4], criterion)用它
嵌入式 PyTorch 工具
pythonimport torch
import torch.nn as nn
model = nn.Sequential(
nn.Linear(768, 256),
nn.ReLU(),
nn.Linear(256, 10),
)
with torch.autograd.detect_anomaly():
output = model(input_tensor)
loss = criterion(output, target)
loss.backward()
for name, param in model.named_parameters():
if param.grad is not None:
print(f"{name}: grad_mean={param.grad.abs().mean():.2e}")重量与偏差的整合
pythonimport wandb
wandb.init(project="debug-training")
for epoch in range(100):
loss = train_one_epoch()
wandb.log({
"loss": loss,
"lr": optimizer.param_groups[0]["lr"],
"grad_norm": torch.nn.utils.clip_grad_norm_(model.parameters(), float("inf")),
})
for name, param in model.named_parameters():
if param.grad is not None:
wandb.log({f"grad/{name}": wandb.Histogram(param.grad.cpu().numpy())})电压板
pythonfrom torch.utils.tensorboard import SummaryWriter
writer = SummaryWriter("runs/debug_experiment")
for epoch in range(100):
loss = train_one_epoch()
writer.add_scalar("Loss/train", loss, epoch)
for name, param in model.named_parameters():
writer.add_histogram(f"weights/{name}", param, epoch)
if param.grad is not None:
writer.add_histogram(f"gradients/{name}", param.grad, epoch)检查清单 (在充分训练之前)
- 试试一次,如果失败,就停止.
- 打印模型总结-- 验证参数数量是合理的.
- 运行一个随机数据的前进传输--检查输出形状.
- 列车5个时代,检查损失减少.
- 检查激活数据,没有死层,没有爆炸.
- 检查梯度流量,没有消失,没有爆炸.
- 检查数据管道-- 打印5个随机样本,
运送它
这一课产生了:
outputs/prompt-nn-debugger.md-- 诊断神经网络训练失败的提示outputs/skill-debug-checklist.md-- 调试训练问题决策树检查清单
调试的主要部署模式:
- 添加监控子到生产训练脚本中
- 每次N步骤的W&B或TensorBoard日志激活和梯度统计
- 实现自动警告,即纳损失,死神经元 (>80%零) 或梯度爆炸
- 总是在修改架构或数据管道时,始终执行过度配件一批测试
运动
- Add an exploding gradient detector.修改
NetworkDebugger检测梯度超过门时,并自动提示梯度切割值. 在20层网络上测试,没有正常化.
- Build a dead neuron resurrector.写一个识别死 ReLU 神经元的函数 (总是输出0),并通过凯明初始化重新启动其进来的重量. 显示这恢复了神经元的70%以上死亡的网络.
- Implement the learning rate finder with plotting.延长时间
find_learning_rate通过 matplotlib 保存结果作为 CSV,并写一个单独的脚本,该脚本读取CSV,并显示LR与损失曲线.在CIFAR-10上确定ResNet-18的最佳LR.
- Create a data pipeline validator.写一个检查数据中的函数:在列车/测试分区间中复制样本,标签分布不平衡 (>10:1比率),输入正常化 (平均接近0, std接近1),以及数据中的NaN/Inf值. 运行在故意破坏的数据集上.
- Debug a real failure.根据10课的迷你框架,引入一个微妙的错误 (例如,将权重矩阵转换向后),并使用梯度检查,以确定哪个参数有不正确的梯度. 记录调试过程.
关键词
| Term | What people say | What it actually means |
|---|---|---|
| Silent bug | "It runs but gives bad results" | A bug that produces no error but degrades model quality -- the dominant failure mode in ML |
| Dead ReLU | "The neurons died" | A ReLU neuron whose input is always negative, so it outputs 0 and receives 0 gradient permanently |
| Vanishing gradients | "Early layers stop learning" | Gradients shrink exponentially through layers, making weights in early layers effectively frozen |
| Exploding gradients | "Loss went to NaN" | Gradients grow exponentially through layers, causing weight updates so large they overflow |
| Gradient checking | "Verify backprop is correct" | Comparing analytical gradients from backprop to numerical gradients from finite differences |
| Overfit-one-batch | "The most important debug test" | Training on a single small batch to verify the model CAN learn -- if it cannot, something is fundamentally broken |
| LR finder | "Sweep to find the right learning rate" | Exponentially increasing the learning rate over one epoch and picking the rate just before loss diverges |
| Data leakage | "Test data leaked into training" | When information from the test set contaminates training, producing artificially high accuracy |
| Activation statistics | "Monitor layer health" | Tracking mean, std, and zero-fraction of each layer's output to detect dead, saturated, or exploding neurons |
| Gradient clipping | "Cap the gradient magnitude" | Scaling gradients down when their norm exceeds a threshold, preventing exploding gradient updates |
进一步阅读
- 史密斯, "神经网络培训周期性学习率" (2017) - - 引入学习率范围测试的论文 (LR寻找器)
- 诺斯卡特等",测试组中的普遍标签错误破坏机器学习基准" (2021) -- 证明ImageNet,CIFAR-10和其他主要基准中的3-6%的标签是错误的
- 张等人",理解深度学习需要重新思考通用化" (2017) -- 论文显示神经网络可以记住随机标签,
- 关于 PyTorch 的文件
torch.autograd.detect_anomaly其他torch.autograd.set_detect_anomaly用于内置的NAN/Inf检测
This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.
Browse the complete course catalog or open this lesson on GitHub.