终端和
Type: Learn
Languages: --
Prerequisites: Phase 0, Lesson 01
Time: ~35 minutes
学习目标
- 使用管道,转向,
grep从命令行过和处理训练日志 - 创建多个面板的持续tmux会议,同时进行训练和GPU监控
- 监控系统和GPU资源
htop现在nvtop其他nvidia-smi - 通过SSH将本地和远程机器之间的文件转移,
scp其他rsync
问题
你将在终端里花更多的时间,比任何编辑器. 训练运行,GPU监控,日志追踪,远程SSH会议,环境管理. 每个人工智能工作流都触及了贝. 如果你在这里慢,你在任何地方都会慢.
这堂课涵盖人工智能工作所需的终端技能. 没有Unix的历史. 没有深入的Bash脚本.
概念
graph TD
subgraph tmux["tmux session: training"]
subgraph top["Top row"]
P1["Pane 1: Training run<br/>python train.py<br/>Epoch 12/100 ..."]
P2["Pane 2: GPU monitor<br/>watch -n1 nvidia-smi<br/>GPU: 78% | Mem: 14/24G"]
end
P3["Pane 3: Logs + experiments<br/>tail -f logs/train.log | grep loss"]
end连接三件事,一个终端,你可以脱离,回家,再回 SSH,再连接.
建立它
步骤1:了解你的
检查你运行的炮弹:
bashecho $SHELL大多数系统使用bash或zsh两者都很好,这门课程中的命令都很好.
重要的事情:
bash# Move around
cd ~/projects/ai-engineering-from-scratch
pwd
ls -la
# History search (most useful shortcut you'll learn)
# Ctrl+R then type part of a previous command
# Press Ctrl+R again to cycle through matches
# Clear terminal
clear # or Ctrl+L
# Cancel a running command
# Ctrl+C
# Suspend a running command (resume with fg)
# Ctrl+Z步骤2:管道和转向
管道将命令连接在一起. 这就是你处理日志,过输出和链工具的方式.
bash# Count how many times "loss" appears in a log
cat train.log | grep "loss" | wc -l
# Extract just the loss values from training output
grep "loss:" train.log | awk '{print $NF}' > losses.txt
# Watch a log file update in real time, filtering for errors
tail -f train.log | grep --line-buffered "ERROR"
# Sort experiments by final accuracy
grep "final_accuracy" results/*.log | sort -t= -k2 -n -r
# Redirect stdout and stderr to separate files
python train.py > output.log 2> errors.log
# Redirect both to the same file
python train.py > train_full.log 2>&1你需要的三个转向:
| Symbol | What it does |
|---|---|
> | Write stdout to file (overwrite) |
>> | Append stdout to file |
2> | Write stderr to file |
2>&1 | Send stderr to same place as stdout |
| | Send stdout of one command as stdin to the next |
步骤3:背景过程
训练需要几个小时,你不想一直保持终端开放.
bash# Run in background (output still goes to terminal)
python train.py &
# Run in background, immune to hangup (closing terminal won't kill it)
nohup python train.py > train.log 2>&1 &
# Check what's running in background
jobs
ps aux | grep train.py
# Bring a background job to foreground
fg %1
# Kill a background process
kill %1
# or find its PID and kill that
kill $(pgrep -f "train.py")之间的区别&现在nohup其他screen现在,我们要去.tmux其他:
| Method | Survives terminal close? | Can reattach? |
|---|---|---|
command & | No | No |
nohup command & | Yes | No (check log file) |
screen / tmux | Yes | Yes |
任何超过几分钟的时间,使用tmux.
步骤4:
通过 tmux,您可以创建多个面板的持续终端会议.
bash# Install
# macOS
brew install tmux
# Ubuntu
sudo apt install tmux
# Start a named session
tmux new -s training
# Split horizontally
# Ctrl+B then "
# Split vertically
# Ctrl+B then %
# Navigate between panes
# Ctrl+B then arrow keys
# Detach (session keeps running)
# Ctrl+B then d
# Reattach
tmux attach -t training
# List sessions
tmux ls
# Kill a session
tmux kill-session -t training典型的人工智能工作流程:
bashtmux new -s train
# Pane 1: start training
python train.py --epochs 100 --lr 1e-4
# Ctrl+B, " to split, then run GPU monitor
watch -n1 nvidia-smi
# Ctrl+B, % to split vertically, tail the logs
tail -f logs/experiment.log
# Now detach with Ctrl+B, d
# SSH out, go get coffee, come back
# tmux attach -t train步骤5:使用 htop 和 nvtop 监测
bash# System processes (better than top)
htop
# GPU processes (if you have NVIDIA GPU)
# Install: sudo apt install nvtop (Ubuntu) or brew install nvtop (macOS)
nvtop
# Quick GPU check without nvtop
nvidia-smi
# Watch GPU usage update every second
watch -n1 nvidia-smi
# See which processes are using the GPU
nvidia-smi --query-compute-apps=pid,name,used_memory --format=csvhtop您将使用的键链:
F6或>按列排序 (按内存排序以查找内存泄漏)F5切换树视图 (见儿童过程)F9杀死一个过程/搜索过程名称
步骤 6:远程GPU盒的SSH
当你租用云GPU (Lambda, RunPod,Vast.ai) 时,你通过SSH连接.
bash# Basic connection
ssh user@gpu-box-ip
# With a specific key
ssh -i ~/.ssh/my_gpu_key user@gpu-box-ip
# Copy files to remote
scp model.pt user@gpu-box-ip:~/models/
# Copy files from remote
scp user@gpu-box-ip:~/results/metrics.json ./
# Sync a whole directory (faster for many files)
rsync -avz ./data/ user@gpu-box-ip:~/data/
# Port forward (access remote Jupyter/TensorBoard locally)
ssh -L 8888:localhost:8888 user@gpu-box-ip
# Now open localhost:8888 in your browser
# SSH config for convenience
# Add to ~/.ssh/config:
# Host gpu
# HostName 192.168.1.100
# User ubuntu
# IdentityFile ~/.ssh/gpu_key
#
# Then just:
# ssh gpu步骤7:人工智能工作的有用称
加入这些~/.bashrc或~/.zshrc其他:
bashsource phases/00-setup-and-tooling/10-terminal-and-shell/code/shell_aliases.sh别人可以复制你想要的.
bash# GPU status at a glance
alias gpu='nvidia-smi --query-gpu=index,name,utilization.gpu,memory.used,memory.total,temperature.gpu --format=csv,noheader'
# Kill all Python training processes
alias killtraining='pkill -f "python.*train"'
# Quick virtual environment activate
alias ae='source .venv/bin/activate'
# Watch training loss
alias watchloss='tail -f logs/*.log | grep --line-buffered "loss"'看到code/shell_aliases.sh对于整个集.
步骤8:常见的AI终端模式
这些问题在实践中反复出现:
bash# Run training, log everything, notify when done
python train.py 2>&1 | tee train.log; echo "DONE" | mail -s "Training complete" you@email.com
# Compare two experiment logs side by side
diff <(grep "accuracy" exp1.log) <(grep "accuracy" exp2.log)
# Find the largest model files (clean up disk space)
find . -name "*.pt" -o -name "*.safetensors" | xargs du -h | sort -rh | head -20
# Download a model from Hugging Face
wget https://huggingface.co/model/resolve/main/model.safetensors
# Untar a dataset
tar xzf dataset.tar.gz -C ./data/
# Count lines in all Python files (see how big your project is)
find . -name "*.py" | xargs wc -l | tail -1
# Check disk space (training data fills disks fast)
df -h
du -sh ./data/*
# Environment variable check before training
env | grep -i cuda
env | grep -i torch用它
在课程中,每种工具都会发挥作用:
| Tool | When you use it |
|---|---|
| tmux | Every training run (Phases 3+) |
tail -f + grep | Monitoring training logs |
nohup / & | Quick background tasks |
htop / nvtop | Debugging slow training, OOM errors |
SSH + rsync | Working on cloud GPUs |
| Piping + redirects | Processing experiment results |
| Aliases | Saving time on repetitive commands |
运动
- 安装tmux,创建一个三个面板的会议,然后运行
htop在一个,watch -n1 date在另一个,和Python脚本在第三个. - 添加来自的号
code/shell_aliases.sh让你的子配置和重新充电source ~/.zshrc(或~/.bashrc) - 创建一个假的训练日志
for i in $(seq 1 100); do echo "epoch $i loss: $(echo "scale=4; 1/$i" | bc)"; sleep 0.1; done > fake_train.log然后使用grep现在tail其他awk只有输出值. - 设置一个SSH配置输入,为您访问 (或使用) 的服务器
localhost实践语法).
关键词
| Term | What people say | What it actually means |
|---|---|---|
| Shell | "The terminal" | The program that interprets your commands (bash, zsh, fish) |
| tmux | "Terminal multiplexer" | A program that lets you run multiple terminal sessions inside one window, and detach/reattach |
| Pipe | "The bar thing" | The | operator that sends one command's output as input to another |
| PID | "Process ID" | A unique number assigned to every running process, used to monitor or kill it |
| nohup | "No hangup" | Runs a command immune to the hangup signal, so closing the terminal won't kill it |
| SSH | "Connecting to the server" | Secure Shell, an encrypted protocol for running commands on a remote machine |
This free lesson is part of the AI Engineering from Scratch curriculum. Read the full explanation, run the lesson code, and verify the result in the interactive reader or from the repository source.
Browse the complete course catalog or open this lesson on GitHub.