08 · 性能优化:CPU 与 GPU 是两条异步时间线
结论:五类瓶颈——launch / Python / sync-bound(CPU 侧)、memory / compute-bound(GPU 侧)——优化是把瓶颈从一类推到另一类。
%%{init: {"flowchart": {"wrappingWidth": 160}}}%%
flowchart LR
tl["时间线<br/>torch.profiler / nsys"] --> q1{"GPU 泳道形态?"}
q1 -->|"稀疏,大段空闲"| q2{"CPU 在忙什么?"}
q1 -->|"与 CPU 交替空洞"| sync["Sync-bound<br/>批量 .item()、pinned + non_blocking"]
q1 -->|"密集,首尾相接"| q3{"CUDA 时间被谁占?"}
q2 -->|"cudaLaunchKernel 占大头"| launch["Launch-bound<br/>增大 batch、融合、CUDA Graphs"]
q2 -->|"Python / 框架逻辑"| py["Python-bound<br/>向量化、compile"]
q3 -->|"逐元素与归约"| mem["Memory-bound<br/>融合 / SDPA、bf16"]
q3 -->|"mm / bmm / conv"| comp["Compute-bound<br/>Tensor Core、减少计算量"]
classDef cpu fill:#fde9d9,stroke:#c0392b
classDef gpu fill:#e8f5e9,stroke:#1e8449
classDef mid fill:#fff4d6,stroke:#b9770e
class launch,py cpu
class mem,comp gpu
class sync mid
COMMENTS
评论存放在 GitHub Discussions, 用 GitHub 账号登录即可发表,支持 Markdown。 想针对正文某句话说?选中那段文字,点浮出的「评论」即可划线评论;觉得哪里写错了,发表时勾上「同时提交 Issue」。 有人回复你时 GitHub 会按你的通知设置发邮件,不用守在这里。