08 · FlashAttention:省的是 IO 不是 FLOPs
结论:不物化 S、P,靠 online softmax 逐块累加,HBM 132 MiB → 66 MiB(实际近 4 MiB),从 memory-bound 变 compute-bound;FLOPs 略多于标准实现;decode 每步读全部 KV,分页只改地址不改字节。
%%{init: {"flowchart": {"wrappingWidth": 190}}}%%
flowchart TB
subgraph std["标准:三个 kernel,S 与 P 往返 HBM"]
direction LR
q1["Q, K 各 1 MiB"] --> g1["GEMM #1"] -- "写 / 读 S 32 MiB" --> sm1["softmax"] -- "写 / 读 P 32 MiB" --> g2["GEMM #2"] --> o1["O 1 MiB"]
end
subgraph fa["FlashAttention:一个 kernel,S / P 只在片上"]
direction LR
q2["Q_i, K_j, V_j tile"] --> g3["S_ij 寄存器"] --> sm2["online softmax<br/>更新 m, l"] --> g4["O_i += P̃_ij V_j"] -- "遍历完写 1 次" --> o2["O 1 MiB"]
end
std ~~~ fa
classDef hbm fill:#fde2e2,stroke:#c0392b
classDef chip fill:#dff5e1,stroke:#1e8449
class q1,o1,q2,o2 hbm
class g3,sm2,g4 chip
COMMENTS
评论存放在 GitHub Discussions, 用 GitHub 账号登录即可发表,支持 Markdown。 想针对正文某句话说?选中那段文字,点浮出的「评论」即可划线评论;觉得哪里写错了,发表时勾上「同时提交 Issue」。 有人回复你时 GitHub 会按你的通知设置发邮件,不用守在这里。