-
The Details of Flash Attention
From the online-softmax derivation to eight CUDA kernels — a 121.8× speedup, and the two optimizations that made things slower.
-
Tensor Parallel + Sequence Parallel — A Deep Dive
-
ZeRO-1 Distributed Optimizer - A Deep Dive
-
Optimizing softmax on GPU
-
The details of flash attention - algorithm