Guanxiong Luo
  • About
  • Projects
  • Blog
  • Publications
  • generative models
  • •

  • diffusion models
  • •

  • linear algebra
  • •

  • computational imaging
  • •

  • inverse problems
  • •

  • tools
  • •

  • coding
  • •

  • The Details of Flash Attention

    From the online-softmax derivation to eight CUDA kernels — a 121.8× speedup, and the two optimizations that made things slower.

    19 min read   ·   June 22, 2026

    2026

  • Tensor Parallel + Sequence Parallel — A Deep Dive

    17 min read   ·   April 30, 2026

    2026

  • ZeRO-1 Distributed Optimizer - A Deep Dive

    19 min read   ·   April 30, 2026

    2026

  • Optimizing softmax on GPU

    7 min read   ·   December 29, 2025

    2025   ·   cuda   self-attention   numerical computation   coding  

  • The details of flash attention - algorithm

    5 min read   ·   December 17, 2025

    2025   ·   generative models   self-attention   coding  

  • Newer
  • 1
  • 2
  • 3
  • Older
© Copyright 2026 Guanxiong Luo. Powered by Jekyll with al-folio theme. Hosted by GitHub Pages. Last updated: September 22, 2026.