pytorch.
3 writings found
Latest Archives
Why Your Attention Kernel is 3.7x Slower Than You Think
Understanding PyTorch attention backends through profiler traces reveals why naive implementations beat optimized ones, and what it means for LLM performance.
Why FlashAttention Breaks the Profiler (And Why That's Good)
FlashAttention shows low GPU occupancy yet outperforms all other attention backends. Here's what the profiler isn't telling you about modern kernel design.
Profiling Attention in PyTorch: From Naive to FlashAttention
A deep dive into profiling attention mechanisms in PyTorch, comparing naive, SDPA math, efficient, and FlashAttention backends using GPU traces.
Prev
Page 1 of 1 Next