New research: FlashAttention-4 FlashAttention-4 achieves up to 1.3x speedup over cuDNN 9.13 and 2.7x over Triton on B200 GPUs with BF16. FlashAttention-4 co-designs algorithms and kernel pipelines for Blackwell GPUs, where tensor core throughput doubles but memory bandwidth and
FlashAttention-4 Achieves 2.7x Speedup on Blackwell GPUs
By
–
