Meta 团队利用 Triton Low-level Extensions (TLX) 重新设计了 Jagged Flash Attention (JFA) 内核,旨在解决 Meta Generative Ads Model (GEM) 在 NVIDIA Blackwell (B200) 芯片上的性能瓶颈。该方案通过显式硬件控制(如异步任务、显式显存管理)替代传统手写 CUDA,将代码量减少 3 倍,同时在前向传播和反向传播上分别提升约 13% 和 50% 的性能,实现了高性能与高开发效率的平衡。
代码行数
3.2K lines
相比 FA4 的 ~10K 行代码,减少约 3 倍
前向传播性能提升
~13%
在锯齿状序列形状上的优化效果
反向传播性能提升
~50%
在锯齿状序列形状上的优化效果
目标硬件
NVIDIA Blackwell (B200)
基于 Triton Low-level Extensions (TLX)
Meta 发布 TLX 优化版 Jagged Flash Attention:Blackwell 架构下 GEM 模型性能新突破
更新背景:Blackwell 时代的性能挑战
随着 NVIDIA Blackwell (B200) 架构的商用,大模型训练对算力的需求达到了新的高度。Meta 的 Generative Ads Model (GEM) 在处理非结构化用户序列时,面临严峻的性能挑战:注意力机制 (Attention) 是其最慢的内核,且处理“锯齿状” (Jagged/Ragged) 序列时,若采用传统填充 (Padding) 策略,高达 50% 的计算资源将被浪费。
此前,为了在 Blackwell 上达到峰值性能,团队不得不依赖手写 CuteDSL 或 CUDA 代码。这种方式虽然性能优异,但开发周期长、扩展性差,难以快速迭代新的注意力变体(如滑动窗口、块稀疏等)。Meta 此次发布的最新工作,正是为了解决这一“性能与开发效率”的矛盾。
“TLX closes this gap on both fronts. On development efficiency, the TLX attention kernel is about 3.2K lines of concise Triton-level code — roughly 3× less than the ~10K-line CuteDSL kernels of the state-of-the-art FlashAttention-4 (FA4). On performance, it outperforms FA4 (May 2026 version) on the jagged shapes that matter for GEM — by ~13% on the forward pass and ~50% on the backward pass.”
— Meta PyTorch 团队