Skip to main content

Share story

AI

PyTorch reports faster attention kernel for Meta ads workloads on Blackwell

An attention kernel built with Triton Low-level Extensions outperforms the May 2026 version of FlashAttention-4 on variable-length workloads used by Meta's Generative Ads Model, PyTorch reports. Benchmarks on NVIDIA B200 GPUs show approximately 13% better forward-pass performance and approximately 50% better backward-pass performance. The kernel processes packed sequences without adding padding and explicitly coordinates memory transfers and computation to reduce stalls. Its code is available on GitHub. The implementation uses about 3.2K lines, compared with approximately 10K lines for the FlashAttention-4 kernels. The performance comparisons use bfloat16 arithmetic and workload shapes relevant to Meta's ads model, limiting broader conclusions.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from PyTorch and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Capital
AI Capital

SoftBank and Nvidia complete $30B OpenAI funding pledges

SoftBank and Nvidia have each made a final $10 billion investment to complete their respective $30 billion pledges to OpenAI's last funding round, The Information reported. The payments complete the two investors' commitments to t...

AI Products
AI Products

SpaceXAI releases Grok iOS app for Intune-managed organizations

SpaceXAI has released Grok for Intune on the App Store, giving organizations a dedicated iOS version of its AI assistant that supports Microsoft's application management system. The company says the app honors workplace restrictio...