# PyTorch reports faster attention kernel for Meta ads workloads on Blackwell

_Published Friday, October 2, 2026 at 12:11 AM EDT · AI · Latest · Tier 2 — Notable_

An attention kernel built with Triton Low-level Extensions outperforms the May 2026 version of FlashAttention-4 on variable-length workloads used by Meta's Generative Ads Model, PyTorch reports. Benchmarks on NVIDIA B200 GPUs show approximately 13% better forward-pass performance and approximately 50% better backward-pass performance.

The kernel processes packed sequences without adding padding and explicitly coordinates memory transfers and computation to reduce stalls. Its code is available on GitHub. The implementation uses about 3.2K lines, compared with approximately 10K lines for the FlashAttention-4 kernels. The performance comparisons use bfloat16 arithmetic and workload shapes relevant to Meta's ads model, limiting broader conclusions.

## Sources

- [PyTorch](https://pytorch.org/blog/optimizing-jagged-flash-attention-with-tlx-the-road-toward-sota-fa4-on-blackwell/)

---
Canonical: https://techandbusiness.org/newswire/g70sTzkxDlzCeYNDIFnh6L
Published: 2026-10-02T04:11:01.956Z
Story chronology: 2026-10-01T22:26:57.000Z
Retrieved: 2026-10-02T08:25:20.960Z
Publisher: Tech & Business (techandbusiness.org)
