# PyTorch post details open-sourced FP8 FlashAttention work for Blackwell

_Published Wednesday, September 16, 2026 at 4:07 PM EDT · AI, Science · Latest · Tier 2 — Notable_

A PyTorch engineering post said code for an MXFP8 extension to FlashAttention-4 has been open sourced for Blackwell GPUs. The work adds block-scaled FP8 forward and backward attention and, on internal LLM shapes, reported gains of up to 1.6 times in forward performance and 1.52 times in backward performance versus BF16. The post said the accompanying zero-gather jagged module is used internally at Meta for GEM training. Reported performance is based on internal shapes and implementation-specific measurements.

## Sources

- [PyTorch](https://pytorch.org/blog/low-precision-flash-attention-4-end-to-end-block-scaled-attention-for-blackwell/)

---
Canonical: https://techandbusiness.org/newswire/gyQWOkQsqB4tsL1PMLdem9
Published: 2026-09-16T20:07:32.459Z
Story chronology: 2026-09-16T18:55:21.000Z
Retrieved: 2026-09-16T22:16:27.699Z
Publisher: Tech & Business (techandbusiness.org)
