# Nvidia reports 10x MoE training gain with JAX Transformer Engine path

_Monday, September 14, 2026 at 12:39 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![Nvidia reports 10x MoE training gain with JAX Transformer Engine path — Primary](https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube.webp)

Nvidia said its optimized JAX and Transformer Engine path achieved a 10x end-to-end throughput gain training DeepSeek-V3 671B, compared with an unoptimized baseline. The company attributes the result to grouped GEMM, optimized expert-parallel dispatch and combine, MXFP8 quantization, host activation offloading and multistream collectives. Nvidia says the stack sustained 97 percent efficiency at 1,024 GPUs and ships in its NGC MaxText container. The configuration is specific to DeepSeek-V3 and requires model-specific tuning.

## Sources

- [NVIDIA Technical Blog](https://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/)

---
Canonical: https://techandbusiness.org/newswire/UKUCT2JfB5fxTSlZYk15Cq
Retrieved: 2026-09-15T01:30:43.260Z
Publisher: Tech & Business (techandbusiness.org)
