# Nvidia adds multi-GPU TensorRT inference to Dynamo-Triton

_Published Monday, September 21, 2026 at 7:08 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![Nvidia adds multi-GPU TensorRT inference to Dynamo-Triton — Primary](https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1.webp)

Nvidia released multi-device TensorRT inference support in Dynamo-Triton 26.07, allowing one served model to execute across multiple GPUs behind a single gRPC endpoint. TensorRT 11.0 supplies the distributed capability, while the server manages execution contexts, CUDA streams and NCCL communications for each GPU rank.

In Nvidia's eight-GPU Cosmos 3 Nano test, end-to-end video-generation latency fell from 156.595 seconds to 34.183 seconds. The benchmark excluded model loading and MP4 encoding and did not measure concurrent throughput, cost per video or total ownership cost.

## Sources

- [NVIDIA Technical Blog](https://developer.nvidia.com/blog/simplifying-model-serving-across-multiple-gpus-with-nvidia-tensorrt-multi-device-integration-in-nvidia-dynamo-triton/)

---
Canonical: https://techandbusiness.org/newswire/6c9vB1-DxMkL67LzhNqGi0
Published: 2026-09-21T23:08:20.502Z
Story chronology: 2026-09-21T21:51:06.000Z
Retrieved: 2026-09-22T00:57:40.354Z
Publisher: Tech & Business (techandbusiness.org)
