Skip to main content

Share story

AI Infrastructure

Nvidia adds multi-GPU TensorRT inference to Dynamo-Triton

Nvidia adds multi-GPU TensorRT inference to Dynamo-Triton Image: Primary
Nvidia released multi-device TensorRT inference support in Dynamo-Triton 26.07, allowing one served model to execute across multiple GPUs behind a single gRPC endpoint. TensorRT 11.0 supplies the distributed capability, while the server manages execution contexts, CUDA streams and NCCL communications for each GPU rank. In Nvidia's eight-GPU Cosmos 3 Nano test, end-to-end video-generation latency fell from 156.595 seconds to 34.183 seconds. The benchmark excluded model loading and MP4 encoding and did not measure concurrent throughput, cost per video or total ownership cost.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from NVIDIA Technical Blog and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Capital Infrastructure
Capital Infrastructure

Rainmaker raises $100 million to scale measured cloud seeding

Cloud-seeding startup Rainmaker raised $100 million in a Series B backed by NOA VC, Upfront Ventures, DCVC, Lowercarbon Capital and Dream Ventures. The company plans to expand its weather-research team, scale operations in the Ame...

Infrastructure Products
Infrastructure Products

Amazon commits $20 million to Colorado River conservation effort

Amazon launched a Colorado River Basin Collaborative and committed $20 million to it, with a goal of raising $100 million over two years for conservation projects. Participating companies may contract independently with projects v...

Security
Security

Greenberg Traurig clients sue over data breach

A proposed class of Greenberg Traurig clients has sued the law firm in New York federal court, alleging that it failed to protect their personal information before a preventable data breach. The plaintiffs also allege that the fir...