Skip to main content

Share story

AI

vLLM Adds Support for NVIDIA Nemotron 3 Super for Multi-Agent AI

Run Highly Efficient and Accurate Multi-Agent AI with NVIDIA Nemotron 3 Super Using vLLM Image: Primary
vLLM has added support for the NVIDIA Nemotron 3 Super model. The model is part of the Nemotron 3 family of open models and is optimized for complex multi agent applications. Agentic AI systems use multiple models to plan, reason and execute multi step tasks. Nemotron 3 Super is a hybrid Mixture of Experts model with 120 billion total parameters but only 12 billion active at inference. It features a 1 million token context window to manage excessive token generation from history and tool outputs. The architecture also delivers up to four times higher throughput to address costs associated with reasoning intensive agents. The model is fully open with available weights, datasets and recipes. It supports multi token prediction and a thinking budget for accuracy with fewer reasoning tokens. Supported GPUs include the B200, H100, DGX Spark and RTX 6000. Model weights in BF16, FP8 and NVFP4 formats can be downloaded from Hugging Face. vLLM serves the model via an OpenAI compatible API, with configurations available for different hardware setups.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from vLLM Blog and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI
AI

Contrastive-LM releases open model for scoring AI agent actions

Contrastive-LM has released CLM-8B, an open model that scores possible actions for an AI agent instead of generating text. It compares a description of the current situation with candidate actions, then returns probabilities that ...

AI Infrastructure
AI Infrastructure

Oracle seeks payment delay protection for New Mexico data center

Oracle has sent a force majeure notice to the developer of Project Jupiter, a planned New Mexico AI data center, Bloomberg reported, citing people familiar with the matter. Oracle wants the right to delay payments if the campus mi...

AI Infrastructure
AI Infrastructure

AWS makes GPU-aware inference routing available on SageMaker HyperPod

AWS says its SageMaker HyperPod Inference Gateway is available for routing model requests within a cluster in supported regions. It installs as an EKS managed add-on on existing HyperPod infrastructure and works with OpenAI-compat...

Products AI
Products AI

Court filings show OpenAI cut forecasts for Apple ChatGPT integration

OpenAI cut its forecasts for users gained through Apple's ChatGPT integration a month after the feature launched in December 2024, according to court documents described by the Financial Times. By summer 2025, OpenAI had concluded...