Skip to main content

Share story

AI

Helion integration reports more than 10% inference gains for some vLLM workloads

A PyTorch project reports that its Helion integration into vLLM improved end-to-end inference performance on NVIDIA Hopper GPUs, with more than 10% higher throughput for some workloads. The backend automatically tunes the matrix calculations used by quantized language models rather than requiring separate manually specialized implementations. The implementation uses Helion for small workloads executed through CUDA Graphs, which replay GPU operations with less CPU overhead, and falls back to existing kernels for larger workloads. Tests used an NVIDIA H100 80GB HBM3 GPU. Fine-grained tuning can still take hours, and compilation increases cold-start latency; caching compiled artifacts largely removes that cost on warm starts.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from PyTorch and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Robotics Capital
Robotics Capital

Fortem raises $50 million with Lockheed Martin participating

Counter-drone company Fortem has raised $50 million in Series B funding, with Lockheed Martin participating, Bloomberg reported. Fortem is seeking to expand manufacturing and bookings rapidly. CEO Jon Gruen said its relationship w...

AI Infrastructure
AI Infrastructure

Broadcom syndicate seeks $60 billion for AI chip financing

Broadcom's Wall Street syndicate is starting to gather $60 billion in fresh AI chip financing, Bloomberg reported, citing sources. The financing effort would help Anthropic and other companies access chips and other infrastructure...

Infrastructure Products
Infrastructure Products

Rocket Lab signs 20-launch agreement with Synspective

Rocket Lab signed a 20-mission Electron launch agreement with Japanese radar satellite operator Synspective, Ars Technica reported, citing Aviation Week & Space Technology. Rocket Lab calls it its largest Electron launch service a...

AI Capital
AI Capital

Halluminate raises $30M for AI financial training environments

Halluminate raised a $30M Series A led by Oak HC/FT, bringing its total funding to $38.5M, Fortune reported. The nine-person San Francisco startup builds environments for training AI to perform complex financial work, giving the f...

AI Infrastructure
AI Infrastructure

OpenAI deploys Jalapeño chips with AMD Turin host processors

OpenAI is deploying its Jalapeño custom AI chips internally alongside AMD EPYC Turin host processors, each with 1.5TB of memory, Tom's Hardware reported. OpenAI hardware chief Richard Ho said the company chose Turin to reduce depl...