Skip to main content
AI Infrastructure

AWS launches GPU-aware SageMaker HyperPod Inference Gateway

AWS launches GPU-aware SageMaker HyperPod Inference Gateway Image: Primary
AWS announced general availability of the SageMaker HyperPod Inference Gateway, a Kubernetes-native routing addon that places inference requests using real-time GPU signals such as KV cache utilization, queue depth, LoRA adapter residency and prefix cache hit rate. AWS says the addon installs on existing HyperPod/EKS clusters with no application or model-server changes and exposes an OpenAI-compatible endpoint. In AWS benchmarks across models from 8B to 235B parameters, the company reported first-token latency reductions of up to 82 percent versus a Kubernetes round-robin baseline, with gains concentrated in mixed-GPU, bursty and shared-prefix workloads. A second-tier Global Inference Router for cross-cluster routing is described as coming soon.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from AWS Machine Learning Blog and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Infrastructure
AI Infrastructure

Huawei outlines 2027 Ascend AI-chip launches

Huawei plans to launch its Ascend 960DT AI chip in the first quarter of 2027 and the Ascend 960PR in the third quarter, Reuters reported. Rotating chairman David Wang said UnifiedBus technology would be key to Huawei's next-genera...

AI Infrastructure
AI Infrastructure

GMI Cloud reportedly seeks $300 million chip loan for Thailand

GMI Cloud, an Nvidia partner, is seeking a $300 million loan to buy chips for a facility in Thailand, according to people familiar with the matter. The prospective financing would add to similar Asian deals supporting AI-compute i...

AI Products
AI Products

AWS adds Gemma 4 open-weight models to Bedrock in EU sovereign cloud

Amazon Web Services said Gemma 4, an Apache 2.0 open-weight model family, is generally available on Amazon Bedrock's next-generation inference engine in the AWS European Sovereign Cloud. AWS called it the first open-weight family...

AI
AI

OpenAI launches Astra for Law with GPT-6 for select firms

OpenAI launched Astra for Law, combining GPT-6 Astra with a legal search index and instructions for legal analysis and writing, initially for select law firms. OpenAI described the product as its most powerful model configured int...

Infrastructure
Infrastructure

Amazon Leo adds six Ariane 6 launch flights

Amazon said Amazon Leo has added six Arianespace Ariane 6 flights to its launch plan, capacity it said is sufficient for more than 800 satellites. The company said it has deployed nearly 400 satellites across 14 missions under a m...

Infrastructure Products
Infrastructure Products

AWS opens general availability of T8i burstable EC2 instances

Amazon Web Services said burstable EC2 T8i instances are generally available, using custom sixth-generation Intel Xeon Scalable processors (Granite Rapids) that the company said are available only on AWS. AWS said the instances a...