# AWS launches GPU-aware SageMaker HyperPod Inference Gateway

_Published Friday, September 18, 2026 at 9:18 AM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![AWS launches GPU-aware SageMaker HyperPod Inference Gateway — Primary](https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21822-featured-image-2.png)

AWS announced general availability of the SageMaker HyperPod Inference Gateway, a Kubernetes-native routing addon that places inference requests using real-time GPU signals such as KV cache utilization, queue depth, LoRA adapter residency and prefix cache hit rate.

AWS says the addon installs on existing HyperPod/EKS clusters with no application or model-server changes and exposes an OpenAI-compatible endpoint. In AWS benchmarks across models from 8B to 235B parameters, the company reported first-token latency reductions of up to 82 percent versus a Kubernetes round-robin baseline, with gains concentrated in mixed-GPU, bursty and shared-prefix workloads.

A second-tier Global Inference Router for cross-cluster routing is described as coming soon.

## Sources

- [AWS Machine Learning Blog](https://aws.amazon.com/blogs/machine-learning/introducing-amazon-sagemaker-hyperpod-inference-gateway/)

---
Canonical: https://techandbusiness.org/newswire/AuwrsoRwixk18u2O16S-Ow
Published: 2026-09-18T13:18:37.198Z
Story chronology: 2026-09-18T13:08:34.000Z
Retrieved: 2026-09-18T16:07:44.556Z
Publisher: Tech & Business (techandbusiness.org)
