# AWS makes GPU-aware inference routing available on SageMaker HyperPod

_Published Thursday, September 24, 2026 at 8:07 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

AWS says its SageMaker HyperPod Inference Gateway is available for routing model requests within a cluster in supported regions. It installs as an EKS managed add-on on existing HyperPod infrastructure and works with OpenAI-compatible model servers without application code changes.

The gateway reads the requested model and selects a GPU server using signals including queue depth and cache use. AWS says this approach cuts first-token latency by up to 82% and reduces p99 first-token latency by 97-98% in mixed-hardware and burst-traffic scenarios. Routing across clusters and regions is still planned.

## Sources

- [Recent Announcements](https://aws.amazon.com/about-aws/whats-new/2026/09/sagemaker-hyperpod-inference-gateway/)

---
Canonical: https://techandbusiness.org/newswire/IRdaqDy-53Oa9rBmnt3LlX
Published: 2026-09-25T00:07:52.744Z
Story chronology: 2026-09-24T08:00:00.000Z
Retrieved: 2026-09-25T01:53:03.384Z
Publisher: Tech & Business (techandbusiness.org)
