Skip to main content

Share story

AI Infrastructure

AWS makes GPU-aware inference routing available on SageMaker HyperPod

AWS says its SageMaker HyperPod Inference Gateway is available for routing model requests within a cluster in supported regions. It installs as an EKS managed add-on on existing HyperPod infrastructure and works with OpenAI-compatible model servers without application code changes. The gateway reads the requested model and selects a GPU server using signals including queue depth and cache use. AWS says this approach cuts first-token latency by up to 82% and reduces p99 first-token latency by 97-98% in mixed-hardware and burst-traffic scenarios. Routing across clusters and regions is still planned.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from Recent Announcements and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Capital Infrastructure
Capital Infrastructure

Zero Infinity Partners announces $156 million infrastructure technology fund

Zero Infinity Partners announced $156 million in commitments for its second fund, which will invest in early-stage companies applying technology to physical infrastructure. The firm said the fund is 50% larger than its predecessor...

Infrastructure Capital
Infrastructure Capital

Clastix raises €2.9 million for Kubernetes infrastructure software

Italian infrastructure software company Clastix has raised €2.9 million in its first external funding round, led by CDP Venture Capital with participation from Mistral and Vertis SGR. Clastix plans to use the seed capital for prod...

Robotics Products
Robotics Products

Cognex agrees to acquire RealSense for about $500 million

Cognex has agreed to buy RealSense for about $500 million in cash, adding cameras and software that help robots judge the distance and position of objects. Cognex plans to fund the purchase with cash and investments on its balance...