# AWS adds model caching to SageMaker HyperPod, cutting inference cold starts

_Friday, September 11, 2026 at 2:25 PM EDT · AI, Products · Latest · Tier 2 — Notable_

Amazon Web Services said SageMaker HyperPod now supports model caching, an inference optimization that pre-loads model weights and container images onto cluster nodes so pods start in seconds rather than minutes.

The feature is generally available in all regions where HyperPod is available. AWS describes two independent capabilities: a weights cache that stores model weights on local NVMe instead of pulling from S3 or FSx over the network, and an image cache that pre-pulls container images to skip ECR downloads.

Pods landing on a node without a warm cache fall back to the original source automatically. AWS reports benchmarks across models from 57 GB to 145 GB showing roughly 60% faster scale-out, with the image cache cutting over two minutes of image-pull time, a 97% reduction.

Customers enable it through the HyperPod Inference Operator.

## Sources

- [Recent Announcements](https://aws.amazon.com/about-aws/whats-new/2026/09/sgm-hyperpod-model-caching-inf/)

---
Canonical: https://techandbusiness.org/newswire/2-AIj7Q_KlhMSN4dF7dU8o
Retrieved: 2026-09-12T02:47:46.103Z
Publisher: Tech & Business (techandbusiness.org)
