# AWS makes SageMaker HyperPod model caching generally available

_Thursday, September 10, 2026 at 5:37 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![AWS makes SageMaker HyperPod model caching generally available — Primary](https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/08/ML-21820-featured-image.png)

Amazon Web Services said model caching for SageMaker Inference on HyperPod is now generally available in all regions where HyperPod is offered.

The feature pre-loads model weights and inference-server container images onto cluster nodes before pods are scheduled, so pods read from local NVMe storage at roughly 7 GB/s instead of pulling weights and images over the network. AWS said benchmarks across 57-145 GB models showed about 60 percent faster scale-out with weights caching, while image caching removed over two minutes of cold image-pull time.

AWS said the first cache population still requires a remote download, and weights caches are per-node, so NVMe consumption scales with node count.

## Sources

- [Artificial Intelligence](https://aws.amazon.com/blogs/machine-learning/reduce-inference-cold-starts-on-amazon-sagemaker-hyperpod-with-model-caching/)

---
Canonical: https://techandbusiness.org/newswire/F4s29mz0mFCCub1cT3tHeq
Retrieved: 2026-09-11T03:31:46.373Z
Publisher: Tech & Business (techandbusiness.org)
