Amazon SageMaker HyperPod now supports model caching for faster inference autoscaling and reduced cold starts

Published
September 11, 2026
https://aws.amazon.com/about-aws/whats-new/2026/09/sgm-hyperpod-model-caching-inf/

Amazon SageMaker HyperPod Model Caching

Amazon SageMaker HyperPod now supports model caching, an inference optimization that pre-loads model weights and container images onto cluster nodes to reduce cold start times from minutes to seconds.

Model caching includes:

  • Weights Cache: Stores model weights on local NVMe for faster access.
  • Image Cache: Pre-pulls container images to skip ECR downloads.

Benchmarks show:

  • 60% faster scale-out for models from 57 GB to 145 GB.
  • 97% reduction in image-pull time.

What to do

  • Enable model caching by adding a modelCacheConfig section to your InferenceEndpointConfig or JumpStartModel resource.
  • No manual setup or cleanup required; the HyperPod Inference Operator handles the lifecycle.

Model caching is now generally available in all regions where SageMaker HyperPod is available. To get started, see the SageMaker HyperPod documentation.




If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.

Follow our blog

Get the latest insights and advice on AWS services from our experts.

By clicking Sign Up you're confirming that you agree with our Terms and Conditions.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.