Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference

Published
September 24, 2026
https://aws.amazon.com/about-aws/whats-new/2026/09/sagemaker-hyperpod-inference-gateway/

Amazon SageMaker HyperPod Inference Gateway

Introducing the Kubernetes-native, GPU-aware routing system for SageMaker HyperPod. It reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios.

Core Components

  • Envoy Endpoint: Terminates HTTPS traffic and exposes a single private endpoint per cluster.
  • Body-Based Router: Routes requests to the correct GPU pool based on model name.
  • Endpoint Picker: Continuously scores model server pods across 6 inference-level signals to select the optimal pod for each request.

Compatibility

Works with any OpenAI-compatible model server, including vLLM and SGLang, with no application code changes required.

Availability

Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Upcoming features include cross-cluster and cross-region routing, global rate limiting, and cost-tier-aware traffic shaping.

What to do

Source: AWS release notes




If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.

Follow our blog

Get the latest insights and advice on AWS services from our experts.

By clicking Sign Up you're confirming that you agree with our Terms and Conditions.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.