Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference

Amazon SageMaker HyperPod Inference Gateway
Introducing the Kubernetes-native, GPU-aware routing system for SageMaker HyperPod. It reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios.
Core Components
- Envoy Endpoint: Terminates HTTPS traffic and exposes a single private endpoint per cluster.
- Body-Based Router: Routes requests to the correct GPU pool based on model name.
- Endpoint Picker: Continuously scores model server pods across 6 inference-level signals to select the optimal pod for each request.
Compatibility
Works with any OpenAI-compatible model server, including vLLM and SGLang, with no application code changes required.
Availability
Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Upcoming features include cross-cluster and cross-region routing, global rate limiting, and cost-tier-aware traffic shaping.
What to do
- Read the launch blog
- Explore the documentation
Source: AWS release notes
If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.



