Announcing region expansion of G7e instances on SageMaker AI inference

Amazon EC2 G7e Instances Now Available in New Regions
Amazon EC2 G7e instances are now available in Asia Pacific (Seoul), Europe (London), and Asia Pacific (Tokyo) on Amazon SageMaker AI inference. These instances feature up to 8 NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs with 96 GB of memory per GPU, 5th Generation Intel Xeon processors, and up to 1,600 Gbps of Elastic Fabric Adapter networking bandwidth, delivering up to 2.3x inference performance compared to previous-generation G6e instances.
With this region expansion, you can deploy inference endpoints closer to your end users in Asia and Europe, reducing latency for generative AI workloads. G7e instances provide up to 768 GB of total GPU memory on a single instance, enabling you to serve medium-to-large language models of up to 70B parameters with FP8 precision without multi-node configurations. These instances are well suited for LLM inference, image and video generation, spatial computing, and scientific computing workloads that require high GPU memory capacity and bandwidth.
What to do
- Deploy inference endpoints on G7e instances in the new regions to reduce latency for your AI workloads.
- Serve medium-to-large language models with FP8 precision without multi-node configurations.
- Utilize the high GPU memory capacity and bandwidth for LLM inference, image and video generation, spatial computing, and scientific computing.
Source: AWS release notes
If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.



