Amazon SageMaker HyperPod now supports deep health checks for Slurm clusters with continuous provisioning

Amazon SageMaker HyperPod Updates
Amazon SageMaker HyperPod now supports deep health checks for Slurm-orchestrated clusters created with continuous provisioning, enabling proactive verification of GPU accelerator health on running instances at any time. This feature allows for quick training starts and asynchronous scaling of instance groups without all-or-nothing failures, paired with comprehensive hardware validation as instances come online.
With deep health checks, you can target entire instance groups or specific instances to run comprehensive hardware stress tests and connectivity tests before committing compute resources to a job. This ensures that instances are validated before scheduling jobs on them, without interrupting workloads on healthy nodes. Progress and results are visible at both the instance group and instance level through the SageMaker console and APIs, providing complete visibility into GPU health, network connectivity, and multi-node communication performance.
Instances undergoing checks are automatically isolated from workload scheduling and returned to service upon passing. When paired with HyperPod's automatic node recovery capability, instances that fail are automatically rebooted or replaced, ensuring cluster health.
What to do
- Enable deep health checks for your Slurm-orchestrated clusters to proactively verify GPU health.
- Monitor progress and results through the SageMaker console and APIs.
- Utilize continuous provisioning for quick training starts and asynchronous scaling.
- Ensure cluster health with automatic node recovery for failed instances.
Source: AWS release notes
If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.



