AWS announces aws-bench, an open-source benchmark for AI agents on AWS

AWS Research Preview: aws-bench
AWS has introduced a research preview of aws-bench, an open-source benchmark for evaluating AI agents on AWS tasks. This tool provides a standardized way to measure performance and diagnose failures for agents operating on AWS infrastructure.
Key Features
- Public suite of test cases derived from real AWS usage.
- Each test pairs a natural-language query with a cloud resource state and a ground-truth answer.
- Easy-to-use CLI tool for testing environments, execution, scoring, and resource state reset.
What to do
- Visit the aws-bench GitHub page for more information.
- Follow the setup instructions on the README to get started.
Source: AWS release notes
If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.



