Serverless Storage on Amazon EMR Serverless now supports terabyte-scale shuffle

Amazon EMR Serverless Updates
Amazon EMR Serverless now supports up to 1TB shuffle operations, raising the previous 200 GB per-job limit. This enhancement enables enterprise customers to run production-scale Apache Spark workloads that require processing large volumes of shuffle data during complex operations such as joins, aggregations, and sorting.
Enterprise data teams can now confidently migrate production workloads that routinely process terabyte-scale datasets without worrying about storage constraints. This enhancement is particularly valuable for workloads involving large table joins across multi-terabyte datasets, and complex aggregations on high-cardinality data that require extensive data shuffling.
The addition of spill support ensures that jobs can seamlessly handle memory-intensive operations by offloading data to disk when necessary, improving job reliability and success rates for demanding analytical workloads.
What to do
- Migrate production workloads that process large volumes of shuffle data.
- Utilize spill support for memory-intensive operations.
- Check the Amazon EMR documentation for supported Regions and limits.
Source: AWS release notes
If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.



