Updated the tokens-per-day quota on the bedrock-runtime endpoint to a cross-model, per-account quota: Replaced the per-model Model invocation max tokens per day quota with a Cross-Model Max Tokens Per Day quota that is scoped per account, per Region across

Published
September 22, 2026
https://docs.aws.amazon.com/bedrock/latest/userguide/quotas-runtime.html

Quotas for the bedrock-runtime endpoint

The bedrock-runtime.region.amazonaws.com endpoint is the primary inference endpoint for Amazon Bedrock. Inference traffic to this endpoint is governed by per-model token-based quotas. You can view these quotas in the Service Quotas console by selecting Amazon Bedrock as the service.

Quota types

Inference on the bedrock-runtime endpoint is governed by the following per-model quotas:

  • Cross-Region InvokeModel tokens per minute for ${model} - Per model, per Region. The maximum number of tokens per minute (input + output, combined) that your account can use for the model when invoked through a cross-Region inference profile.
  • On-demand InvokeModel tokens per minute for ${model} - Per model, per Region. The maximum number of tokens per minute (input + output, combined) that your account can use for the model when invoked on-demand in a single Region.
  • Cross-Model Max Tokens Per Day - Per account, per Region. Maximum tokens per day across all supported Amazon Bedrock models for this account.
  • InvokeModel requests per minute for ${model} - Per model, per Region. The maximum number of inference requests per minute that your account can submit for the model.

Requesting a quota increase

Before requesting a quota increase, verify that the model is not in a Legacy or Deprecated lifecycle status. If a quota is marked as Yes, you can adjust it by following the steps at Requesting a Quota Increase in the Service Quotas User Guide.

What to do

  • Check the model's lifecycle status on the Model lifecycle page.
  • Consider migrating to the successor model if the current model is scheduled for retirement.
  • Request an increase for the Cross-Region InvokeModel tokens per minute for ${model} quota to also increase the other related quotas.

Source: AWS release notes




If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.

Follow our blog

Get the latest insights and advice on AWS services from our experts.

By clicking Sign Up you're confirming that you agree with our Terms and Conditions.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.