Added India geographic cross-Region inference for Kimi K3: Kimi K3 is now available through India geographic cross-Region inference (the in.moonshotai.kimi-k3 inference profile) in the Asia Pacific (Mumbai) ( ap-south-1 ) and Asia Pacific (Hyderabad) ( ap-

Published
October 7, 2026
https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html

Kimi K3

Model Details

Kimi K3 is Moonshot AI's most capable open-weight model, combining native vision with a 1-million-token context window for long-running coding and knowledge workflows that sustain context across large repositories, documents, and images.

APIs supported

bedrock-runtime, bedrock-mantle

Capabilities and Features

Explicit Prompt Caching supported

  • Min tokens per cache checkpoint: 1,024
  • Cache retention (TTL): At least 30 minutes

Pricing

All prices are per 1 million tokens. Pricing shown is for the Standard tier.

  • Global CRIS: $3.00 input, $15.00 output, $0.30 cache read, $3.75 cache write (30 min)
  • US CRIS: $3.30 input, $16.50 output, $0.33 cache read, $4.125 cache write (30 min)
  • IN CRIS: $3.30 input, $16.50 output, $0.33 cache read, $4.125 cache write (30 min)

Regional Availability

Kimi K3 is available through US Geo cross-Region inference, India geographic cross-Region inference, and Global cross-Region inference.

Quotas and Limits

Your AWS account has default quotas to maintain the performance of the service and to ensure appropriate usage of Amazon Bedrock. See your default quotas in Service Quotas and request limit increases as necessary.

Usage Considerations and Limitations

  • Prefer the OpenAI-compatible APIs over Converse
  • Video inputs are not supported
  • Place images before text for combined inputs
  • Image detail parameter is honored only on the Chat Completions API

What to do

  • Use the bedrock-runtime endpoint for new applications
  • Use explicit prompt caching to improve cache hit rate and reduce latency and cost
  • Choose the appropriate service tier based on your workload requirements
  • Select the appropriate inference option based on your data residency requirements
  • Test both orderings for image and text inputs to achieve higher-quality answers

Source: AWS release notes




If you need further guidance on AWS, our experts are available at AWS@westloop.io. You may also reach us by submitting the Contact Us form.

Follow our blog

Get the latest insights and advice on AWS services from our experts.

By clicking Sign Up you're confirming that you agree with our Terms and Conditions.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.