AWS brings Kimi K3 to Bedrock with a 1M-token context window
AWS has added Moonshot AI's Kimi K3 to Amazon Bedrock, offering managed API access to vision input, a one-million-token context window and prompt caching.
AWS has added Moonshot AI’s Kimi K3 to Amazon Bedrock. The September 18 launch gives teams a managed API route to the model, which accepts image and text input, returns text, and supports a one-million-token context window and prompt caching.
Teams access Kimi K3 through US Geo or Global cross-Region inference profiles; AWS does not list a single-Region option. The Bedrock model card says developers can use the Responses, Chat Completions, Converse or Invoke APIs, with AWS recommending Chat Completions. Responses and Chat Completions offer Standard, Priority and Flex service tiers. Converse and Invoke are limited to Standard on-demand inference.
Implicit prompt caching is enabled by default. Teams can also set explicit cache controls through the Responses and Chat Completions APIs. AWS requires at least 1,024 tokens for each explicit cache checkpoint and says cached content remains available for at least 30 minutes. According to the company, reusing long prompt prefixes can improve cache-hit rates while reducing latency and cost.
At the Standard tier, AWS prices Global cross-Region inference at $3 per million input tokens, $15 per million output tokens, $0.30 per million cache-read tokens and $3.75 per million tokens written to a 30-minute cache. The corresponding US cross-Region prices are $3.30, $16.50, $0.33 and $4.125.
Moonshot AI describes Kimi K3 as a 2.8-trillion-parameter mixture-of-experts model that activates 104 billion parameters. The developer reports roughly 2.5 times better overall scaling efficiency than Kimi K2, though the opened sources provide no independent Bedrock benchmark for that claim. DataPhoenix previously covered Kimi K3’s behavior in a UK AISI evaluation sandbox.
AWS says inference data for open-weight models stays within the AWS data boundary, is not shared with the model provider or used to train the underlying model, and is protected by zero data retention and zero operator access.
AWS also flags limits for teams testing the managed route. Reusing prior-turn reasoning content in a multi-turn Converse request can trigger an InternalServerException, while Converse rejects attached document inputs including PDF and HTML files. Bedrock Knowledge Bases and intelligent prompt routing are not supported for Kimi K3.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
