- What changed
- OpenAI GPT-5.6 models are now generally available on Amazon Bedrock with support for explicit prompt caching. This feature enables precise control over cached prompt segments to help reduce inference costs for existing workloads.
- Why you should care
- Bringing new flagship models and granular prompt caching to managed cloud platforms directly impacts the operational costs and architecture decisions of enterprise AI deployments.
- Your move
- Watch. Review current Bedrock workload configurations to implement explicit prompt caching for repetitive or large context prompts to optimize inference costs.
- What to watch next
- Watch for further verified reporting or independent evidence.
- Event
- release
- Event date
- Jul 30, 2026
- Relevant to
- AI Engineers, Cloud Architects, Enterprise Developers