releases.sh

Prompt caching live for GPT-OSS 120B—50% input token savings

October 21, 2025ChangelogView original ↗
2 features1 enhancementThis release2 featuresNew capabilities1 enhancementImprovements to existing featuresAI-tallied from the release notes

Added Automatic prompt caching now live for openai/gpt-oss-120b: 50% cost savings on cached input tokens ($0.075/M vs $0.15/M), lower latency, and higher effective rate limits since cached tokens don't count toward limits. Zero setup required.

Changed Python SDK updated to v0.33.0, TypeScript SDK to v0.34.0 — improved prompt caching support and added annotation/citation support to chat completion messages and streamed deltas.

Fetched August 11, 2026

Prompt caching live for GPT-OSS 120B—50% input token… — releases.sh