Prompt caching live for GPT-OSS 120B—50% input token savings
Added
Automatic prompt caching now live for openai/gpt-oss-120b: 50% cost savings on cached input tokens ($0.075/M vs $0.15/M), lower latency, and higher effective rate limits since cached tokens don't count toward limits. Zero setup required.
Changed Python SDK updated to v0.33.0, TypeScript SDK to v0.34.0 — improved prompt caching support and added annotation/citation support to chat completion messages and streamed deltas.
Fetched August 11, 2026


