releases.sh

Changelog

$npx @buildinternet/releases get groq-changelog
Mon
Wed
Fri
SepOctNovDecJanFebMarAprMayJunJulAugSep
Less
More
Releases4Avg1/mo

Python SDK v1.2.0 preserves hardcoded query params, improves multipart file-copy performance, and adds binary request streaming with a custom JSON encoder; Python 3.9 is deprecated. TypeScript SDK v1.1.2 restores streaming in defaultParseResponse and fixes an abort-signal memory leak and minimatch CVEs.

Read more →

MiniMax M2.5 general-purpose model and Qwen3-VL 32B Instruct vision-language model are now available on GroqCloud for Enterprise customers. Python SDK v1.2.0 preserves hardcoded query params when merging with user params, ensures file data is sent as a single parameter, and improves multipart request file-copy performance. TypeScript SDK v1.1.2 restored streaming support in defaultParseResponse, fixed an abort-signal memory leak, and patched a minimatch CVE.

Read more →

Text-to-speech has migrated platform-wide from PlayAI to Orpheus models from Canopy Labs, which offer enhanced expressiveness with vocal direction controls, faster inference, and improved audio quality. Users still on playai-tts or playai-tts-arabic must migrate before the shutdown date listed on the deprecations page.

Read more →

OpenAI's GPT-OSS-Safeguard 20B, an open-weight reasoning model for safety classification, is now available on Groq. The model offers a 131K context window, up to 65K output tokens, ~1000 TPS throughput, prompt caching for 50% cost savings, and supports tool use, code execution, and multiple response formats for Trust & Safety content moderation and policy-based classification tasks.

Read more →

Automatic prompt caching is now available for openai/gpt-oss-120b, reducing cached input token costs to $0.075/M (from $0.15/M), lowering latency, and increasing effective rate limits since cached tokens do not count toward limits. Python SDK updated to v0.33.0 and TypeScript SDK to v0.34.0 with improved prompt caching support and annotation/citation support for chat completion messages and streamed deltas.

Read more →

Remote MCP server integration is available in beta on GroqCloud, connecting AI models to thousands of external tools via Anthropic's open MCP standard. The implementation is fully compatible with OpenAI Responses API and OpenAI remote MCP spec, enabling zero-code-change migration from OpenAI to Groq. Supported across eight models including Llama, Qwen, Kimi, and proprietary OSS variants, with launch partner tutorials for BrowserBase, Browser Use, Exa, Firecrawl, HuggingFace, Parallel, Stripe, and Tavily.

Read more →

Moonshot AI's Kimi K2-0905 model is now available on GroqCloud with day zero support, featuring a 256K context window (largest on GroqCloud), prompt caching for up to 50% savings, and improved agentic coding reliability in multi-turn interactions. Pricing is $1.00/M input and $3.00/M output tokens.

Read more →

Compound and Compound Mini move from beta to general availability as production-ready agentic AI systems integrating web search, code execution, and browser automation in a single API call. Both models deliver ~25% higher accuracy and ~50% fewer mistakes than OpenAI's Web Search Preview and Perplexity Sonar. Python SDK v0.31.1 and TypeScript SDK v0.32.0 add support for new Compound tool types including Wolfram Alpha and browser automation.

Read more →

Improved chat completion message type definitions in Python SDK v0.31.1 and TypeScript SDK v0.32.0 for better OpenAI compatibility, fixing errors with certain message formats. Added support for new Groq Compound tool types (Wolfram Alpha, Browser Automation, Visit Website).

Read more →

OpenAI's open-source Mixture-of-Experts models GPT-OSS 20B and GPT-OSS 120B are now available, featuring 131K context window, 32K max output tokens, reasoning, built-in browser search and code execution, and structured outputs support. Also launched Responses API (Beta), fully compatible with OpenAI's Responses API with support for text/image inputs, stateful conversations, and function calling. Python and TypeScript SDKs added reasoning_effort control and new tool types (browser_search, code_interpreter).

Read more →

Groq now supports structured outputs with JSON schema for moonshotai/kimi-k2-instruct, meta-llama/llama-4-maverick-17b-128e-instruct, and meta-llama/llama-4-scout-17b-16e-instruct. Model responses are guaranteed to conform strictly to a provided JSON schema, eliminating the need for complex parsing logic.

Read more →

Moonshot AI's Kimi K2 Instruct, a trillion-parameter MoE model with 32 billion activated parameters, is now available, offering a 131K token context window and 16K max output tokens. It surpasses GPT-4.1 on agentic and coding benchmarks and excels at tool use and autonomous problem-solving.

Read more →
Last Checked
11d ago
Latest
Apr 18, 2026
Tracking since Apr 14, 2025