---
name: Groq
slug: groq
domain: groq.com
description: High-speed LLM inference API and hardware for low-latency AI applications.
category: ai
sources: 1
total_releases: 23
releases_last_30d: 0
avg_releases_per_week: 0
last_updated: 2026-08-11
tracking_since: 2025-04-14
canonical: https://releases.sh/groq
tags:
  - ai
  - inference
  - llm
accounts:
  github: "groq"
  x: "GroqInc"
---

<Source name="Changelog" slug="groq-changelog" type="scrape" releases="23" latest-date="2026-04-18T00:00:00.000Z" primary="true" url="https://releases.sh/groq/groq-changelog" />

## Recent Releases

_Summaries below — fetch the release's `canonical` URL for full content, or `url` for the original source._

<Release source="groq-changelog" date="April 18, 2026" published="2026-04-18T00:00:00.000Z" url="https://console.groq.com/docs/changelog#minimax-m25-and-qwen3vl-32b-instruct-enterprise" canonical="https://releases.sh/release/rel_2kz-r5h-_YW-5VRgLA9UH" truncated="true">
### MiniMax M2.5 and Qwen3-VL 32B Instruct (Enterprise)

MiniMax M2.5 general-purpose model and Qwen3-VL 32B Instruct vision-language model are now available on GroqCloud for Enterprise customers. Python SDK v1.2.0 preserves hardcoded query params when merging with user params, ensures file data is sent as a single parameter, and improves multipart request file-copy performance. TypeScript SDK v1.1.2 restored streaming support in defaultParseResponse, fixed an abort-signal memory leak, and patched a minimatch CVE.
</Release>

<Release source="groq-changelog" date="January 30, 2026" published="2026-01-30T00:00:00.000Z" url="https://console.groq.com/docs/changelog#platformwide-migration-from-playai-to-orpheus-tts" canonical="https://releases.sh/release/rel_slohPX8XO2tKT7GyKNH-b" truncated="true">
### Platform-wide Migration from PlayAI to Orpheus TTS

Text-to-speech has migrated platform-wide from PlayAI to Orpheus models from Canopy Labs, which offer enhanced expressiveness with vocal direction controls, faster inference, and improved audio quality. Users still on playai-tts or playai-tts-arabic must migrate before the shutdown date listed on the deprecations page.
</Release>

<Release source="groq-changelog" date="December 1, 2025" published="2025-12-01T00:00:00.000Z" url="https://console.groq.com/docs/changelog#mcp-connectors-beta" canonical="https://releases.sh/release/rel_LUkKXRCb5HmCWr3VHCH-g" truncated="true">
### MCP Connectors (Beta)

Groq now offers pre-built MCP Connectors for Google Workspace applications—Gmail, Google Calendar, and Google Drive—with zero configuration and OAuth 2.0 authentication. Available tools include email search and reading, calendar event viewing, and file search and access.
</Release>

<Release source="groq-changelog" date="October 29, 2025" published="2025-10-29T00:00:00.000Z" url="https://console.groq.com/docs/changelog#openai-gptosssafeguard-20b" canonical="https://releases.sh/release/rel_QFRiyluxZeQXt11y7s8KR" truncated="true">
### OpenAI GPT-OSS-Safeguard 20B

OpenAI's GPT-OSS-Safeguard 20B, an open-weight reasoning model for safety classification, is now available on Groq. The model offers a 131K context window, up to 65K output tokens, ~1000 TPS throughput, prompt caching for 50% cost savings, and supports tool use, code execution, and multiple response formats for Trust & Safety content moderation and policy-based classification tasks.
</Release>

<Release source="groq-changelog" date="October 21, 2025" published="2025-10-21T00:00:00.000Z" url="https://console.groq.com/docs/changelog#prompt-caching-enabled-for-gptoss-120b" canonical="https://releases.sh/release/rel_o8ua9YNzzWNiq1d950yRf" truncated="true">
### Prompt Caching Enabled for GPT-OSS 120B

Automatic prompt caching is now available for openai/gpt-oss-120b, reducing cached input token costs to $0.075/M (from $0.15/M), lowering latency, and increasing effective rate limits since cached tokens do not count toward limits. Python SDK updated to v0.33.0 and TypeScript SDK to v0.34.0 with improved prompt caching support and annotation/citation support for chat completion messages and streamed deltas.
</Release>

<Release source="groq-changelog" date="September 25, 2025" published="2025-09-25T00:00:00.000Z" url="https://console.groq.com/docs/changelog#prompt-caching-enabled-for-gptoss-20b" canonical="https://releases.sh/release/rel_pfXaVh3l5QB34NVgMR04x" truncated="true">
### Prompt Caching Enabled for GPT-OSS 20B

Automatic prompt caching is now available for openai/gpt-oss-20b, reducing cached input token costs by 50% (from $0.075/M to $0.037/M) while lowering latency through automatic prefix matching. No configuration required.
</Release>

<Release source="groq-changelog" date="September 23, 2025" published="2025-09-23T00:00:00.000Z" url="https://console.groq.com/docs/changelog#remote-model-context-protocol-mcp" canonical="https://releases.sh/release/rel_LWJIB2p9S8gv29MlGsN7h" truncated="true">
### Remote Model Context Protocol (MCP)

Remote MCP server integration is available in beta on GroqCloud, connecting AI models to thousands of external tools via Anthropic's open MCP standard. The implementation is fully compatible with OpenAI Responses API and OpenAI remote MCP spec, enabling zero-code-change migration from OpenAI to Groq. Supported across eight models including Llama, Qwen, Kimi, and proprietary OSS variants, with launch partner tutorials for BrowserBase, Browser Use, Exa, Firecrawl, HuggingFace, Parallel, Stripe, and Tavily.
</Release>

<Release source="groq-changelog" date="September 5, 2025" published="2025-09-05T00:00:00.000Z" url="https://console.groq.com/docs/changelog#moonshot-ai-kimi-k2-instruct-0905" canonical="https://releases.sh/release/rel_-sMTwSq6cvmCysLaIQDsK" truncated="true">
### Moonshot AI Kimi K2 Instruct 0905

Moonshot AI's Kimi K2-0905 model is now available on GroqCloud with day zero support, featuring a 256K context window (largest on GroqCloud), prompt caching for up to 50% savings, and improved agentic coding reliability in multi-turn interactions. Pricing is $1.00/M input and $3.00/M output tokens.
</Release>

<Release source="groq-changelog" date="September 4, 2025" published="2025-09-04T00:00:00.000Z" url="https://console.groq.com/docs/changelog#groq-compound-and-compound-mini" canonical="https://releases.sh/release/rel_yIF36XMpvW6tfbSQ3IHZj" truncated="true">
### Groq Compound and Compound Mini

Compound and Compound Mini move from beta to general availability as production-ready agentic AI systems integrating web search, code execution, and browser automation in a single API call. Both models deliver ~25% higher accuracy and ~50% fewer mistakes than OpenAI's Web Search Preview and Perplexity Sonar. Python SDK v0.31.1 and TypeScript SDK v0.32.0 add support for new Compound tool types including Wolfram Alpha and browser automation.
</Release>

<Release source="groq-changelog" date="August 20, 2025" published="2025-08-20T00:00:00.000Z" url="https://console.groq.com/docs/changelog#prompt-caching" canonical="https://releases.sh/release/rel_7qAV5shMeTdXXbSAVyN_V" truncated="true">
### Prompt Caching

Prompt caching automatically reuses computation from recent requests sharing a common prefix, reducing latency and cutting token costs by 50% for cached portions. Rolling out first to Kimi K2 with more models coming; no code changes or additional fees required.
</Release>

## Fetching more

Append `.md` (markdown), `.json` (raw data), or `.atom` (feed) to any URL on this page.

- Per-source history: `https://releases.sh/groq/{source-slug}`
- Atom feed: `https://releases.sh/groq.atom`
- Individual release: `https://releases.sh/release/{release-id}`
