grok-voice-transcribe-1.0 is deprecated and reaches end of life on October 2, 2026. Until then all requests to that slug route to grok-voice-transcribe-2.0 at the same price, with higher accuracy.
Release Notes
npx @buildinternet/releases get release-notesgrok-voice-transcribe-2.0 is now available for speech to text alongside grok-voice-transcribe-1.0, which remains the default.
A new safety_identifier request field accepts an opaque end-user identifier assigned by your application, letting SpaceXAI attribute a policy violation to one of your end users rather than your API key. It is supported on Chat Completions, the Responses API, deferred chat completions, the Batch API, and the gRPC GetCompletionsRequest, and the legacy user field is still accepted.
Grok 4.7 is available on the xAI API as grok-4.7 with a 500k context window, low through xhigh reasoning effort (default high), and pricing of $2/$0.50/$6 per 1M tokens (input/cached input/output) below 200k prompt tokens and $4/$1/$12 above. On the Responses API it always returns reasoning.encrypted_content even when include does not list it.
On November 2, 2026, grok-imagine-image-quality is retired, with requests to the slug served by grok-imagine-image-2.0 with quality set to low at a lower per-image price and no change to the request or response shape. grok-imagine-image (1.0) is unaffected.
Grok 4.6 is available on the xAI API with a 500k context window, text and image inputs with text-only output, and no text output limit. Pricing is $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200k prompt tokens and $4 / $1 / $12 above, with reasoning effort supporting low, medium, high (default), and xhigh.
Grok Bot is now available, offering durable AI teammates that work on a persistent cloud computer with messaging, approvals, connectors, and routines.
The grok-imagine-image-2.0 quality parameter now accepts auto, with the default moving from medium to auto, which uses low for generation and medium for editing. Image editing accepts up to 5 source images per request instead of 3, and both generation and editing add 21:9 and 5:2 aspect ratios.
Grok 4.5 is now available in the API console for EU users.
Speech to Text now accepts a vad_threshold parameter as a streaming query param and batch multipart field to tune the voice-activity gate that skips non-speech audio. Lower values transcribe quieter or noisier speech for narrowband telephony, and 0 disables the gate.
grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video with optional preset voices, with native 1080p for text-to-video and image-to-video. Text-to-video on this model runs as text-to-image then image-to-video.
grok-voice-think-fast-2.0 is now available with speech to speech, and grok-voice-latest will route to this model starting August 5, 2026.
Grok 4.5, aimed at coding, agentic tasks, and knowledge work, is now available on the xAI API at $2 per 1M input tokens and $6 per 1M output tokens. Reasoning effort is configurable across low, medium, and high, defaulting to high.
Text inference endpoints (Chat Completions and Responses) now accept service_tier: "priority" to request higher scheduling priority per request. The response's service_tier field reports the tier actually applied, and priority rates are billed only when priority is used.
Any file in Files API storage can now be turned into a permanent, unauthenticated URL, revocable at any time or set to auto-expire between 1 hour and 30 days. Imagine endpoints accept image_file_id, video_file_id, and reference_image_file_ids in place of URL inputs, and storage_options persists generated assets to Files storage, optionally publishing a shareable link in one round trip.
xAI's fast coding model trained specifically for agentic coding, currently in early access. The model slug is grok-build-0.1.
Grok Build is now available in beta. Use the interactive TUI, run headlessly in scripts, or build apps and orchestrators with the Agent Client Protocol.
Install with a single command:
curl -fsSL https://x.ai/cli/install.sh | bashThe Responses API can now be driven over a single, long-lived WebSocket connection for lower end-to-end latency on tool-heavy agent workloads.
The Context Compaction API is now available, letting you shrink long conversations into a shorter context and reuse it in follow-up requests for lower cost, faster time-to-first-token, and sharper responses on long agent loops.
You can now clone a voice from a short audio clip and use it across the Text-to-Speech and Voice Agent APIs. Create and manage your voice catalog from the xAI console.


