Llama 4 Scout and Maverick models now available
Meta's Llama 4 Scout (17Bx16MoE) and Maverick (17Bx128E) models for image understanding and text generation are now available via Groq API, supporting a 128K token context window, up to 5 image inputs, function calling/tool use, and JSON mode.
Performance (per Artificial Analysis): Llama 4 Scout (meta-llama/llama-4-scout-17b-16e-instruct) — 607 TPS; Llama 4 Maverick (meta-llama/llama-4-maverick-17b-128e-instruct) — 297 TPS.
Fetched August 11, 2026


