MoE models that compute router logits now return them when output_router_logits=True, changing existing outputs, and the "paged|" prefix for SDPA and flash attention is deprecated in favor of setting the regular attention implementation. Adds EmbeddingGemma2, a multimodal embedding model handling text, images, audio, and video in a shared 768-dimensional space with Matryoshka truncation.
Transformers
npx @buildinternet/releases get transformers-releasesNew model additions include Nemotron 3 Diarization (streaming speaker diarization, up to eight speakers), NemotronH Omni (joint text/image/video/audio reasoning), HyperCLOVAX Vision V2, and the GTE BERT-style encoder family. The release also carries flagged breaking changes across ROCm gpt-oss attention routing, DETR image processing, indexer layer_type remapping, vLLM video token counting, DINO modernization, and a Kernels version bump, plus a long tail of bug fixes spanning MoE expert parallelism, Qwen2.5-VL temporal RoPE, and DeepSeek-OCR-2 patching.
Adds support for seven new architectures: HY4-Preview (780B MoE), VibeVoice, NeoMME, Fun-ASR-Nano, KimiLinear, and Canary-1B-v2. Vision rotary embeddings (2D/3D) are standardized into a unified RoPE module, requiring migration for custom vision models.
Adds support for GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. Restores backward compatibility for the tensor-parallel API and pins an ESMFold2 kernel commit for security reasons.
This release adds four new models: Qwen4-Exp, GraniteSpeech5, Step3.7-Flash, and CohereCompass. It also replaces the legacy tensor-parallel implementation with a DTensor-native backend, a breaking change for existing TP users, and fixes several cache and generation bugs.
Fixed DFlash candidate token device mismatch with device_map='auto', aligned logit distributions for sampling-based candidate generators, fixed MTP config when mlp_layer_types is absent, and added a fallback from Lanczos to bicubic filtering on CUDA so images process on accelerator. Also fixed gemma4 video device placement.
Adds support for Meta's Muse Glimmer 30B multimodal model plus GraniteSWA, GraniteMoeSWA, A.X-K1/K2, and Cosmos3 Edge. Kernels for linear attention models (Mamba, GDN, Conv-only) are now opt-in rather than mandatory, cache cropping only accepts negative offsets, and T5 family gains SDPA support with a possible default attention change.
Fixed assisted decoding for models using EncoderDecoderCache (e.g., OlmoHybrid) and a SDPA prefill issue with position_bias during StaticCache. Also bumps FP8 kernels and fixes DeepGEMM on multiple devices.
Added Inkling, a 975B-parameter multimodal model, and TIPSv2. GPTNeoX now remaps embed_out to lm_head and GPTBigCode has attention backend changes for vLLM compatibility. Multi-Token Prediction decoding support, SDPA prefill with FlashAttention for StaticCache (up to 260% faster), and numerous bug fixes across MoE, cache, and generation.
Fixes custom model compatibility with the latest vllm release by being more defensive with remap_legacy_layer_types and handling cases where custom code doesn't know about the new linear layer type names. Also fixed a key type assertion in _LazyAutoMapping.register.
This release adds support for seven new model architectures: Kimi 2.5–2.7 (multimodal agentic coding), MiMo-V2-Flash (256K context MoE model), Nemotron 3.5 ASR and Nemotron ASR Streaming (multilingual speech recognition with configurable latency-accuracy tradeoffs), Qwen3 ASR with forced aligner, ZAYA1 (MoE language model), VideoPrism (video understanding encoder), and RADIO (vision foundation model family).
Fixed mistral tokenizer resolution when mistral-common is installed and updated the lower bound for PEFT. This is similar to v5.10.3 minus fixes already in the main release.
This patch release fixes several regressions introduced by previous changes, including issues with {image/video/audio}_token_ids in ProcessorMixin, InternVL models, and offsets in processing. It also addresses a regression in the Mistral common backend and updates the peft lower bound.
This release introduces the MiniMax-M3-VL vision-language model, the PP-OCRv6 OCR system, and the Parakeet-RNNT model for speech processing. Several bug fixes and improvements were also made, including changes to CI, stop string matching, and model documentation.
New models DiffusionGemma and DeepSeek-V3.2 have been added, featuring optimizations for inference speed and efficient long-context handling. The Kernels API was extended for module fusion and parameter transformation, with added support for fp8/fp4 Triton kernels. Model parallel beam search bugs in Qwen2-VL model families were fixed.
Fixed a conversion bug for CLIP models that affected downstream models like SAM3.
Added Gemma4 12B Unified, an encoder-free multimodal model that projects raw vision and audio inputs directly into language model space; Sapiens2, a vision transformer family for human-centric tasks; DeepSeek-OCR-2 for document understanding; and Mellum, a code-focused mixture-of-experts model. Fixed numerous model parallelism bugs across tensor and expert parallelism, beam search under parallel settings, and loss over-counting; also fixed encoder-decoder cache initialization regression and BitsAndBytes quantization tensor-dropping bug.
Added support for Cohere2Moe (a Mixture-of-Experts model with sliding window and full attention), HRM-Text (hierarchical reasoning model with two transformer stacks), and Parakeet tdt speech model. SAM3, EdgeTAM, and SAM3-Lite-Text now expect full text embeddings instead of pooler outputs, requiring input updates. Fixed generation issues including inputs_embeds handling for Gemma4, an AttributeError in RAG's generate() caused by missing config fields, memory leaks from lru decorators in vision models, and improved audio/vision encoder compilability.
Fixed Deepseek V4 integration issues including CSA mask collapse and WeightConverter regex incorrectly matching shared_experts as experts. Also added fatal_error to ContinuousBatchingManager for serving operations.
Release v5.8.0
New Model additions
DeepSeek-V4
DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model…

