{"id":"src_IUEmstpqhDfKSwUnkELwV","slug":"transformers-releases","name":"Transformers","type":"github","url":"https://github.com/huggingface/transformers","orgId":"org_GDdYeYynEgCEBNBwy-m6s","productId":"prod_766UloABFcuw8w2LvR2iy","productSlug":"transformers","org":{"id":"org_GDdYeYynEgCEBNBwy-m6s","slug":"hugging-face","name":"Hugging Face"},"isPrimary":false,"isHidden":false,"discovery":"curated","metadata":"{\"evaluatedMethod\":\"github\",\"evaluatedAt\":\"2026-04-07T17:19:13.059Z\",\"changelogDetectedAt\":\"2026-04-07T17:27:13.693Z\",\"wellKnownSweptAt\":\"2026-10-01T06:01:26.362Z\",\"sourceActor\":{\"nextAlarmAt\":\"2026-10-08T19:08:44.028Z\",\"lastAlarmAt\":\"2026-10-08T15:08:45.363Z\",\"managed\":true}}","notice":null,"kind":"sdk","stars":167059,"starsFetchedAt":"2026-10-08T15:11:23.289Z","releaseCount":109,"releasesLast30Days":3,"avgReleasesPerWeek":0.8,"latestVersion":"v5.19.0","latestDate":"2026-10-06T16:39:23.000Z","changelogUrl":null,"hasChangelogFile":false,"lastFetchedAt":"2026-10-08T15:11:23.289Z","lastPolledAt":"2026-10-08T15:11:10.625Z","changeDetectedAt":null,"trackingSince":"2024-04-23T22:01:20.000Z","releases":[{"id":"rel_T6OWDvvEInddmCJEd2FHi","version":"v5.19.0","type":"feature","title":"Release v5.19.0","summary":"MoE models that compute router logits now return them when output_router_logits=True, changing existing outputs, and the \"paged|\" prefix for SDPA and flash attention is deprecated in favor of setting the regular attention implementation. Adds EmbeddingGemma2, a multimodal embedding model handling text, images, audio, and video in a shared 768-dimensional space with Matryoshka truncation.","titleGenerated":"Transformers v5.19.0 adds EmbeddingGemma2 and breaking MoE router logits changes","titleShort":"MoE router logits now returned; EmbeddingGemma2 added","breaking":"major","importance":4,"content":"# Release v5.19.0\r\n\r\n\r\n## New Model additions\r\n\r\n### EmbeddingGemma2\r\n\r\n<img width=\"2716\" height=\"2308\" alt=\"image\" src=\"https://github.com/user-attachments/assets/84734e74-163d-4d12-b166-ffcf6749d563\" />\r\n\r\nEmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture. It encodes text, images, audio, and video, individually or combined in one input, into a shared 768-dimensional vector space for cross-modal retrieval, semantic similarity, clustering, and classification. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. It also offers configurable visual and video token budgets, and unused vision or audio towers can be disabled at load time to save memory.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/embedding_gemma2)\r\n* Smthn smthn (#49364) by @vasqu in [#49364](https://github.com/huggingface/transformers/pull/49364)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nAll MoE models whose routers compute logits now return them when `output_router_logits=True`, following the Qwen3-MoE pattern (a `router_logits` recorder on the base model, `MoeModelOutputWithPast` from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.\r\n* 🚨 Return router logits from every MoE model that computes them (#48920) by @qgallouedec\r\n\r\n`Owlv2ForObjectDetection.embed_image_query` now selects the query box with the highest objectness score, as in the original OWLv2 notebook, instead of the OWL-ViT heuristic, so image-guided query embeddings and detections may differ from earlier releases.\r\n* 🚨 Select the OWLv2 image query by objectness (#49200) by @qgallouedec\r\n\r\nThe `\"paged|\"` prefix for SDPA and flash attention implementations is deprecated, so users should set the regular attention implementation (e.g. `sdpa` or `flash_attention_2`) for continuous batching instead of `paged|sdpa` or `paged|flash_attention_2`.\r\n* 🚨 Attention 🚨 Deprecate \"paged|\" prefix for SDPA and flash  (#49112) by @remi-or\r\n\r\nThe regular flash and SDPA attention functions (`flash_attention.py`, `sdpa_attention.py`) now support continuous batching directly, and `\"paged|...\"` implementations for these are redirected to them, while eager still requires the `\"paged|eager\"` prefix.\r\n* 🚨 Attention 🚨 Make regular attention support CB (#49101) by @remi-or\r\n\r\nIn continuous batching, the cache update for the index-based and block-table paths is now fused into a single call, which slightly changes the cache update function's behavior and affects any custom code that calls the separate update paths.\r\n* 🚨 [CB] 🚨 Fuse update for index and block table path (#49088) by @remi-or\r\n\r\nContinuous batching internals changed in preparation for removing `\"paged\"`: `max\r\n* 🚨 [CB] 🚨 Little fixes before removing \"paged\" (#49069) by @remi-or\r\n\r\n\r\n\r\n## Parallelization\r\n\r\nExpert parallelism gains a token-dispatch implementation, selected via the new `ep_dispatch_experts` plan rule and now the default for Qwen3 MoE and Mellum, which removes the requirement that EP size equal TP size. The `Trainer` was also adapted to work with expert parallelism, and the docs now note that PEFT adapters support tensor parallelism. A CI-related fix for pipeline-parallel chart2table inference was also included.\r\n\r\n\r\n* [`distributed`]: adapt Trainer to work Expert Parallel (#48873) by @3outeille in [#48873]\r\n* [`distributed`] Add expert-parallel token dispatch, default for Qwen3 MoE (#48865) by @3outeille in [#48865]\r\n* Fix pp chart2table inference (#49225) by @zucchini-nlp in [#49225]\r\n* [docs] TP for PEFT adapters (#49060) by @stevhliu in [#49060]\r\n\r\n\r\n## Cache\r\n\r\nThis release fixes quantized cache handling: `generate` no longer mutates the user's `cache_config`, and `QuantizedLayer.reorder_cache` is repaired. It also adds per-layer cache configuration, so `DynamicCache` and `StaticCache` initialize each layer from its own config (sliding window, attention chunk size, conv states, and attention head counts) to better support heterogeneous models. Separately, the deprecation cycle on mask and cache\r\n\r\n\r\n* Fix quantized cache (#48700) by @jiqing-feng in [#48700]\r\n* [CI] check_bad_commit: use EFS cache to avoid Xet FUSE OOM (exit 137) (#49273) by @ydshieh in [#49273]\r\n* Support per-layer cache configuration (#48178) by @eladsegal in [#48178]\r\n* Remove deprecation cycle on mask and cache (#49231) by @Cyrilvallez in [#49231]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Revert \"[CI] Temporarily disable AMD scheduled CI trigger\" (#49271) (#49362) by @ydshieh in [#49362]\r\n* ALM testx fIxing... (#49355) by @zucchini-nlp in [#49355]\r\n* Fix DataCollatorForLanguageModeling ignoring seed=0 (#49349) by @akanyaani in [#49349]\r\n* Fix explicit kernelization modes for evaluation models (#49339) by @DimensionSTP in [#49339]\r\n* fix(cohere_compass): add video to input_modalities (#49320) (#49323) by @destopianpirate in [#49323]\r\n* GPT-OSS: read the clamped SwiGLU's alpha and limit from the config (#49346) by @IlyasMoutawwakil in [#49346]\r\n* [`distributed`] Default MoE `ep_plans` to token dispatch (#49160) by @3outeille in [#49160]\r\n* Move CI to Python 3.11, with the version set in `.python-version` (#49311) by @tarekziade in [#49311]\r\n* Bump doc-builder main docs workflow pin (#49350) by @paulinebm in [#49350]\r\n* [`distributed`] decoupled tp_plan and ep_plan (#48859) by @3outeille in [#48859]\r\n* [`distributed`] Add dense and expert device mesh (#48857) by @3outeille in [#48857]\r\n* Bump doc-builder workflow pin (#49342) by @paulinebm in [#49342]\r\n* Allow registering new processing backends (#48984) by @zucchini-nlp in [#48984]\r\n* Fix WatermarkDetector repeated-ngram counting (#49315) by @LE0-Lin in [#49315]\r\n* Fix minicpm 4.6V video batching and CI (#49259) by @zucchini-nlp in [#49259]\r\n* remove colwise_gather_output for lm_head for DeepseekV4  (#49314) by @3outeille in [#49314]\r\n* Fix the `labels` docstring of CLIPSegForImageSegmentation (#49001) by @qgallouedec in [#49001]\r\n* [CI] Temporarily disable AMD scheduled CI trigger (#49271) by @ydshieh in [#49271]\r\n* Remove the unused pandas entry from setup.py deps (#49329) by @tarekziade in [#49329]\r\n* Fix run_glue.py label ids becoming strings with pandas>=3 (#49326) by @tarekziade in [#49326]\r\n* [Fix] Remove old test file with two ancient tests (#49280) by @remi-or in [#49280]\r\n* Fix chunked prefill with multi-axis position ids (Qwen3-VL) (#49319) by @dacorvo in [#49319]\r\n* Why override when you can fix in `Mixin` (#49105) by @zucchini-nlp in [#49105]\r\n* Do RoPE with mul instead of matmul (#49265) by @Rocketknight1 in [#49265]\r\n* Deprecated pipelines should redirect; removed pipelines should raise (#49308) by @LysandreJik in [#49308]\r\n* Enable model test on upcoming `torch_tpu` backend (#49207) by @tengomucho in [#49207]\r\n* Fix encoder repetition penalty source rows after batch expansion (#49262) by @gitedmond in [#49262]\r\n* Preserve masked EOS scores in exponential decay length penalty (#49260) by @gitedmond in [#49260]\r\n* [CI] Avoid fetching all artifacts when only specific ones are needed (#49282) by @ydshieh in [#49282]\r\n* Revert \"[CI] ssh-runner: add optional cache_type input to switch between bucket and EFS runners (#49159)\" (#49270) by @ydshieh in [#49270]\r\n* Add Trainer.loss_is_scaled_for_ga to declare whether compute_loss already scales for gradient accumulation (#49240) by @qgallouedec in [#49240]\r\n* fix precomputed flash kwargs clashing (#47928) by @IlyasMoutawwakil in [#47928]\r\n* Reapply modular examples (#49230) by @Cyrilvallez in [#49230]\r\n* Remove some stuff that does not exist (#49232) by @Cyrilvallez in [#49232]\r\n* Remove rotaries deprecation cycle (#49228) by @Cyrilvallez in [#49228]\r\n* [`GTE`] Add memory mixin (#49234) by @vasqu in [#49234]\r\n* [CB] Make CB more device agnostic and add support for XPU (#49156) by @remi-or in [#49156]\r\n* correct final metrics due to resuming from checkpoint (#49208) by @SunMarc in [#49208]\r\n* [ESM/Persimmon/NllbMoe/ESMFold] use safetensors repos to fix Xet FUSE SIGBUS/IOError on bucket CI runners (#49194) by @ydshieh in [#49194]\r\n* Dev update (#49212) by @vasqu in [#49212]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @vasqu\r\n    * Smthn smthn (#49364)\r\n    * [`GTE`] Add memory mixin (#49234)\r\n    * Dev update (#49212)\r\n* @remi-or\r\n    * [Fix] Remove old test file with two ancient tests (#49280)\r\n    * [CB] Make CB more device agnostic and add support for XPU (#49156)\r\n    * 🚨 Attention 🚨 Deprecate \"paged|\" prefix for SDPA and flash  (#49112)\r\n    * 🚨 Attention 🚨 Make regular attention support CB (#49101)\r\n    * 🚨 [CB] 🚨 Fuse update for index and block table path (#49088)\r\n    * [Refactor] Make some flash-attention utils more readable (#49071)\r\n    * 🚨 [CB] 🚨 Little fixes before removing \"paged\" (#49069)","publishedAt":"2026-10-06T16:39:23.000Z","fetchedAt":"2026-10-06T18:43:32.034Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.19.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/84734e74-163d-4d12-b166-ffcf6749d563","r2Key":"releases/c8af3522bec5ad8c8d3c28f347b9d130ed85902fe0e5bd5aaa94da3343452167.png","r2Url":"https://media.releases.sh/releases/c8af3522bec5ad8c8d3c28f347b9d130ed85902fe0e5bd5aaa94da3343452167.png"}],"coverageCount":0},{"id":"rel_B2cwEAtWYUrjtsDJKXGzy","version":"v5.18.0","type":"feature","title":"Release 5.18.0","summary":"New model additions include Nemotron 3 Diarization (streaming speaker diarization, up to eight speakers), NemotronH Omni (joint text/image/video/audio reasoning), HyperCLOVAX Vision V2, and the GTE BERT-style encoder family. The release also carries flagged breaking changes across ROCm gpt-oss attention routing, DETR image processing, indexer layer_type remapping, vLLM video token counting, DINO modernization, and a Kernels version bump, plus a long tail of bug fixes spanning MoE expert parallelism, Qwen2.5-VL temporal RoPE, and DeepSeek-OCR-2 patching.","titleGenerated":"Transformers v5.18.0 adds Nemotron 3 Diarization, NemotronH Omni, HyperCLOVAX Vision V2, and GTE","titleShort":"Nemotron 3 Diarization, NemotronH Omni models land in Transformers","breaking":"major","importance":4,"content":"## New Model additions\r\n\r\n\r\n### Nemotron 3 Diarization\r\n\r\n<img width=\"1680\" height=\"900\" alt=\"image\" src=\"https://github.com/user-attachments/assets/fe735cb3-9e5b-43ad-8f60-9dec8425aec7\" />\r\n\r\nNemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine \"who spoke when\" in real-world audio. It supports both streaming and offline inference, handles up to eight speakers, and orders speaker outputs by each speaker's first arrival in the input audio.\r\n\r\nThe model uses the Arrival-Order Speaker Cache (AOSC) [1](https://huggingface.co/papers/2507.18446) and FIFO queue introduced for Streaming Sortformer [1](https://huggingface.co/papers/2507.18446), [2](https://huggingface.co/papers/2409.06656). A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms. With chunked inference, the maximum audio duration is not limited.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/nemotron3_diarization)\r\n* Add Nemotron3Diarization (#49056) by @eustlb in [#49056](https://github.com/huggingface/transformers/pull/49056)\r\n\r\n### NemotronH Omni\r\n\r\nNemotronH Omni is a multimodal reasoning model from NVIDIA that pairs the [NemotronH](https://huggingface.co/docs/transformers/main/en/model_doc/nemotron_h) hybrid\r\nMamba-Transformer language model with a [RADIO](https://huggingface.co/docs/transformers/main/en/model_doc/radio) vision encoder and an optional Parakeet-based sound encoder.\r\nImage (and video) patches are projected through a RADIO tower and a pixel-shuffle MLP into the language model's\r\nembedding space at the `<image>` / `<video>` context-token positions; audio clips are projected in the same way at\r\n`<audio>` positions. The result is a single autoregressive model that reasons jointly over text, images, video and\r\nsound.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/nemotron_h_omni)\r\n* Add support for Nemotron Omni (#46509) by @meatybobby in [#46509](https://github.com/huggingface/transformers/pull/46509)\r\n\r\n### HyperCLOVAX Vision V2\r\n\r\nHyperCLOVAX Vision V2 is a multimodal vision-language model developed by NAVER. It combines the [HyperClovaX](https://huggingface.co/docs/transformers/main/en/model_doc/hyperclovax) language model backbone with a [Qwen2.5-VL](https://huggingface.co/docs/transformers/main/en/model_doc/qwen2_5_vl) vision encoder. The model supports text, image, and video inputs and is capable of chain-of-thought reasoning via built-in thinking tokens (`<think>...</think>`).\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/hyperclovax_vision_v2)\r\n* add HyperClovaX Vision (#44314) by @jp1924 in [#44314](https://github.com/huggingface/transformers/pull/44314)\r\n\r\n### GTE\r\n\r\nGTE was proposed in [mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval](https://huggingface.co/papers/2407.19669) by Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li and Min Zhang.\r\n\r\nGTE is a BERT-style bidirectional encoder that replaces absolute position embeddings with RoPE, uses a gated MLP, and applies layer normalization after each residual connection. The same architecture backs Alibaba's `gte-*-v1.5`, `gte-multilingual-*` and `gte-en-mlm-*` checkpoints as well as Snowflake's `snowflake-arctic-embed-m-v2.0`.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/gte)\r\n* model: Add GTE to Transformers (#48416) by @harshaljanjani in [#48416](https://github.com/huggingface/transformers/pull/48416)\r\n\r\n\r\n## Breaking changes\r\n\r\n* 🚨 [ROCm] gpt-oss: route FA3 to aiter-flash-attn, generate ROCm fixtures (#46837) by @Abdennacer-Badaoui\r\n* 🚨 Speed up detr image processing (#48066) by @guarin\r\n* 🚨 Remap indexers layer_type (#48974) by @Cyrilvallez\r\n* 🚨 [vLLM] Fix video token counting for Transformers backend video inputs (Part 1) (#48894) by @harshaljanjani\r\n* 🚨 🚨Bring some dinos to modern standards (#46266) by @molbap\r\n* :rotating_light: [`Kernels`] Bump version (#48714) by @vasqu\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* fix incorrect hub tokenizer class (#48641) by @itazap\r\n* [CI] Deduplicate Nvidia/AMD CI reply comments and add headers (#48655) by @ydshieh\r\n* Update dev (#48654) by @vasqu\r\n* [tests] Fix NougatModelIntegrationTest: pin artifact revision and update golden values (#48638) by @ydshieh\r\n* [fix] Fix GlmOcr integration tests: wrong token IDs and image token decode bug (#48650) by @ydshieh\r\n* [GLM 5.3 Flash] Preserve original names when saving text checkpoints (#48676) by @Dovis01\r\n* Register activation kernel layers on XPU (#47858) by @jiqing-feng\r\n* Mirror the registered pytree flatten when building automatic dynamic shapes (#48578) by @IlyasMoutawwakil\r\n* Fix `AutoImageProcessor` requiring torchvision when only Pillow is installed (#48616) by @blipbyte\r\n* Assert cached decode matches recomputing without a cache (#48289) by @IlyasMoutawwakil\r\n* [higgs_audio_v2] Fix: use config.num_codebooks in audio labels tensor (Simple Fix!) (#48560) by @b-re-w\r\n* [`MoE`] Fix eager EP (#48653) by @vasqu\r\n* [MusicgenMelody] Fix conditioning silently dropped at generation step 0 (#48679) by @ydshieh\r\n* Fix deepspeed ci (#48640) by @SunMarc\r\n* Synthetic test assets (#48589) by @tarekziade\r\n* Skip the expert-parallel sentinel masking when expert parallelism is off (#48201) by @qgallouedec\r\n* [Fix] yolos offload issue (#48688) by @molbap\r\n* [serge] Fix 2 integration tests for model `minimax` failing with `output_mismatch` (tensor values differ (2)) (#48515) by @sergereview[bot]\r\n* [serge] Fix 2 integration tests for model `mistral` failing with `other` (other (2)) (#48429) by @sergereview[bot]\r\n* Fix Qwen2.5-VL temporal RoPE for fractional video intervals (#48669) by @yeyeyeping\r\n* Fix slow integration tests on XPU (#48611) by @jiqing-feng\r\n* QA: Added a `MemoryCleanupMixin` class for tests (#48681) by @tarekziade\r\n* [serge] Fix 1 integration test for model `flex_olmo` failing with `other` (other (1)) (#48668) by @sergereview[bot]\r\n* Fix D-FINE / RT-DETR main loss being computed over the denoising queries (#48528) by @stefan-it\r\n* [CI] Replace hardcoded username allowlists with author_association check in workflow triggers (#48712) by @ydshieh\r\n* Add flash_attention_4 in attn_implementation AutoModel docstring (#48684) by @3manifold\r\n* Keep the expert-parallel sentinel slots out of the router gradient (#48689) by @qgallouedec\r\n* [DeepseekV3] OOM cascade root-cause investigation (generator ref leak in conversion_mapping) (#48720) by @ydshieh\r\n* Use huggingface_hub httpx export (#48685) by @Wauplin\r\n* unifying device_mesh init to enable PP + TP inference (#48155) by @3outeille\r\n* Enable FSDP2 + expert parallelism via a 2-D (fsdp, tp) device mesh (#48516) by @qgallouedec\r\n* Switch daily CI to torch 2.14 — update expected outputs (#48750) by @ydshieh\r\n* fix videomae load error (#48675) by @sywangyi\r\n* Fix silently random-initializing `RTDetrModel`/`SEWDForCTC` loads (wrong `base_model_prefix`) (#48744) by @<NOT FOUND>\r\n* Fix Glm4vMoeIntegrationTest: offload_folder + MemoryCleanupMixin (#48776) by @ydshieh\r\n* [InternVL] Normalize num_patches before np.cumsum (#48469) by @lorenzozanee\r\n* Kernel api doc, part 2 (#46889) by @michaelbenayoun\r\n* Fix hidden state selection on Gemma4 assistant's first prefill step (#48704) by @glistening\r\n* Fix inkling embedding norm (#48786) by @Cyrilvallez\r\n* Fix TextToAudioPipeline crash for tokenizer-only models (#48505) by @jiqing-feng\r\n* Cast pixel values to the patch embedding dtype in DeepSeek-OCR-2 (#48632) by @jiqing-feng\r\n* add kernel mapping entry for RMSNormGated, KDA, Conv1D on XPU (#48702) by @kaixuanliu\r\n* Bump `peft` version requirement (#48716) by @shniubobo\r\n* [serge] Fix 12 integration tests for model `edgetam` failing with `import_or_config` (other (12)) (#48322) by @sergereview[bot]\r\n* Restore legacy tensor-parallel initialization for compatibility (#48797) by @3outeille\r\n* Fix kosmos flaky test (#48800) by @IlyasMoutawwakil\r\n* Add integration tests for MuseGlimmerAssistantModel (#48796) by @ydshieh\r\n* Fix `AutoModel.from_pretrained` not restoring `modules_to_save` weights (#48595) by @shniubobo\r\n* [docs] mrope and axial rope (#48717) by @stevhliu\r\n* Fix decoder only path for older bert variants (#48785) by @Cyrilvallez\r\n* [Fix] Clean-up ternaries in the DeepSeek family (#48447) by @remi-or\r\n* [utils] Add MUSA support for Flash Attention 2 (#48612) by @XiaomingFun233\r\n* [generate] Drop attention mask early without padding (#48814) by @Cyrilvallez\r\n* Pin utf-8 in test_can_init_all_missing_weights source read (Windows non-UTF-8 locale fix) (#48819) by @dltsum\r\n* Relax static-cache tolerance in test_generate_with_static_cache (1e-5 → 5e-5) (#48815) by @ydshieh\r\n* Detect nested rope_parameters without relying on layer_types (#48798) by @hmellor\r\n* Support for MoE in the GGUF integration (#48529) by @SunMarc\r\n* [docs] gguf (#46357) by @stevhliu\r\n* [apply_chat_template] pass sampling_rate to call (#48794) by @eustlb\r\n* [CI] Add link checker (#48160) by @stevhliu\r\n* Qwen3.8 GGUF (#48660) by @SunMarc\r\n* docs: fix broken #combining-with-fsdp2 anchor in expert_parallelism.md (#48854) by @ydshieh\r\n* [GPTNeoXJapanese] Fix RoPE ignoring partial_rotary_factor (#48652) by @blipbyte\r\n* QA: deactivate rule 41 (#48852) by @tarekziade\r\n* Fix `reset` on the dynamic cache layers (#48809) by @jiqing-feng\r\n* [generate] Make all methods and logits processors agnostic to lm_head output size (#48846) by @Cyrilvallez\r\n* Fix kosmos (#48853) by @Cyrilvallez\r\n* Add test for causal only variant of some encoder-decoder models (#48760) by @nandan2003\r\n* Fix ESMFold2 ligand iPLDDT weighting: mol_type non-polymer code is 3, not 4 (#48831) by @faustomilletari\r\n* Fix flaky RfDetr test_save_load (#48869) by @ydshieh\r\n* fix(ci): harden GitHub Actions workflows (#48160) (#48834) by @hf-security-analysis[bot]\r\n* Fix Glm4MoeIntegrationTest: split into 3 tests, switch to GLM-4.5-Air (#48820) by @ydshieh\r\n* Fix `StaticCache` for Mllama and enable `torch.compile` (#48141) by @jiqing-feng\r\n* Fix Minimax M2 partial_rotary_factor by remapping legacy rotary_dim (#48486) by @dajiaohuang\r\n* Fix Pixtral image processor do_pad option (#48872) by @jzakrzew\r\n* [CB] Add pause mechanism (#48462) by @remi-or\r\n* Fix DFlash sampled candidates losing the batch dimension (#48281) by @VaggelisGian\r\n* Fix sliding window cache when assistant model calls generate (#48280) by @VaggelisGian\r\n* Fix GLM rope_parameters dict shared mutation across test instances (#48895) by @ydshieh\r\n* [`feat`] Allow untying `hidden_states[-1]` from `last_hidden_state` via the model config (#48087) by @tomaarsen\r\n* Fix Idefics2 padding when the first example has no image (#48753) by @Arnavsharma2\r\n* Don't fail model loading when the accelerator can't report free memory (#48136) by @studioego\r\n* [Improvement] Rework the docstring and comments of mHC (#48888) by @remi-or\r\n* [serge] Fix 3 integration tests for model `deepseek_vl` failing with `other` (other (3)) (#48536) by @sergereview[bot]\r\n* Add auto_docstring task model overrides (#47815) by @guarin\r\n* Make LongcatFlashConfig self-consistent without building the model (#48899) by @hmellor\r\n* [`DSA`] Only save latents on dsa with indexer as well (#48876) by @vasqu\r\n* [Chat Parsing] Coerce oneOf tool arguments (#48719) by @yonigozlan\r\n* Always materialize the causal mask in Doge so sdpa stays causal (#48821) by @blipbyte\r\n* Fix docstring argument names that don't match signatures (#48474) by @IshaanPotle\r\n* Mark test_generate_with_static_cache as flaky for olmo and bigbird_pegasus (#48856) by @ydshieh\r\n* Clean up some MoE models' integration tests (#48833) by @ydshieh\r\n* Fix T5 tied weights order for lm_head (#48238) by @vaibhavmashal\r\n* Size the device_map buffer from the largest leaf module (#47211) by @dhruv7477\r\n* Fix compile cache tests (#48923) by @Cyrilvallez\r\n* Update arXiv citation in NeoMME model doc (#48926) by @tonywu71\r\n* Standardise Aria's MoE onto the experts interface (#48907) by @hmellor\r\n* Fix possessive typo in Whisper long-form warning (#48877) by @erikdw\r\n* Modular conversion small fix (#48934) by @zucchini-nlp\r\n* fix(vibevoice-asr): use integer ceiling division for audio token count (#48864) by @ege-arhan\r\n* Improve lazy import error messages (#48602) by @karatarassul4-max\r\n* Fix tests due to dropping attn mask (#48903) by @SunMarc\r\n* docs: fix docstring parameters that do not match signatures (#48904) by @simpleqt\r\n* DOC: Improve documentation for remap_legacy_layer_types function (#48651) by @mathewOracle\r\n* Align special tokens on the text config (#48847) by @qgallouedec\r\n* docs: fix dead doc links in i18n READMEs and the xlnet docstring (#48905) by @simpleqt\r\n* Fix AXK2 integration test: update CUDA expected text and rename class (#48941) by @ydshieh\r\n* Resolve the Hub revision once per load instead of passing a private _commit_hash around (#47611) by @Wauplin\r\n* Keep image processor backends in sync on keys and dtypes (#48739) by @yupengtang\r\n* Fix infeasible cost matrix errors in hugarian matcher losses (#47730) by @guarin\r\n* Fix more stuff (#48979) by @zucchini-nlp\r\n* Skip an unnecessary image copy in the torchvision image normalization path (#48897) by @jzakrzew\r\n* Better attn default for gguf (#48935) by @SunMarc\r\n* Add pose estimation keypoint preprocessing to Sapiens2ImageProcessor (#47199) by @Sainava\r\n* Fix image processor class-level size mutation and min_pixels handling (#48916) by @Gracy769\r\n* Contract the mamba2 chunk scan with einsum instead of broadcast-then-sum (#48978) by @tarekziade\r\n* Fix typos in DeepseekV4 comments (#48953) by @Janiarafath\r\n* Reduce peak memory in examples_torch CI job (OOM fix) (#48983) by @ydshieh\r\n* Fix vibevoice TTS batched audio index (#48902) by @ebezzam\r\n* [CB] [Major] Upgrade the cache to support different attention types (#47809) by @remi-or\r\n* Fix SwitchTransformers Top1 router: raw logits, expert capacity accounting, and router losses (#48421) by @yurekami\r\n* Fix Qwen3OmniMoeIntegrationTest OOM (#48987) by @ydshieh\r\n* Let `prefix_allowed_tokens_fn` override model `-inf` and raise an exception on unsatisfiable generation constraints. (#48927) by @ksh108405\r\n* [generate] Simplify candidate generators by removing required `update_candidate_strategy` (#48982) by @Cyrilvallez\r\n* [generate] Always correctly restrict assisted decoding with max length/eos token  (#48981) by @Cyrilvallez\r\n* QA: applied ruff rule PLW1514 (#48990) by @tarekziade\r\n* Update tokenizer gguf support  (#48656) by @SunMarc\r\n* [vLLM] Fix video token counting for Transformers backend video inputs (Part 2) (#48900) by @harshaljanjani\r\n* Update ggml kernels path  (#48991) by @SunMarc\r\n* Enable compressed-tensors FP8 kernels on MPS (torch >= 2.15) (#48985) by @Isalia20\r\n* Fix mps autocast handling in rotary embeddings (#49006) by @Isalia20\r\n* Fix static cache per layer head shapes (#48619) by @dacorvo\r\n* Summarization examples: download NLTK punkt_tab, not punkt (#49014) by @davanstrien\r\n* [docs] Fix legacy hf CLI references (transformers) (#48988) by @Wauplin\r\n* Fix RecurrentGemma compiled generation with StaticCache (#48961) by @sywangyi\r\n* updated GraniteMoeHybrid expectations (CPU) (#49016) by @tarekziade\r\n* OpenVINO HF Exporter (#47003) by @IlyasMoutawwakil\r\n* fix noisy comments (#49013) by @tarekziade\r\n* Fix assisted decoding for VLM due to dropping attn mask  (#49019) by @SunMarc\r\n* Deprecate min-max pixels (#49021) by @zucchini-nlp\r\n* Keep special token ids the tokenizer does not define (#48708) by @albertvillanova\r\n* Added more usage of MemoryCleanupMixin (#49011) by @tarekziade\r\n* Shorten noisy OpenVINO SDPA comment (#49033) by @tarekziade\r\n* PEFT x Dtensor-based TP integration (#48485) by @michaelbenayoun\r\n* fix(zamba):  add use_associative_scan config flag to avoid torch.compile slowdown (#48331) by @msnliu\r\n* Scope GITHUB_TOKEN permissions per job (#49046) by @hf-security-analysis[bot]\r\n* Fix mask creation not being skipped under `torch.compile` (#48975) by @jiqing-feng\r\n* [Chat] Loading GGUF models served with the Chat CLI (#49031) by @ariG23498\r\n* Pin GitHub Actions to commit SHAs (#49049) by @hf-security-analysis[bot]\r\n* Only seed numpy in BigBird block-sparse attention during training (#49037) by @jiqing-feng\r\n* QA: Fix llama4 leak (#49042) by @tarekziade\r\n* Fix startup failures: drop permissions reusable-workflow callers cannot grant (#49055) by @paulinebm\r\n* Fix startup failures: drop pull-requests: read from check_failed_tests.yml (#49057) by @paulinebm\r\n* Register the remaining mamba-ssm kernel layers on XPU (#49035) by @jiqing-feng\r\n* Add workflow for building XPU CI Docker images (#49039) by @regisss\r\n* Fix assistant masks for processors (#48793) by @Rocketknight1\r\n* Add MPS maintainer (#49052) by @Rocketknight1\r\n* [docs] Loading behavior (#49023) by @stevhliu\r\n* [Nemotron3Diarization] nit: hub pr merged to main (#49065) by @eustlb\r\n* fix(ci): harden GitHub Actions workflows (#49057) (#49059) by @hf-security-analysis[bot]\r\n* Honor `config.output_router_logits` in the MoE VLM wrappers (#48885) by @qgallouedec\r\n* Fix XPU Docker image build (#49073) by @regisss\r\n* Fix assisted eos token condition (#49075) by @Cyrilvallez\r\n* [serge] Fix OOM in Moshi integration tests with MemoryCleanupMixin (#48839) by @sergereview[bot]\r\n* Fix odd head_dim validation for RoPE configurations (#48524) by @somuai\r\n* Keep already-decoded array arguments unchanged (#49062) by @yonigozlan\r\n* Fix NaN in Parakeet eager attention with padded batches (#49070) by @ArthurZucker\r\n* Bump huggingface_hub upper bound to <3.0 (transformers) (#49083) by @Wauplin\r\n* Keep `attention_mask` as `None` in OPT's causal mask creation (#49002) by @jiqing-feng\r\n* Add `Trainer.end` (#48875) by @qgallouedec\r\n* Add image processing tester init (#48829) by @guarin\r\n* [Executorch] Add MLX recipe (#48910) by @metascroy\r\n* [Zamba] Fix associative scan breaking ONNX export and OOM in integration test (#49092) by @ydshieh\r\n* Fix RGB early-return skipping PNG tRNS compositing (#49005) by @cs-fisha\r\n* Use namespaced dataset ids in docs and PyTorch examples (#49017) by @davanstrien\r\n* Fix MaskFormerSwin attention mask dtype to follow hidden states (#49032) by @kaixuanliu\r\n* [NemotronH-Omni] Fix device mismatch in test tensor creation (#49102) by @ydshieh\r\n* [serge] Fix 2 integration tests for model `cvt` failing with `output_mismatch` (tensor values differ (2)) (#49076) by @sergereview[bot]\r\n* [serge] Fix 2 integration tests for model `hy_v3` failing with `output_mismatch` (tensor values differ (2)) (#49068) by @sergereview[bot]\r\n* [serge] Fix 2 integration tests for model `pvt_v2` failing with `other` (#49067) by @sergereview[bot]\r\n* Keep the eos ids the config declares (#49082) by @albertvillanova\r\n* Use grouped_mm on TPU devices under torch.compile (#49097) by @salkan0\r\n* [serge] Fix 2 integration tests for model `jamba` failing with `other` (other (2)) (#49044) by @sergereview[bot]\r\n* Remap the legacy Gemma 1 hidden_act in the config post-init (#49084) by @PCfVW\r\n* Fix exporters import on torch < 2.9 (is_contiguous_or_false) (#49124) by @ydshieh\r\n* [gemma3n] fix audio test fixture: use hf_hub_download instead of load_dataset (#49128) by @ydshieh\r\n* qwen3 models map to wrong tokenizer class on the hub (#49116) by @itazap\r\n* Restore the Unicode whitespace set in the GPT-SW3 tokenizer (#48912) by @David-Wu1119\r\n* [AMD] Fix some integration tests (#49153) by @Abdennacer-Badaoui\r\n* [Gemma] Fix test_model_7b_fp16_static_cache expected value for cuda 8 after #49084 (#49133) by @ydshieh\r\n* Fix UMT5 decoder self-attention not being causal (#49135) by @the-cross-art\r\n* Skip flash tests that fall back to a hub kernel when kernels is missing (#49129) by @Abdennacer-Badaoui\r\n* Auto generate model inits (#47829) by @guarin\r\n* Fix get_json_schema dropping items/enum for unions of list/dict/Literal types (#49136) by @JoeyTan21\r\n* Pick the default flash implementation based on the current hardware (#49109) by @Abdennacer-Badaoui\r\n* Deprecate the use_mamba_kernels config flag that no longer has any effect (#49155) by @albertvillanova\r\n* [CI] ssh-runner: add optional cache_type input to switch between bucket and EFS runners (#49159) by @ydshieh\r\n* [generation] Encode multimodal data only once (#45783) by @zucchini-nlp\r\n* No more -hf repo names for ESMC (#49158) by @Rocketknight1\r\n* [MPS] Remove cu_seqlens_k clone workaround for metal-flash-sdpa (#49091) by @Isalia20\r\n* [PerceptionLM] Restore test_inputs_embeds overrides to fix flaky test (#49165) by @ydshieh\r\n* [CB] Fix failing tests discovered when using the B200 (#49171) by @remi-or\r\n* [imagegpt/vilt/trocr] fix cache fixture tests: use hf_hub_download instead of load_dataset (#49178) by @ydshieh\r\n* Fix lost call (#49163) by @molbap\r\n* Document image_like_kwargs (#49180) by @guarin\r\n* [CacheHardIntegrationTest] use dedicated safetensors repo to fix Xet bucket cache corruption (#49182) by @ydshieh\r\n* Document running the example scripts on Hugging Face Jobs (#49050) by @davanstrien\r\n* Fix backslash handling in generate flag values in transformers chat (#48709) by @JHC56\r\n* Finish removing the MPS autocast workaround (#49157) by @rubenG1009\r\n* Fix Trainer checkpoint resume crashing on CPU with multiple processes (#49123) by @neevmodh\r\n* Fix command syntax for optimum-cli export (#46451) by @Ahwar\r\n* QA: fix noisy comment checker (#49184) by @tarekziade\r\n* update to torch 2.14 (#49199) by @sywangyi\r\n* Video processors - general maintenance (#48251) by @zucchini-nlp\r\n* [Parakeet] Convert NeMo's stochastic depth to layerdrop (#49191) by @Deep-unlearning\r\n* [MPS] Let mps sdpa handle grouped query attention directly (#49187) by @Isalia20\r\n* Replace datasets that no longer load with maintained uploads (#49018) by @davanstrien\r\n* Fail fast on eval OOM under `auto_find_batch_size` (#49198) by @qgallouedec\r\n* Fix SequenceBiasLogitsProcessor edge cases: token id 0 and prefix equal to context (#49117) by @lucaluo925\r\n* Map bare list, tuple and dict annotations to the right JSON schema type in get_json_schema (#49145) by @825pranav\r\n* [`Kernels`] Sync mamba version (#49205) by @vasqu\r\n* gguf user defined tokens (#49008) by @SunMarc\r\n* QA: restore masking_utils export comment with a noqa (#49209) by @tarekziade\r\n* Fix deepstack features for mixed-input (#49177) by @zucchini-nlp\r\n* Fix stale _added_tokens_encoder entries in cpmant and wav2vec2 (#47440) by @ishan-1010\r\n* Fix additional_special_tokens data loss with extra_special_tokens (#47848) by @erichanwang\r\n* Fix MPS GQA version gating (#49210) by @Isalia20\r\n* Add Strix Halo (gfx1151) Atlas Inference Hub-kernel path for Qwen3.5/3.6/3.8 Gated DeltaNet (#49127) by @AzeezIsh\r\n* Fix MiniMax M3 partial 3D vision rotary embeddings (#49164) by @cuichenx\r\n* Fix missing router_logits in Qwen3.5-MoE and other MoE models (#49179) by @zucchini-nlp\r\n* [Nemotron3Diarization] fix streaming last stft frame dropped (#49167) by @eustlb\r\n\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @harshaljanjani\r\n    * model: Add GTE to Transformers (#48416)\r\n    * [vLLM] Fix video token counting for Transformers backend video inputs (Part 2) (#48900)\r\n    * 🚨 [vLLM] Fix video token counting for Transformers backend video inputs (Part 1) (#48894)\r\n* @eustlb\r\n    * [Nemotron3Diarization] fix streaming last stft frame dropped (#49167)\r\n    * [Nemotron3Diarization] nit: hub pr merged to main (#49065)\r\n    * Add Nemotron3Diarization (#49056)\r\n    * [apply_chat_template] pass sampling_rate to call (#48794)\r\n* @tarekziade\r\n    * QA: restore masking_utils export comment with a noqa (#49209)\r\n    * QA: fix noisy comment checker (#49184)\r\n    * QA: Fix llama4 leak (#49042)\r\n    * Shorten noisy OpenVINO SDPA comment (#49033)\r\n    * Added more usage of MemoryCleanupMixin (#49011)\r\n    * fix noisy comments (#49013)\r\n    * updated GraniteMoeHybrid expectations (CPU) (#49016)\r\n    * QA: applied ruff rule PLW1514 (#48990)\r\n    * Contract the mamba2 chunk scan with einsum instead of broadcast-then-sum (#48978)\r\n    * QA: deactivate rule 41 (#48852)\r\n    * QA: Added a `MemoryCleanupMixin` class for tests (#48681)\r\n    * Synthetic test assets (#48589)\r\n* @ydshieh\r\n    * [CacheHardIntegrationTest] use dedicated safetensors repo to fix Xet bucket cache corruption (#49182)\r\n    * [imagegpt/vilt/trocr] fix cache fixture tests: use hf_hub_download instead of load_dataset (#49178)\r\n    * [PerceptionLM] Restore test_inputs_embeds overrides to fix flaky test (#49165)\r\n    * [CI] ssh-runner: add optional cache_type input to switch between bucket and EFS runners (#49159)\r\n    * [Gemma] Fix test_model_7b_fp16_static_cache expected value for cuda 8 after #49084 (#49133)\r\n    * [gemma3n] fix audio test fixture: use hf_hub_download instead of load_dataset (#49128)\r\n    * Fix exporters import on torch < 2.9 (is_contiguous_or_false) (#49124)\r\n    * [NemotronH-Omni] Fix device mismatch in test tensor creation (#49102)\r\n    * [Zamba] Fix associative scan breaking ONNX export and OOM in integration test (#49092)\r\n    * Fix Qwen3OmniMoeIntegrationTest OOM (#48987)\r\n    * Reduce peak memory in examples_torch CI job (OOM fix) (#48983)\r\n    * Fix AXK2 integration test: update CUDA expected text and rename class (#48941)\r\n    * Clean up some MoE models' integration tests (#48833)\r\n    * Mark test_generate_with_static_cache as flaky for olmo and bigbird_pegasus (#48856)\r\n    * Fix GLM rope_parameters dict shared mutation across test instances (#48895)\r\n    * Fix Glm4MoeIntegrationTest: split into 3 tests, switch to GLM-4.5-Air (#48820)\r\n    * Fix flaky RfDetr test_save_load (#48869)\r\n    * docs: fix broken #combining-with-fsdp2 anchor in expert_parallelism.md (#48854)\r\n    * Relax static-cache tolerance in test_generate_with_static_cache (1e-5 → 5e-5) (#48815)\r\n    * Add integration tests for MuseGlimmerAssistantModel (#48796)\r\n    * Fix Glm4vMoeIntegrationTest: offload_folder + MemoryCleanupMixin (#48776)\r\n    * Switch daily CI to torch 2.14 — update expected outputs (#48750)\r\n    * [DeepseekV3] OOM cascade root-cause investigation (generator ref leak in conversion_mapping) (#48720)\r\n    * [CI] Replace hardcoded username allowlists with author_association check in workflow triggers (#48712)\r\n    * [MusicgenMelody] Fix conditioning silently dropped at generation step 0 (#48679)\r\n    * [fix] Fix GlmOcr integration tests: wrong token IDs and image token decode bug (#48650)\r\n    * [tests] Fix NougatModelIntegrationTest: pin artifact revision and update golden values (#48638)\r\n    * [CI] Deduplicate Nvidia/AMD CI reply comments and add headers (#48655)\r\n* @molbap\r\n    * Fix lost call (#49163)\r\n    * 🚨 🚨Bring some dinos to modern standards (#46266)\r\n    * [Fix] yolos offload issue (#48688)\r\n* @remi-or\r\n    * [CB] Fix failing tests discovered when using the B200 (#49171)\r\n    * [CB] [Major] Upgrade the cache to support different attention types (#47809)\r\n    * [Improvement] Rework the docstring and comments of mHC (#48888)\r\n    * [CB] Add pause mechanism (#48462)\r\n    * [Fix] Clean-up ternaries in the DeepSeek family (#48447)\r\n* @jiqing-feng\r\n    * Keep `attention_mask` as `None` in OPT's causal mask creation (#49002)\r\n    * Register the remaining mamba-ssm kernel layers on XPU (#49035)\r\n    * Only seed numpy in BigBird block-sparse attention during training (#49037)\r\n    * Fix mask creation not being skipped under `torch.compile` (#48975)\r\n    * Fix `StaticCache` for Mllama and enable `torch.compile` (#48141)\r\n    * Fix `reset` on the dynamic cache layers (#48809)\r\n    * Cast pixel values to the patch embedding dtype in DeepSeek-OCR-2 (#48632)\r\n    * Fix TextToAudioPipeline crash for tokenizer-only models (#48505)\r\n    * Fix slow integration tests on XPU (#48611)\r\n    * Register activation kernel layers on XPU (#47858)\r\n* @Wauplin\r\n    * Bump huggingface_hub upper bound to <3.0 (transformers) (#49083)\r\n    * [docs] Fix legacy hf CLI references (transformers) (#48988)\r\n    * Resolve the Hub revision once per load instead of passing a private _commit_hash around (#47611)\r\n    * Use huggingface_hub httpx export (#48685)\r\n* @meatybobby\r\n    * Add support for Nemotron Omni (#46509)\r\n* @Sainava\r\n    * Add pose estimation keypoint preprocessing to Sapiens2ImageProcessor (#47199)\r\n* @jp1924\r\n    * add HyperClovaX Vision (#44314)","publishedAt":"2026-09-30T16:46:27.000Z","fetchedAt":"2026-09-30T17:14:51.408Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.18.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/fe735cb3-9e5b-43ad-8f60-9dec8425aec7","r2Key":"releases/2378aeb9395625b573224b4ef4e40e1c2f4f69717567511fca3fa4a6aa87d0c1.png","r2Url":"https://media.releases.sh/releases/2378aeb9395625b573224b4ef4e40e1c2f4f69717567511fca3fa4a6aa87d0c1.png"}],"coverageCount":0},{"id":"rel_kLwE0cXxxSuumg3fAfjWx","version":"v5.17.0","type":"feature","title":"Release 5.17.0","summary":"Adds support for seven new architectures: HY4-Preview (780B MoE), VibeVoice, NeoMME, Fun-ASR-Nano, KimiLinear, and Canary-1B-v2. Vision rotary embeddings (2D/3D) are standardized into a unified RoPE module, requiring migration for custom vision models.","titleGenerated":"Transformers v5.17.0 adds HYV4, VibeVoice, and six new model architectures","titleShort":"Seven new models; vision RoPE refactor breaking","breaking":"minor","importance":4,"content":"# Release v5.17.0\r\n\r\n\r\n## New Model additions\r\n\r\n### HYV4\r\n\r\n<img width=\"1503\" height=\"827\" alt=\"image\" src=\"https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9\" />\r\n\r\n\r\nHy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per\r\ntoken. Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every\r\ntoken to 8 of them. The context window is 1M tokens.\r\n\r\nThe architecture combines four features:\r\n\r\n- **Multi-head Latent Attention (MLA)** compresses keys and values into a low-rank latent\r\n  (`kv_lora_rank`) that `kv_b_proj` expands back to one key/value per query head.\r\n- **DeepSeek Sparse Attention (DSA)** selects `index_topk` keys per query with a lightweight indexer.\r\n  Following [IndexShare](https://huggingface.co/papers/2603.12201), only the layers marked `\"full\"`\r\n  in `indexer_types` run an indexer; `\"shared\"` layers reuse the previous full layer's selection.\r\n- **Gated MLA with learnable attention sinks**, where each head owns a sink logit that participates\r\n  in the softmax and contributes no value, as in [GPT-OSS](./gpt_oss).\r\n- **Independent Hyper-Connections (iHC)** replace the plain residual path with `hc_mult` parallel\r\n  residual streams that are collapsed before, and redistributed after, every sublayer.\r\n\r\nThe implementation does not execute the multi-token prediction (MTP) layers. Released checkpoints\r\nkeep those weights so that other runtimes can use them for speculative decoding; they are ignored\r\nat load time.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/hy_v4)\r\n* Add h4 (#48473) by @ArthurZucker in [#48473](https://github.com/huggingface/transformers/pull/48473)\r\n\r\n### VibeVoice\r\n\r\n<img width=\"2140\" height=\"1188\" alt=\"image\" src=\"https://github.com/user-attachments/assets/29ccea01-a855-4d4f-a6af-61bc4fc883a4\" />\r\n\r\n[VibeVoice](https://huggingface.co/papers/2508.19205) is a novel framework for synthesizing high-fidelity, long-form speech with multiple speakers by employing a next-token diffusion approach within a Large Language Model (LLM) structure. It's designed to capture the authentic conversational \"vibe\" and is particularly suited for generating audio content like podcasts and multi-participant audiobooks.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/vibevoice)\r\n* Implement VibeVoice  (#40546) by @pengzhiliang in [#40546](https://github.com/huggingface/transformers/pull/40546)\r\n\r\n### NeoMME\r\n\r\nNeoMME is a family of efficient 260M and 800M parameter multimodal-native multilingual foundation encoders from H Company. It processes multilingual text tokens and raw image patches in a single bidirectional Transformer encoder, without a separately pretrained vision tower or causal language model.\r\n\r\nNeoMME-Retriever is a model fine-tuned from the NeoMME backbone for visual document retrieval with joint late-interaction and dense objectives. It takes text queries and documents (text or page screenshots) and produces multi-vector embeddings for MeanMaxSim scoring (late-interaction) and mean-pooled embeddings for cosine similarity (dense).\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/neomme)\r\n* Add NeoMME and NeoMME-Retriever (#47992) by @tonywu71 in [#47992](https://github.com/huggingface/transformers/pull/47992)\r\n\r\n### Fun-ASR-Nano\r\n\r\nFun-ASR-Nano is an 800M-parameter end-to-end speech recognition model developed by Alibaba DAMO Academy's FunAudioLLM team. It achieves state-of-the-art performance on Chinese, English, and Japanese ASR benchmarks while being significantly smaller than comparable models.\r\n\r\nKey features are\r\n- **Chinese, English, and Japanese**, including 7 Chinese dialects and 26 regional accents\r\n- **Hotword customization** for domain-specific vocabulary\r\n- **Native punctuation** output (no separate punctuation model needed)\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/fun_asr_nano)\r\n* Add Fun-ASR-Nano model (#46180) by @LauraGPT in [#46180](https://github.com/huggingface/transformers/pull/46180)\r\n\r\n### KimiLinear\r\n\r\nKimi Linear is a hybrid linear attention architecture from Moonshot AI, introduced in\r\n[Kimi Linear: An Expressive, Efficient Attention Architecture](https://huggingface.co/papers/2510.26692).\r\n\r\nAt its core is **Kimi Delta Attention (KDA)**, a refinement of [Gated DeltaNet](https://huggingface.co/papers/2412.06464)\r\nthat gives each key channel its own forget gate, so the recurrent state decays per channel instead of per head. KDA is\r\nused in most layers; every fourth layer keeps a full-attention block that reuses DeepSeek-V3's Multi-head Latent\r\nAttention (MLA), and the feed-forward blocks are DeepSeek-V3-style MoE with a shared expert.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/kimi_linear)\r\n* Kimi linear (#48250) by @remi-or in [#48250](https://github.com/huggingface/transformers/pull/48250)\r\n\r\n### Canary\r\n\r\nCanary-1B-v2, a fast, robust multilingual model for Automatic Speech Recognition (ASR) and Speech-to-Text Translation (AST):\r\n\r\nCanary reuses the [Fast Conformer](https://huggingface.co/papers/2305.05084) encoder from [Parakeet](./parakeet.md) (loaded through [`ParakeetEncoder`] / [`ParakeetEncoderConfig`]) and pairs it with a Transformer decoder that uses fixed sinusoidal positional embeddings, cross-attention to the encoder outputs and tied input/output embeddings. The task is selected through a decoder prompt prefix built by [`CanaryProcessor`] of the form `<|startofcontext|> <|startoftranscript|> <|emo:undefined|> <source_lang> <target_lang> <pnc|nopnc> <|noitn|> <|notimestamp|> <|nodiarize|>`, where `source_lang == target_lang` selects transcription and otherwise selects translation.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/canary)\r\n* model: Add NVIDIA Canary-1B-v2 to Transformers (#46825) by @harshaljanjani in [#46825](https://github.com/huggingface/transformers/pull/46825)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nVision rotary embeddings (2D/3D) have been standardized into a unified RoPE frequency computation module, so users with custom vision models relying on attention-layer-level or model-specific RoPE grid interleaving logic must migrate to the new centralized `modeling_rope_utils.py` implementation.\r\n* :rotating_light: Vision (2d/3d) rotary embeddings  (#48105) by @zucchini-nlp\r\n\r\n\r\n\r\n## Generation\r\n\r\nGeneration improvements include a performance optimization that avoids unnecessary accelerator synchronization on every decode step (reducing per-step overhead), and a fix to prevent unconditional downloading of remote hub files during generation. Several correctness fixes were also applied, including enforcing auto-compile cache checks for encoder-decoder models, standardizing `past_key_values` naming in AfMoE, and resolving flaky export and integration test failures.\r\n\r\n\r\n* [`Generate`] Avoid unconditionally downloading remote hub file (#48620) by @vasqu in [#48620]\r\n* [generate] stop synchronizing the accelerator on every decode step (#47975) by @SunMarc in [#47975]\r\n* [AfMoE] Standardize past_key_values argument naming across forward and generate (#48430) by @shenhuaqingshi in [#48430]\r\n* Fix MTP generation test regex gate for escaped layer ignore keys (#48003) (#48262) by @Noxtimo in [#48262]\r\n* fix(generation): Enforce the auto-compile cache check for encoder-decoder models (#48364) by @harshaljanjani in [#48364]\r\n* [serge] Fix 4 integration tests for model `generation` failing with `output_mismatch` (list output differs (4)) (#48133) by @sergereview[bot] in [#48133]\r\n* [VibeVoice] Skip generate export tests (flaky) (#48396) by @ydshieh in [#48396]\r\n\r\n\r\n## Cache\r\n\r\nFixed several cache-related bugs, including a quantized cache issue in VibeVoice, incorrect rejection of non-static cache implementations in VoxtralRealtime, missing auto-compile cache checks for encoder-decoder models, and a silent failure when paged attention is called without a cache. Documentation was also updated to clarify `ContinuousBatchingConfig` usage and sliding window model limitations.\r\n\r\n\r\n* vibevoice: fix bug for quant cache (#48487) by @kaixuanliu in [#48487]\r\n* Fix VoxtralRealtime rejecting non-static cache implementations (#48082) by @jiqing-feng in [#48082]\r\n* Raise when a paged attention forward is called with no cache (#48297) by @qgallouedec in [#48297]\r\n* [docs] Pass ContinuousBatchingConfig and sliding window models  (#48381) by @stevhliu in [#48381]\r\n* Retry get_daily_ci_runs on stale GitHub API cache (#48374) by @ydshieh in [#48374]\r\n\r\n\r\n## Kernels\r\n\r\nKernel support was improved with fixes for nested FLA kernel imports when only `fla-core` is installed, a warning when hub-kernel functions silently fall back to slower pure-PyTorch reference implementations, and the ability to register standalone functions (e.g., RoPE) in `KernelConfig` with optional non-inheritance of default mappings. Additional fixes include corrected repository paths for ESMFold2 kernels and updated documentation for `KernelConfig` customization.\r\n\r\n\r\n* Support nested FLA kernel imports for fla-core (#48221) by @DimensionSTP in [#48221]\r\n* Warn once when a hub-kernel function falls back to its reference PyTorch path (#48185) by @qgallouedec in [#48185]\r\n* [docs] Kernel updates (#48465) by @stevhliu in [#48465]\r\n* [`Kernels`] Enable functions into kernels registry and allow non inheritance (#48443) by @vasqu in [#48443]\r\n* Fix kernel commit and repo paths for ESMFold2 (#48186) by @Rocketknight1 in [#48186]\r\n\r\n\r\n## Quantization\r\n\r\nFixed several quantization bugs, including a quant cache issue in VibeVoice, incorrect FP8 embedding handling for Qwen models, missing FP8 tensor parallelism layer overrides, and unnecessary MXFP4 weight dequantization on XPU devices.\r\n\r\n\r\n* fix qwen4exp-fp8 ple embedding (#48368) by @JJJYmmm in [#48368]\r\n* Keep MXFP4 weights quantized on XPU when use_kernels is set (#47923) by @jiqing-feng in [#47923]\r\n* Fix missing FP8 TP layer overrides (#48343) by @changwangss in [#48343]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* MRoPE continued (#48594) by @zucchini-nlp in [#48594]\r\n* [fix] Update stale expected strings in HunYuanVL integration tests (#48646) by @ydshieh in [#48646]\r\n* [Quantizaiton]support 5/6/7 bits in AutoRound (#48481) by @wenhuach21 in [#48481]\r\n* [fix] Update stale golden values and fix expected_logits shape in FlavaForPreTraining integration tests (#48639) by @ydshieh in [#48639]\r\n* Fix YOLOS device mismatch with device_map=\"auto\" (#46886) by @swankystark in [#46886]\r\n* Honor `shift_labels` in decoder-only LLM/VLM losses (#48493) by @qgallouedec in [#48493]\r\n* [KimiLinear] Fix test_cpu_offload: set num_local_experts=4 in model tester (#48624) by @ydshieh in [#48624]\r\n* [docs] Per-layer config (#48601) by @stevhliu in [#48601]\r\n* Fix `generate_flags` parsing in `transformers chat` (#48597) by @SunMarc in [#48597]\r\n* [fix] Fix how we read package versions - triggered by torch 2.14+ (#48615) by @ydshieh in [#48615]\r\n* [Docker] Upgrade CPU torch to <=2.14.0, torchcodec to <=0.16.0 (#48614) by @ydshieh in [#48614]\r\n* Fix AttributeError in gradient_checkpointing_enable(offload=True) (#48590) by @tarekziade in [#48590]\r\n* Another day fixing CI (#48591) by @zucchini-nlp in [#48591]\r\n* esmfold2: keep `distogram_head` in fp32 as well (#48488) by @kaixuanliu in [#48488]\r\n* Add `supports_context_parallel` to `PreTrainedModel` (#48442) by @qgallouedec in [#48442]\r\n* [docs] mlinter reference (#48460) by @stevhliu in [#48460]\r\n* [docs] Add a LiteRT page under community integrations (#48540) by @john-rocky in [#48540]\r\n* [GLM 5.3 Flash] Fix NaN gradients in chunked KDA (#48455) by @imvladikon in [#48455]\r\n* [docs] Fix [[autodoc]] directives in ALBERT model documentation (#48593) by @samyuktahegde in [#48593]\r\n* [Fix] Fix A10 expectations for a test (#48454) by @remi-or in [#48454]\r\n* Add PR comment CI for AMD (MI300) (#48065) by @ydshieh in [#48065]\r\n* docs: fix docstring parameter names that do not match signatures (#48575) by @simpleqt in [#48575]\r\n* docs: remove phantom parameters from docstrings (#48576) by @simpleqt in [#48576]\r\n* extend some case to xpu as well (#48502) by @sywangyi in [#48502]\r\n* [serge] Fix 2 integration tests for model `glm4_moe` failing with `OOM` (other (2)) (#48551) by @sergereview[bot] in [#48551]\r\n* [serge] Fix 2 integration tests for model `nemotron` failing with `import_or_config` (other (2)) (#48582) by @sergereview[bot] in [#48582]\r\n* Compress the agent conventions file and document two modular pitfalls (#48586) by @tarekziade in [#48586]\r\n* Guard against a None video processor class when the backend is unavailable (#48557) by @caiotheodoro in [#48557]\r\n* [serge] Fix 2 integration tests regressed by commit 83d46aa2a2c4 (PR #47625) (#48580) by @sergereview[bot] in [#48580]\r\n* [nit] use requires_backends (#47576) by @eustlb in [#47576]\r\n* [serge] Fix 2 integration tests for model `kosmos2` failing with `import_or_config` (other (2)) (#48552) by @sergereview[bot] in [#48552]\r\n* Fix failing tests for cohere_compass (#48005) by @kaixuanliu in [#48005]\r\n* [Fix] Use dedicated helpers for DeepGEMM and SonicMoE tests (#48523) by @remi-or in [#48523]\r\n* CI: Point test fixtures at hf-internal-testing copies we already host (#48521) by @tarekziade in [#48521]\r\n* fix failed test cases for glm5_next (#48497) by @kaixuanliu in [#48497]\r\n* Pass kwargs to the Mamba2 mixer in Nemotron-H, Falcon-H1 and Mamba2 (#48490) by @kfastino in [#48490]\r\n* Retire test_multi_gpu_data_parallel_forward (#48508) by @tarekziade in [#48508]\r\n* Processing tests [part 2] (#47922) by @zucchini-nlp in [#47922]\r\n* QA: Add noisy comment checker (#48484) by @tarekziade in [#48484]\r\n* Allow nested rope params for tiny models (#48435) by @zucchini-nlp in [#48435]\r\n* Fix sliding-window mask `layer_idx` in Gemma3/Gemma4 `create_masks_for_vision_model` (#48482) by @jiqing-feng in [#48482]\r\n* add xpu expectations for hunyuan_vl model tests (#48504) by @kaixuanliu in [#48504]\r\n* [`Qwen 3.5 Moe`] Fix decorators (#48436) by @vasqu in [#48436]\r\n* Infinite loop in dependency search (#48393) by @zucchini-nlp in [#48393]\r\n* [serge] Fix 2 integration tests for model `fsmt` failing with `output_mismatch` (tensor values differ (2)) (#48496) by @sergereview[bot] in [#48496]\r\n* Fix some tests by removing the deprecation cycle (#48503) by @Cyrilvallez in [#48503]\r\n* Remove deprecation (#48500) by @Cyrilvallez in [#48500]\r\n* Fix Pix2StructTextAttention init using hidden_size instead of d_kv (#47558) by @<NOT FOUND> in [#47558]\r\n* Fix pre patch release utility (#48499) by @Cyrilvallez in [#48499]\r\n* Update dev version (#48498) by @Cyrilvallez in [#48498]\r\n* Fix Inkling inputs_embeds and add more tests (#47827) by @Cyrilvallez in [#47827]\r\n* Simplify and fix qwen4 tests (#48340) by @Cyrilvallez in [#48340]\r\n* [docs] Partial checkpointing and group_by_length (#48463) by @stevhliu in [#48463]\r\n* doc: fix syntax error and typos in VibeVoice documentation (#48489) by @VimalN2005 in [#48489]\r\n* [`Qwen4 Exp`] Use partial to avoid skipping mask more easily (#48456) by @vasqu in [#48456]\r\n* Add support for NeuCodec (#47143) by @harryjulian in [#47143]\r\n* [serge] Fix 1 integration tests regressed by commit bd9509355c8a (PR #47493) (#48426) by @sergereview[bot] in [#48426]\r\n* Fix rotary embedding regression (#48477) by @Cyrilvallez in [#48477]\r\n* Remove deprecated mask functions (#48476) by @Cyrilvallez in [#48476]\r\n* [MTP] Save memory by only capturing the last layer's hidden_states (#48475) by @Cyrilvallez in [#48475]\r\n* Allow capturing only necessary hidden_states with capture_outputs (#48081) by @sywangyi in [#48081]\r\n* Support per-layer MTP configuration (#48264) by @eladsegal in [#48264]\r\n* [docs] Fix code snippets (#47772) by @stevhliu in [#47772]\r\n* [Fix] Sparse TikToken tokenizers silently fail (#48446) by @remi-or in [#48446]\r\n* No inherit decorator for NeoMME (#48457) by @zucchini-nlp in [#48457]\r\n* Batch Rebalance Data Sampler (#47340) by @delock in [#47340]\r\n* [serge] Fix 2 integration tests for model `cwm` failing with `import_or_config` (other (2)) (#48414) by @sergereview[bot] in [#48414]\r\n* fix: Add DEIMv2 attribution (#48448) by @harshaljanjani in [#48448]\r\n* [fix] inkling: mps + cuda mel spec extraction (#47432) by @eustlb in [#47432]\r\n* fix: decode() batch path respects self.clean_up_tokenization_spaces (#47793) by @lorenzozanee in [#47793]\r\n* Grounding dino fp16 dtype [backlog] (#48438) by @molbap in [#48438]\r\n* Init the process group with a load-scaled timeout for sharded loading (#48228) by @qgallouedec in [#48228]\r\n* Raise a clear error when a token is both forced and suppressed (#47511) by @qgallouedec in [#47511]\r\n* Clarify device placement in pipelines (#47367) by @LysandreJik in [#47367]\r\n* Add offload to gradient checkpointing (#48444) by @qgallouedec in [#48444]\r\n* [MiniCPMV4_6] Update test_small_model_vision_generation_batch expected output (value drift) (#48406) by @ydshieh in [#48406]\r\n* [serge] Fix 2 integration tests for model `hyperclovax` failing with `other` (other (2)) (#48440) by @sergereview[bot] in [#48440]\r\n* fix(models): Drop the position-indexed token type lookup in RoPE encoders (#48407) by @harshaljanjani in [#48407]\r\n* Document image_hidden_states/pixel_values mutual exclusivity for SmolVLM/Idefics2/Idefics3 (#47714) by @verma8076 in [#47714]\r\n* Re-order a bit for easier navigation (#48434) by @zucchini-nlp in [#48434]\r\n* Deprecated stuff gone (#48367) by @zucchini-nlp in [#48367]\r\n* [serge] Fix 6 integration tests for model `seamless_m4t_v2` failing with `other` (other (6)) (#48425) by @sergereview[bot] in [#48425]\r\n* [Docs]: Update GLM 5.3 (#48401) by @Dovis01 in [#48401]\r\n* Avoid print to stdout that fails the job `check_failed_tests` job (#48391) by @ydshieh in [#48391]\r\n* Fix incorrect tuple return annotations on forward methods returning a Tensor (#48359) by @Gronoxx in [#48359]\r\n* Fix interval merge invariant in _find_disjoint (#47860) by @sharmax-vikas in [#47860]\r\n* skip mtp slow tests for now (#48328) (#48329) by @tarekziade in [#48329]\r\n* Update Tailscale action version in workflow (#48394) by @glegendre01 in [#48394]\r\n* fix some failure in xpu (#48252) by @sywangyi in [#48252]\r\n* [Improvement] Make gated delta rule more explicit  (#47625) by @remi-or in [#47625]\r\n* [CB] Fix wrong device scoping (#48370) by @remi-or in [#48370]\r\n* Bump transformers-mlinter to 0.1.5 and clear the new findings (#48259) by @tarekziade in [#48259]\r\n* fix: flash-attn fallback failing on torch2.13 (#48388) by @NanoCode012 in [#48388]\r\n* [LongcatFlash] Fix test_longcat_generation_cpu: use device_map=\"cpu\" to avoid MoE disk offload issue (#48377) by @ydshieh in [#48377]\r\n* [Qwen3VLMoe] Update `test_small_model_integration_test_batch` expected output (value drift) (#48376) by @ydshieh in [#48376]\r\n* [ONNX] Skip affected models on torch 2.13 (two dynamo regressions) (#48191) by @ydshieh in [#48191]\r\n* Fix `safe_open` mmap memory exhaustion on Windows by using `pread` backend (#48341) by @eryk-roch in [#48341]\r\n* Fix Zamba2 construction for num_mem_blocks > 1 checkpoints (#48325) by @john-rocky in [#48325]\r\n* [Docs] Change 5.3 Flash pos in toc (#48366) by @Dovis01 in [#48366]\r\n* [conftest] Use get_cpu_ram_total_gib for psutil patch (cgroup-aware) (#48290) by @ydshieh in [#48290]\r\n* Fix flaky test_training_gradient_checkpointing for BigBirdPegasus (fp noise filter) (#48332) by @ydshieh in [#48332]\r\n* Quiet continuous batching at default verbosity (#48314) by @qgallouedec in [#48314]\r\n* Wait for the first request in the async continuous batching bootstrap (#48304) by @qgallouedec in [#48304]\r\n* Ignore a stale best checkpoint recorded in a resumed trainer state (#48319) by @VaggelisGian in [#48319]\r\n* [docs] Fix links and remove TokenizerFast (#47748) by @stevhliu in [#47748]\r\n* Fix incorrect token classification prefix for ESMC (#48348) by @Rocketknight1 in [#48348]\r\n* Create the continuous batching CPU group with local synchronization (#48302) by @qgallouedec in [#48302]\r\n* Resolve continuous batching config against the text config for composite models (#48299) by @qgallouedec in [#48299]\r\n* [`CI`] Unblock fast CI for now (failing tests) (#48344) by @vasqu in [#48344]\r\n* [qwen4_exp] disable torch/onnx export tests due to data-dependent control flow (#48345) by @ydshieh in [#48345]\r\n* [debug] Trace previous CI run selection in get_previous_daily_ci.py (#48338) by @ydshieh in [#48338]\r\n* Normalize HunYuanVL's legacy field aliases via attribute_map (#48261) by @hmellor in [#48261]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @ydshieh\r\n    * [fix] Update stale expected strings in HunYuanVL integration tests (#48646)\r\n    * [fix] Update stale golden values and fix expected_logits shape in FlavaForPreTraining integration tests (#48639)\r\n    * [tests] Fix integration test golden values broken by fast image processor default (PR #41388) (#48637)\r\n    * [KimiLinear] Fix test_cpu_offload: set num_local_experts=4 in model tester (#48624)\r\n    * Fix GPU memory teardown in CLI serve tests (#48618)\r\n    * [fix] Fix how we read package versions - triggered by torch 2.14+ (#48615)\r\n    * [Docker] Upgrade CPU torch to <=2.14.0, torchcodec to <=0.16.0 (#48614)\r\n    * Add PR comment CI for AMD (MI300) (#48065)\r\n    * [MiniCPMV4_6] Update test_small_model_vision_generation_batch expected output (value drift) (#48406)\r\n    * Avoid print to stdout that fails the job `check_failed_tests` job (#48391)\r\n    * [VibeVoice] Skip generate export tests (flaky) (#48396)\r\n    * [LongcatFlash] Fix test_longcat_generation_cpu: use device_map=\"cpu\" to avoid MoE disk offload issue (#48377)\r\n    * [Qwen3VLMoe] Update `test_small_model_integration_test_batch` expected output (value drift) (#48376)\r\n    * [ONNX] Skip affected models on torch 2.13 (two dynamo regressions) (#48191)\r\n    * Retry get_daily_ci_runs on stale GitHub API cache (#48374)\r\n    * [conftest] Use get_cpu_ram_total_gib for psutil patch (cgroup-aware) (#48290)\r\n    * Fix flaky test_training_gradient_checkpointing for BigBirdPegasus (fp noise filter) (#48332)\r\n    * [qwen4_exp] disable torch/onnx export tests due to data-dependent control flow (#48345)\r\n    * [debug] Trace previous CI run selection in get_previous_daily_ci.py (#48338)\r\n* @LauraGPT\r\n    * Add Fun-ASR-Nano model (#46180)\r\n* @tarekziade\r\n    * Fix AttributeError in gradient_checkpointing_enable(offload=True) (#48590)\r\n    * Compress the agent conventions file and document two modular pitfalls (#48586)\r\n    * CI: Point test fixtures at hf-internal-testing copies we already host (#48521)\r\n    * Retire test_multi_gpu_data_parallel_forward (#48508)\r\n    * QA: Add noisy comment checker (#48484)\r\n    * skip mtp slow tests for now (#48328) (#48329)\r\n    * Bump transformers-mlinter to 0.1.5 and clear the new findings (#48259)\r\n* @remi-or\r\n    * [Fix] Fix A10 expectations for a test (#48454)\r\n    * Kimi linear (#48250)\r\n    * [Fix] Use dedicated helpers for DeepGEMM and SonicMoE tests (#48523)\r\n    * [Fix] Sparse TikToken tokenizers silently fail (#48446)\r\n    * [Improvement] Make gated delta rule more explicit  (#47625)\r\n    * [CB] Fix wrong device scoping (#48370)\r\n    * [CB] Fail faster (#48334)\r\n* @ArthurZucker\r\n    * Add h4 (#48473)\r\n* @harryjulian\r\n    * Add support for NeuCodec (#47143)\r\n* @delock\r\n    * Batch Rebalance Data Sampler (#47340)\r\n* @harshaljanjani\r\n    * fix: Add DEIMv2 attribution (#48448)\r\n    * model: Add NVIDIA Canary-1B-v2 to Transformers (#46825)\r\n    * fix(generation): Enforce the auto-compile cache check for encoder-decoder models (#48364)\r\n    * fix(models): Drop the position-indexed token type lookup in RoPE encoders (#48407)\r\n* @tonywu71\r\n    * Add NeoMME and NeoMME-Retriever (#47992)\r\n* @Dovis01\r\n    * [Docs]: Update GLM 5.3 (#48401)\r\n    * [Docs] Change 5.3 Flash pos in toc (#48366)\r\n    * [Glm 5.3 Flash] GLM 5.3 Flash Support (#48342)\r\n* @pengzhiliang\r\n    * Implement VibeVoice  (#40546)","publishedAt":"2026-09-09T15:42:45.000Z","fetchedAt":"2026-09-09T16:05:50.757Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.17.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9","r2Key":"releases/6aaaa7be263758a4d02d2475d0621dadefc1ad2e9336fe49126ec6efc82050bd.png","r2Url":"https://media.releases.sh/releases/6aaaa7be263758a4d02d2475d0621dadefc1ad2e9336fe49126ec6efc82050bd.png"},{"type":"image","url":"https://github.com/user-attachments/assets/29ccea01-a855-4d4f-a6af-61bc4fc883a4","r2Key":"releases/d108b6917b70f388880ab653d8f967ff24fd89c6218d6765592930a0eff14073.png","r2Url":"https://media.releases.sh/releases/d108b6917b70f388880ab653d8f967ff24fd89c6218d6765592930a0eff14073.png"}],"coverageCount":0},{"id":"rel_dBHC58_xnednKbxy_L87i","version":"v5.16.1","type":"feature","title":"Release v5.16.1","summary":"Adds support for GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. Restores backward compatibility for the tensor-parallel API and pins an ESMFold2 kernel commit for security reasons.","titleGenerated":"Hugging Face Transformers v5.16.1 adds GLM-5.3-Flash support","titleShort":"GLM-5.3-Flash arrives; TP BC restored","breaking":"none","importance":4,"content":"# Release v5.16.1\r\n\r\nThis is a special release as we include GLM! (and a few small fixes)\r\n\r\n# GLM-5.3-Flash\r\n\r\n<img width=\"4239\" height=\"2643\" alt=\"image\" src=\"https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209\" />\r\n\r\nGLM-5.3-Flash, the first **natively multimodal model** in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.\r\n\r\nGLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest **30T-token** multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/glm5_next)\r\n* [Glm 5.3 Flash] GLM 5.3 Flash Support (#48342) by @Dovis01 in [#48342](https://github.com/huggingface/transformers/pull/48342)\r\n\r\n\r\n## Small patch fixes\r\n\r\nMainly BC behavior for TP and pinning a hf kernel for security reasons :hugs: \r\n\r\n- Restore BC for the tensor-parallel API (#48300) by @ArthurZucker \r\n- Fix kernel commit and repo paths for ESMFold2 (#48186) by @Rocketknight1 \r\n\r\n**Full Changelog**: https://github.com/huggingface/transformers/compare/v5.16.0...v5.16.1","publishedAt":"2026-08-26T14:50:01.000Z","fetchedAt":"2026-08-26T16:40:47.617Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.16.1","media":[{"type":"image","url":"https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209","r2Key":"releases/db036979bc58c94dbeaa0b11c6215e32154fbe6e1c5e39ddba7662f50b5b341a.png","r2Url":"https://media.releases.sh/releases/db036979bc58c94dbeaa0b11c6215e32154fbe6e1c5e39ddba7662f50b5b341a.png"}],"coverageCount":0},{"id":"rel_lbzs8Z4GN0uS_B0W6mWWy","version":"v5.16.0","type":"feature","title":"Release: v5.16.0","summary":"This release adds four new models: Qwen4-Exp, GraniteSpeech5, Step3.7-Flash, and CohereCompass. It also replaces the legacy tensor-parallel implementation with a DTensor-native backend, a breaking change for existing TP users, and fixes several cache and generation bugs.","titleGenerated":"Transformers v5.16.0 adds Qwen4-Exp, GraniteSpeech5, and other models","titleShort":"Qwen4-Exp, GraniteSpeech5, and more new models land","breaking":"major","importance":4,"content":"# Release v5.16.0\r\n\r\n\r\n## New Model additions\r\n\r\n### Qwen4-Exp\r\n\r\n<img width=\"2241\" height=\"693\" alt=\"image\" src=\"https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671\" />\r\n\r\nQwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).\r\n\r\nGR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream.\r\n\r\nQSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed. This block-level selection reduces indexing overhead and improves memory locality for long sequences. Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads.\r\n\r\nPLE enriches selected decoder layers with layer-specific lexical features derived from hashed token n-grams and a dilated depthwise convolution.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/qwen4_exp)\r\n* Add Qwen4Exp model (#48337) by @Cyrilvallez in [#48337](https://github.com/huggingface/transformers/pull/48337)\r\n\r\n### GraniteSpeech5\r\n\r\n<img width=\"1600\" height=\"1440\" alt=\"image\" src=\"https://github.com/user-attachments/assets/106ef712-9f45-43c9-98b7-ce6e4a7c136d\" />\r\n\r\nGranite Speech 5.0 Turbo CTC is a lightweight (~470M parameters) conformer encoder for automatic speech recognition, trained with Connectionist Temporal Classification (CTC) on BPE targets. It is a fast, encoder-only member of the [Granite Speech](https://huggingface.co/papers/2505.08699) family: transcription requires a single forward pass followed by greedy CTC decoding, with no autoregressive decoder.\r\n\r\nArchitecturally, it extends the Granite Speech conformer CTC encoder with:\r\n\r\n1. **Frame stacking + block-wise time subsampling**: the feature extractor stacks pairs of log-mel(+delta) frames (2x), and the first two conformer blocks each subsample time by 2 through a stride-2 depthwise convolution (with a mean-pooled residual), for a total 8x time reduction at 10 ms mel hop.\r\n\r\n2. **Block attention with Shaw's relative positional embeddings**: attention is computed over fixed-size blocks (the sequence is right-padded to a whole number of blocks, with padded frames masked out), using separate bias-free query/key/value projections.\r\n\r\n3. **Self-conditioned CTC**: the CTC posteriors of the middle layer are projected and fed back into the hidden states, and the CTC head is shared between this mid-layer self-conditioning and the final prediction.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/granite_speech5)\r\n* Add Granite Speech 5.0 - (#48288) by @eustlb in [#48288](https://github.com/huggingface/transformers/pull/48288)\r\n\r\n### Step3p7\r\n\r\nStep-3.7-Flash was proposed in [Step 3.7 Flash](https://static.stepfun.com/blog/step-3.7-flash/) by StepFun. It is a 198B-parameter sparse Mixture-of-Experts vision-language model, pairing a 196B-parameter MoE language backbone with a 1.8B-parameter vision encoder for native image understanding.\r\n\r\nStepFun hasn't published a technical report for Step-3.7-Flash, so the details below are drawn from the released checkpoint's configuration rather than a paper.\r\n\r\n- **Sparse MoE decoder**: all but the first 3 decoder layers route through a MoE block of 288 routed experts (top-8 per token) plus a single shared expert. The router scores experts with a sigmoid and a learned per-expert bias instead of an auxiliary load-balancing loss, the same strategy as [DeepSeek-V3](./deepseek_v3).\r\n- **Gated attention**: each attention layer adds an extra projection whose sigmoid output gates the attention output per head, before the output projection — the same *Gated Attention* mechanism used in [Qwen3-Next](./qwen3_next). A subset of layers use fewer heads and a sliding window instead of full attention.\r\n- **Multi-token prediction**: some checkpoints ship extra decoder layers trained for multi-token prediction, which [`~GenerationMixin.generate`] can use for speculative decoding via `use_mtp=True`.\r\n- **Vision encoder**: a SigLIP-style ViT with 2-D rotary position embeddings and a learned per-layer scale on the attention and MLP branches. Its output is downsampled 4x by two stride-2 convolutions before a linear projector maps it into the text model's hidden size.\r\n- **Dynamic image tiling**: instead of a fixed tile grid, the image processor picks its tiling window from each image's own aspect ratio, producing one downscaled global view plus zero or more local high-resolution crops per image.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/step3p7)\r\n* [new model] step 3.7 (#46658) by @itazap in [#46658](https://github.com/huggingface/transformers/pull/46658)\r\n\r\n### CohereCompass\r\n\r\nCohereCompass is the base architecture for small, specialized (vision-)language models trained by Cohere.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/cohere_compass)\r\n* Add CohereCompass modeling (#47878) by @calpt in [#47878](https://github.com/huggingface/transformers/pull/47878)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nThe legacy tensor-parallel implementation has been replaced with a DTensor-native backend, so users relying on the previous TP API for inference or training must migrate to the new DTensor-based interface.\r\n* 🚨 TP dtensor API inference + training (#47579) by @3outeille\r\n\r\n`attn_implementation=\"sdpa\"` dispatch is now properly supported for wav2vec2_conformer, wav2vec2-bert, and SeamlessM4T/v2 models, which may change initialization behavior for users who previously worked around this limitation.\r\n* 🚨[wav2vec2] Support attn_implementation=sdpa dispatch (#46196) by @YangKai0616\r\n\r\n`FuyuProcessor` no longer returns the `image_patch_indices` output, so any code that depends on this field must be updated to remove references to it.\r\n* :rotating_light: Leftover processors (#47924) by @zucchini-nlp\r\n\r\n\r\n\r\n## Cache\r\n\r\nSeveral cache-related bugs were fixed in this release, including an off-by-one error in the sliding window cache, Whisper speculative decoding cache corruption, CpmAnt use-cache failures, Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches, and compressed-tensors loading for KV-cache-only quantized models. Documentation was also added for cache token removal using negative values, and per-layer cache configuration support (allowing models to use different cache settings per layer) was introduced.\r\n\r\n\r\n* Cpmant fix use cache (#48013) by @jiqing-feng in [#48013]\r\n* [docs] Cache crop (#47950) by @stevhliu in [#47950]\r\n* Revert \"Support per-layer cache configuration and attention-mask selection\" (#48175) by @Cyrilvallez in [#48175]\r\n* Support per-layer cache configuration and attention-mask selection (#47901) by @eladsegal in [#47901]\r\n* Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache (#47872) by @jiqing-feng in [#47872]\r\n* Fix sliding window cache index off-by-one on wraparound (#47708) by @hameedibrh in [#47708]\r\n* [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression (#48000) by @ydshieh in [#48000]\r\n* Fix compressed-tensors loading for KV-cache-only quantized models (#47904) by @kylesayrs in [#47904]\r\n\r\n\r\n## Generation\r\n\r\nThis release fixes several generation bugs across multiple models, including Whisper speculative decoding issues (UnboundLocalError, cache corruption, speed regression, and left-padded batch position IDs), broken image generation in Emu3, garbage output in OLMo/GPTNeoX, and Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches. Additionally, logit distributions for candidate generators using sampling are now aligned by returning logits after applying logit processors.\r\n\r\n\r\n* [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() (#48108) by @ydshieh in [#48108]\r\n* [Whisper] Fix decoder position IDs for left-padded batches in longform generation (#48028) by @ydshieh in [#48028]\r\n* [serge] Fix 2 integration tests for model `generation` failing with `import_or_config` (other (2)) (#48061) by @sergereview[bot] in [#48061]\r\n* Align logit distributions for CandidateGenerators using sampling (#48007) by @Cyrilvallez in [#48007]\r\n* [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) (#47988) by @ydshieh in [#47988]\r\n* [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 (#47948) by @ydshieh in [#47948]\r\n\r\n\r\n## Attention\r\n\r\nSeveral attention-related bug fixes were made in this release, including correcting a SigLIP2 documentation typo, fixing Flash/SDPA attention dispatch tests for xcodec2 and ROCm RDNA GPUs, resolving a GPT2 cross-attention mask being silently discarded, and enabling SDPA support declaration in `TimmWrapper`. Per-layer cache configuration and attention-mask selection support was also introduced, allowing models with heterogeneous layer configurations to use distinct `sliding_window`, `attention_chunk_size`, and `number_of_conv_states` values per layer.\r\n\r\n\r\n*  doc: Fix typo in SigLIP2 Flash Attention code example (#48197) by @VimalN2005 in [#48197]\r\n* [xcodec2] Fix flex attention and flash dispatch tests (#48244) by @jiqing-feng in [#48244]\r\n* Fix ROCm SDPA-flash skip guard that crashes on RDNA GPUs (#47965) by @Abdennacer-Badaoui in [#47965]\r\n* [GPT2] Fix encoder_attention_mask being silently discarded in cross-attention (#47946) by @DavidJohnQuinlan in [#47946]\r\n* Declare sdpa support in `TimmWrapper` (#47939) by @jiqing-feng in [#47939]\r\n\r\n\r\n## Quantization\r\n\r\nQuantization improvements include adding NVFP4 quantization support via HF kernels (enabling on-the-fly BF16 weight quantization with ~50% memory reduction), and fixing several bugs: reverting a regression in `is_quantization_compressed` that caused incorrect module layouts for packed-format checkpoints, fixing CLIP weight initialization failures with quantized checkpoints, and restoring KV-cache quantization setup for KV-cache-only quantized models.\r\n\r\n\r\n* Revert \"[Quantization]: Refactor is_quantization_compressed for format-based detection\" (#48072) by @subin9 in [#48072]\r\n* feat: add nvfp4 quantization (#47883) by @drbh in [#47883]\r\n* [DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization (#47991) by @ydshieh in [#47991]\r\n* Fix CLIP _init_weights when a child module carries quantized weights (#47921) by @Bluear7878 in [#47921]\r\n\r\n\r\n## Parallelization\r\n\r\nIntroduced a naive pipeline parallel inference engine supporting tied/untied weight embeddings with seamless `generate()` integration, while restoring backward compatibility for the tensor-parallel API with a deprecation cycle for `tp_plan` in `from_pretrained()`. Additionally fixed a model parallel bug in the BLT model affecting beam search.\r\n\r\n\r\n* Restore BC for the tensor-parallel API (#48300) by @ArthurZucker in [#48300]\r\n* fix bug for blt model parallel bug (#48327) by @kaixuanliu in [#48327]\r\n* Pipeline parallel naive inference (#47289) by @3outeille in [#47289]\r\n\r\n\r\n## Kernels\r\n\r\nKernel support was improved with documentation updates highlighting supported models, a fix for export crashes on kernel-decorated functions by adding a `is_torchdynamo_exporting` guard, and the default Flash Attention 2 hub kernel version was bumped to v3 to resolve compatibility issues with newer PyTorch versions.\r\n\r\n\r\n* [docs] Kernel supported models (#48258) by @stevhliu in [#48258]\r\n* [Fix] Export crashes on kernel-decorated function (#47808) by @remi-or in [#47808]\r\n* Bump default flash-attn2 hub kernel version to v3 (#47863) by @jiqing-feng in [#47863]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Fix video-llama modular conversion (#48336) by @zucchini-nlp in [#48336]\r\n* CI: gate the hunyuan-moe slow test  (#48330) by @tarekziade in [#48330]\r\n*  Add a regression test for force_accelerate_hooks signature preservation  (#48260) by @wtdcode in [#48260]\r\n* Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor (#48071) by @dkrisman in [#48071]\r\n* Fix `scores` type in stopping criteria docstrings (#47676) by @qgallouedec in [#47676]\r\n* Docstring check didn't match some file - fix it (#48121) by @zucchini-nlp in [#48121]\r\n* Add shared ImageProcessingTester (#47745) by @guarin in [#47745]\r\n* Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models (#48211) by @ydshieh in [#48211]\r\n* [docs] Fix failing doctests (#47687) by @stevhliu in [#47687]\r\n* Fix build_2d_sinusoidal_position_embedding on MPS (#47897) by @guarin in [#47897]\r\n* [Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager (#48236) by @ydshieh in [#48236]\r\n* Let gradient checkpointing skip layers with every_n_layers (#48200) by @qgallouedec in [#48200]\r\n* Disable daily nightly CI (#48292) by @remi-or in [#48292]\r\n* gs (#48288) by @eustlb in [#48288]\r\n* CI: fix muse OOMs (#48284) by @tarekziade in [#48284]\r\n* Fix `BayesianDetectorModel.from_pretrained()` by calling `post_init()` (#48254) by @woojinpaik in [#48254]\r\n* [Fix] Avoid duplicating tests in CI (#48287) by @remi-or in [#48287]\r\n* [`GDN`] Fix recurrent FLA fallback (#48266) by @vasqu in [#48266]\r\n* Fix `tie_word_embeddings` not lifted from `text_config` for some VLM configs (BC regression) (#45857) by @qgallouedec in [#45857]\r\n* replace xpu-smi subprocess call in benchmark_v2 (#48083) by @kaixuanliu in [#48083]\r\n* ignore mlinter ci file (#48267) by @tarekziade in [#48267]\r\n* Fix dtype mismatch in grouped_mm_fallback for LoRA training on Mamba+… (#47933) by @adh-aakriti in [#47933]\r\n* Compute MoE load-balancing loss per layer to avoid giant one-hot materialization −99.7% @ 128k (#48131) by @qgallouedec in [#48131]\r\n* deterministic layer_types buffer registration in multiple models (#48162) by @mowoe in [#48162]\r\n* Fix nemotron_h save_pretrained emitting singular backbone.embedding.weight (#48075) by @yuekaizhang in [#48075]\r\n* `force_accelerate_hooks` should not hide the signature it wraps (#48156) by @SunMarc in [#48156]\r\n* fix(data_collator): align TokenClassification numpy_call with torch_call (#48212) by @<NOT FOUND> in [#48212]\r\n* Fix a typo in a use of a local variable field_ in a test (#48184) by @AleksMat in [#48184]\r\n* [serge] Fix 2 integration tests regressed by commit 16780c86b20a (PR #47622) (#48134) by @sergereview[bot] in [#48134]\r\n* [Gemma4] Fix stale expected values in integration tests (#48233) by @ydshieh in [#48233]\r\n* [TableTransformer, PI0] Fix stale expected values and OOM in integration tests (#48198) by @ydshieh in [#48198]\r\n* Post two CI badges on a PR: CPU PR CI and GPU run-slow (#48190) by @tarekziade in [#48190]\r\n* Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) (#48171) by @ydshieh in [#48171]\r\n* Assign a reviewer even when a codeowner has left, and route models by modality (#48085) by @tarekziade in [#48085]\r\n* [VITS] Un-skip test_model_forward (#46375) by @blipbyte in [#46375]\r\n* Apply context parallelism to the evaluation path (#48167) by @qgallouedec in [#48167]\r\n* Fix `gpt_oss` runs on GPU (#48118) by @tarekziade in [#48118]\r\n* [EsmFold2] Fix stale expected distogram logit values (#48182) by @ydshieh in [#48182]\r\n* Fix Apr 05 integration test regressions (cuda sm_86) (#48170) by @ydshieh in [#48170]\r\n* Add MLU support to is_flash_linear_attention_available (#46995) by @atri2549 in [#46995]\r\n* Fix integration test expected values for cuda sm_86 (Mar 15 regressions) (#48168) by @ydshieh in [#48168]\r\n* Fix DynamicCache reconstruction during ExecuTorch export (#47900) by @eladsegal in [#47900]\r\n* [LLaVA] Fix pixtral integration tests for cuda sm_86 (#48166) by @ydshieh in [#48166]\r\n* Port ESMC and ESMFold2 to Transformers (#46419) by @Rocketknight1 in [#46419]\r\n* [Qwen2.5-Omni] Update stale expected values for cuda sm_86 (#48164) by @ydshieh in [#48164]\r\n* [Mistral3] Fix batched integration tests: padding_side=left + update expected values (#48161) by @ydshieh in [#48161]\r\n* Fix DeepSeek V2 default vocab size (#48159) by @hmellor in [#48159]\r\n* [InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) (#48153) by @ydshieh in [#48153]\r\n* Retry transient network errors (RemoteDisconnected) in github_utils (#48124) by @ydshieh in [#48124]\r\n* fix bugs for clvp model (#47127) by @kaixuanliu in [#47127]\r\n* Always tie embeddings for LongT5 and Pop2Piano (#47620) by @jiqing-feng in [#47620]\r\n* Use generator with seed for LengthGroupedSampler in Trainer._get_eval_sampler for deterministic eval order with per_device_eval_batch_size > 1 (#48025) by @philipshurpik in [#48025]\r\n* Enable mlinter findings artifact for inline PR reviews (#48117) by @ydshieh in [#48117]\r\n* Fix CpmAnt loading: size lm_head to vocab_size (#48012) by @jiqing-feng in [#48012]\r\n* Delete old mlinter review comments before posting new ones (#48107) by @ydshieh in [#48107]\r\n* [Video] Warn and return all frames when num_frames exceeds total_num_frames (#48074) by @carlszk in [#48074]\r\n* Accept artifact dir as argument in post_mlinter_review.py (#48106) by @ydshieh in [#48106]\r\n* Fix mlinter artifact path (#48088) by @ydshieh in [#48088]\r\n* [Video] Fix convert_to_rgb channel slicing and alpha blending for RGBA videos (#48053) by @<NOT FOUND> in [#48053]\r\n* [PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models (#48064) by @ydshieh in [#48064]\r\n* fix(pipeline): preserve model.generation_config precedence over pipeline defaults (#47752) (#47953) by @nithin42 in [#47953]\r\n* [Gemma3] Update integration test expected values for A10G (#48036) by @ydshieh in [#48036]\r\n* Fallback from 'lanczos' to 'bicubic' when on cuda (#48026) by @zucchini-nlp in [#48026]\r\n* [serge] Fix 2 integration tests regressed by commit b9090ae58cda (PR #47096) (#48060) by @sergereview[bot] in [#48060]\r\n* [Fix] Small FA-related test failures in CB (#47341) by @remi-or in [#47341]\r\n* fix: honor empty processor_kwargs={} in multimodal pipelines (#48044) by @<NOT FOUND> in [#48044]\r\n* [CircleCI] Enable CI for private forks, no-op for public repo (#48056) by @ydshieh in [#48056]\r\n* Support `BatchFeature` in length-grouped samplers (#48034) by @qgallouedec in [#48034]\r\n* Fix EOS for candidate generators (#47931) by @<NOT FOUND> in [#47931]\r\n* [Gemma3n] Update integration test expected values for A10G + torch 2.13 (#48035) by @ydshieh in [#48035]\r\n* fix failed test cases for muse_glimmer (#48011) by @kaixuanliu in [#48011]\r\n* Cohere compass tests (#47895) by @zucchini-nlp in [#47895]\r\n* docs: use relative paths for README language menus and add fa/ro entries (#47777) by @Priyans-Lathiya in [#47777]\r\n* Moving mlinter to 0.1.4 (#47918) by @tarekziade in [#47918]\r\n* [MoE] Fix Blackwell GPU crash with torch._grouped_mm on torch <= 2.8 (#48014) by @<NOT FOUND> in [#48014]\r\n* [docs] Muse Glimmer (#47882) by @stevhliu in [#47882]\r\n* [Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype (#48031) by @ydshieh in [#48031]\r\n* Proper separation of tests (#47943) by @zucchini-nlp in [#47943]\r\n* unpin `pytest` in the `examples_torch` deps (#48023) by @tarekziade in [#48023]\r\n* [CI] Fix startup failure in pr_build_doc_with_comment workflow by adding missing get-pr-number dependency (#47971) by @<NOT FOUND> in [#47971]\r\n* Let the GPU verify caller turn on the memory probe (#48001) by @tarekziade in [#48001]\r\n* Fix MTP config when mlp_layer_types is absent (#48015) by @Cyrilvallez in [#48015]\r\n* Remove duplicate block_sparse_moe assignment in GraniteMoeDecoderLayer (#47876) by @Aman2394 in [#47876]\r\n* [ModernVBERT] Fix integration test checkpoint (404 since April) (#48009) by @ydshieh in [#48009]\r\n* Fix cropping (#48006) by @Cyrilvallez in [#48006]\r\n* Fix DFlash candidate token device mismatch with device_map=\"auto\" (#47877) by @sywangyi in [#47877]\r\n* :red_circle: Allow tokenizers 0.23.1 (#46381) by @ArthurZucker in [#46381]\r\n* [Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() (#47997) by @ydshieh in [#47997]\r\n* [OLMoE] Update expected logits for A10G and add torch.no_grad() (#47989) by @ydshieh in [#47989]\r\n* [OLMo] Fix OOM in logits tests by adding torch.no_grad() (#47986) by @ydshieh in [#47986]\r\n* [AXK1] Fix expected logits for CUDA A10G (#47980) by @ydshieh in [#47980]\r\n* [Gemma] Update expected values for A10G (#47976) by @ydshieh in [#47976]\r\n* Fix gemma4 video to device (#47896) by @guarin in [#47896]\r\n* Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test (#47972) by @ydshieh in [#47972]\r\n* Potential fix for code scanning alert no. 267: Artifact poisoning (#47949) by @tarekziade in [#47949]\r\n* Fix GatedDeltaNet A_log dtype to prevent -inf under bfloat16 init (#47944) by @Nkluge-correa in [#47944]\r\n* [serge] Fix 2 integration tests for model `got_ocr2` failing with `other` (other (2)) (#47937) by @sergereview[bot] in [#47937]\r\n* Fix Jinja block endings in CHAT WITH MODELS' Writing a chat template … (#47960) by @ak1for2business-prog in [#47960]\r\n* [tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA (#47968) by @ydshieh in [#47968]\r\n* [serge] Fix 2 integration tests for model `opt` failing with `other` (other (2)) (#47909) by @sergereview[bot] in [#47909]\r\n* Scan a diff in trufflehog, not the whole repo history (#47945) by @tarekziade in [#47945]\r\n* [serge] Fix 2 integration tests for model `vivit` failing with `output_mismatch` (tensor values differ (2)) (#47566) by @sergereview[bot] in [#47566]\r\n* Fix Gemma `sliding_window` being halved on every config save/reload (#47940) by @Bluear7878 in [#47940]\r\n* docs: fix incorrect PEFT anchor link in fine-tuning section (#47927) by @dsulot in [#47927]\r\n* CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners (#47934) by @ydshieh in [#47934]\r\n* Make muse glimmer exportable (#47871) by @IlyasMoutawwakil in [#47871]\r\n* docs: add installation instructions for NVIDIA Spark (ARM64) devices (#47906) by @mfuntowicz in [#47906]\r\n* [CohereCompass] Minor docs fixes (#47903) by @calpt in [#47903]\r\n* fix: correct checkpoints, config annotations, and create_dummy_models improvements (#47902) by @ydshieh in [#47902]\r\n* Transform paths and repeat joining for response parsing (#47648) by @Rocketknight1 in [#47648]\r\n* [docs] Update toctree (#47781) by @stevhliu in [#47781]\r\n* Add `CI_CPU_MEMORY_LIMIT_GB` to check_failed_tests workflow (#47884) by @ydshieh in [#47884]\r\n* Use tiny Hub checkpoint in Qwen3ASR processor test (#47833) by @ydshieh in [#47833]\r\n* Update AutoRound XPU/CPU backend (#47826) by @yiliu30 in [#47826]\r\n* docs(tests): fix typos in test comments (#47859) by @zhaoxinyi02 in [#47859]\r\n* Update version post release (#47870) by @Cyrilvallez in [#47870]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @tarekziade\r\n    * CI: gate the hunyuan-moe slow test  (#48330)\r\n    * CI: fix muse OOMs (#48284)\r\n    * ignore mlinter ci file (#48267)\r\n    * Post two CI badges on a PR: CPU PR CI and GPU run-slow (#48190)\r\n    * Assign a reviewer even when a codeowner has left, and route models by modality (#48085)\r\n    * Fix `gpt_oss` runs on GPU (#48118)\r\n    * Moving mlinter to 0.1.4 (#47918)\r\n    * unpin `pytest` in the `examples_torch` deps (#48023)\r\n    * Let the GPU verify caller turn on the memory probe (#48001)\r\n    * Potential fix for code scanning alert no. 267: Artifact poisoning (#47949)\r\n    * Scan a diff in trufflehog, not the whole repo history (#47945)\r\n* @dkrisman\r\n    * Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor (#48071)\r\n* @jiqing-feng\r\n    * Cpmant fix use cache (#48013)\r\n    * [xcodec2] Fix flex attention and flash dispatch tests (#48244)\r\n    * Always tie embeddings for LongT5 and Pop2Piano (#47620)\r\n    * Fix CpmAnt loading: size lm_head to vocab_size (#48012)\r\n    * Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache (#47872)\r\n    * Bump default flash-attn2 hub kernel version to v3 (#47863)\r\n    * Declare sdpa support in `TimmWrapper` (#47939)\r\n* @ydshieh\r\n    * Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models (#48211)\r\n    * [Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager (#48236)\r\n    * [Gemma4] Fix stale expected values in integration tests (#48233)\r\n    * [TableTransformer, PI0] Fix stale expected values and OOM in integration tests (#48198)\r\n    * Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) (#48171)\r\n    * [EsmFold2] Fix stale expected distogram logit values (#48182)\r\n    * Fix Apr 05 integration test regressions (cuda sm_86) (#48170)\r\n    * Fix integration test expected values for cuda sm_86 (Mar 15 regressions) (#48168)\r\n    * [LLaVA] Fix pixtral integration tests for cuda sm_86 (#48166)\r\n    * [Qwen2.5-Omni] Update stale expected values for cuda sm_86 (#48164)\r\n    * [Mistral3] Fix batched integration tests: padding_side=left + update expected values (#48161)\r\n    * [InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) (#48153)\r\n    * Retry transient network errors (RemoteDisconnected) in github_utils (#48124)\r\n    * [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() (#48108)\r\n    * Enable mlinter findings artifact for inline PR reviews (#48117)\r\n    * Delete old mlinter review comments before posting new ones (#48107)\r\n    * Accept artifact dir as argument in post_mlinter_review.py (#48106)\r\n    * Fix mlinter artifact path (#48088)\r\n    * [Whisper] Fix decoder position IDs for left-padded batches in longform generation (#48028)\r\n    * [PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models (#48064)\r\n    * [Gemma3] Update integration test expected values for A10G (#48036)\r\n    * [CircleCI] Enable CI for private forks, no-op for public repo (#48056)\r\n    * [Gemma3n] Update integration test expected values for A10G + torch 2.13 (#48035)\r\n    * [Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype (#48031)\r\n    * [ModernVBERT] Fix integration test checkpoint (404 since April) (#48009)\r\n    * [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression (#48000)\r\n    * [Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() (#47997)\r\n    * [Whisper] Fix integration test failures on A10G (dtype, stale values, API changes) (#47995)\r\n    * [DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization (#47991)\r\n    * [OLMoE] Update expected logits for A10G and add torch.no_grad() (#47989)\r\n    * [OLMo] Fix OOM in logits tests by adding torch.no_grad() (#47986)\r\n    * [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) (#47988)\r\n    * [AXK1] Fix expected logits for CUDA A10G (#47980)\r\n    * [Gemma] Update expected values for A10G (#47976)\r\n    * [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 (#47948)\r\n    * Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test (#47972)\r\n    * [tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA (#47968)\r\n    * CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners (#47934)\r\n    * fix: correct checkpoints, config annotations, and create_dummy_models improvements (#47902)\r\n    * Add `CI_CPU_MEMORY_LIMIT_GB` to check_failed_tests workflow (#47884)\r\n    * Use tiny Hub checkpoint in Qwen3ASR processor test (#47833)\r\n* @eustlb\r\n    * gs (#48288)\r\n* @eladsegal\r\n    * Fix DynamicCache reconstruction during ExecuTorch export (#47900)\r\n    * Support per-layer cache configuration and attention-mask selection (#47901)\r\n* @YangKai0616\r\n    * 🚨[wav2vec2] Support attn_implementation=sdpa dispatch (#46196)\r\n* @drbh\r\n    * feat: add nvfp4 quantization (#47883)\r\n* @Priyans-Lathiya\r\n    * docs: use relative paths for README language menus and add fa/ro entries (#47777)\r\n* @itazap\r\n    * [new model] step 3.7 (#46658)\r\n* @calpt\r\n    * [CohereCompass] Minor docs fixes (#47903)\r\n    * Add CohereCompass modeling (#47878)","publishedAt":"2026-08-26T12:35:15.000Z","fetchedAt":"2026-08-26T12:38:20.198Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.16.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671"},{"type":"image","url":"https://github.com/user-attachments/assets/106ef712-9f45-43c9-98b7-ce6e4a7c136d","r2Key":"releases/4ef8217aedb5d74c6f53450727c97475ebe757c022e7c7af6a1fc51caa081c56.png","r2Url":"https://media.releases.sh/releases/4ef8217aedb5d74c6f53450727c97475ebe757c022e7c7af6a1fc51caa081c56.png"}],"coverageCount":0},{"id":"rel_wmuvrQ6PISoRhoH7yKzwz","version":"v5.15.1","type":"feature","title":"Patch release: v5.15.1","summary":"Fixed DFlash candidate token device mismatch with device_map='auto', aligned logit distributions for sampling-based candidate generators, fixed MTP config when mlp_layer_types is absent, and added a fallback from Lanczos to bicubic filtering on CUDA so images process on accelerator. Also fixed gemma4 video device placement.","titleGenerated":"Transformers v5.15.1 fixes DFlash and MTP issues, Lanczos fallback","titleShort":"DFlash and MTP fixes; Lanczos falls back on cuda","breaking":"none","importance":2,"content":"# Patch release v5.15.1\r\n\r\nThis patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter.\r\n\r\nIt contains the following commits:\r\n\r\n- Fix DFlash candidate token device mismatch with device_map=\"auto\" (#47877) by @sywangyi and @Cyrilvallez\r\n- Align logit distributions for CandidateGenerators using sampling  (#48007) by @Cyrilvallez\r\n- Fix MTP config when mlp_layer_types is absent (#48015) by @Cyrilvallez \r\n- Fallback from 'lanczos' to 'bicubic' when on cuda (#48026) by @zucchini-nlp\r\n- Fix gemma4 video to device (#47896) by @guarin\r\n","publishedAt":"2026-08-19T10:50:47.000Z","fetchedAt":"2026-08-19T10:54:34.224Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.15.1","media":[],"coverageCount":0},{"id":"rel_XopFHQ5vmoG9cfSb0hr4A","version":"v5.15.0","type":"feature","title":"Release: v5.15.0","summary":"Adds support for Meta's Muse Glimmer 30B multimodal model plus GraniteSWA, GraniteMoeSWA, A.X-K1/K2, and Cosmos3 Edge. Kernels for linear attention models (Mamba, GDN, Conv-only) are now opt-in rather than mandatory, cache cropping only accepts negative offsets, and T5 family gains SDPA support with a possible default attention change.","titleGenerated":"Transformers v5.15.0 adds Muse Glimmer, makes kernels opt-in for linear attention","titleShort":"Muse Glimmer supported; kernels now opt-in for linear attention","breaking":"major","importance":4,"content":"# Release v5.15.0\r\n\r\n## New Model additions\r\n\r\n### Meta Muse Glimmer\r\n\r\nMuse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups.\r\n\r\nMuse Glimmer is a dense 30B parameter model consisting of:\r\n- 2B ViT-style encoder for vision (Perception Encoder)\r\n- 28B parameter text decoder\r\n\r\nWe're covering it in the following blogpost: http://hf.co/blog/muse-glimmer\r\n\r\n<img width=\"960\" height=\"1787\" alt=\"image\" src=\"https://github.com/user-attachments/assets/3d8e548e-f84f-4269-8bd0-a12722d7ab01\" />\r\n\r\n---\r\n\r\n### GraniteMoeSWA & GraniteSWA\r\n\r\n<img width=\"1013\" height=\"389\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2c2b87f0-466a-413a-a4be-25ceae49c9a5\" />\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/granitemoe_swa)\r\n* Add Granite-swa and Granitemoe-swa model support (#47179) by @daviswer in [#47179](https://github.com/huggingface/transformers/pull/47179)\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/granite_swa)\r\n* Add Granite-swa and Granitemoe-swa model support (#47179) by @daviswer in [#47179](https://github.com/huggingface/transformers/pull/47179)\r\n\r\n---\r\n\r\n### A.X-K1 & A.X-K2\r\n\r\n<img width=\"580\" height=\"319\" alt=\"image\" src=\"https://github.com/user-attachments/assets/a20665bb-43ee-4af0-bb6f-80da495ea4f3\" />\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/axk2)\r\n* Add AXK2 from SKT (#47528) by @vasqu in [#47528](https://github.com/huggingface/transformers/pull/47528)\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/axk1)\r\n* add_axk1 (#46867) by @kmswin1 in [#46867](https://github.com/huggingface/transformers/pull/46867)\r\n\r\n---\r\n\r\n### Cosmos3 Edge\r\n\r\n<img width=\"1280\" height=\"720\" alt=\"image\" src=\"https://github.com/user-attachments/assets/16e41b90-11a5-4f22-8c56-11162b9c0a5f\" />\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/cosmos3_edge)\r\n* Add Cosmos3 Edge model support (#47181) by @atharvajoshi10 in [#47181](https://github.com/huggingface/transformers/pull/47181)\r\n\r\n\r\n## Breaking changes\r\n\r\nKernels are now opt-in rather than mandatory for linear attention models (Mamba, GDN, Conv-only, etc.), so users who relied on automatic kernel selection must explicitly enable kernels to maintain previous behavior.\r\n* 🚨 [`Kernels`] Refactor all linear attn models & native kernels fallback (#47630) by @vasqu\r\n\r\nThe cache cropping API now only accepts negative values (relative offsets) instead of absolute sizes, so users calling crop methods directly must update their code to pass negative values accordingly.\r\n* 🚨 [cache] Cropping can only be done with negative values (#47720) by @Cyrilvallez\r\n\r\nT5 and its model family (MT5, LongT5, etc.) now support SDPA and other attention backends via `ALL_ATTENTION_FUNCTIONS`, meaning the default attention implementation may change and users relying on the previous eager-only path should explicitly set `attn_implementation=\"eager\"` if needed.\r\n* 🚨 Enable SDPA (and other attention backends) for T5 and propagate to the T5 family (#47014) by @jiqing-feng\r\n\r\nSeveral small private helper functions (e.g., `_is_url`, `_build_image_tokens`) have been removed from multimodal processor files, so users or downstream libraries that imported these private functions directly must remove or replace those references.\r\n* :rotating_light: Processors update the rest (#46556) by @zucchini-nlp\r\n\r\n\r\n## Attention\r\n\r\nThis release includes several attention fixes and improvements, including correcting Multi-Head Latent Attention (MLA) cache compression, optimizing Flash Attention max sequence length computation in vision models, and fixing bugs in CTRL flex-attention and SDPA prefill with position bias. Additional changes refactor linear attention models for better maintainability, make Gemma 4's heterogeneous attention config explicit, and improve MPS support via metal-flash-sdpa integration.\r\n\r\n\r\n* [Fix] Fix multi-head latent attention (MLA) (#47761) by @remi-or in [#47761]\r\n* Refactor all linear attention models to latest best standards for convolution (#47452) by @Cyrilvallez in [#47452]\r\n* Allow metal-flash-sdpa for OpenAIPrivacyFilter on MPS (#46740) by @ArthurZucker in [#46740]\r\n* Use new `per_layer_config` for Gemma 4 so that heterogeneous attention config is explicit (#47384) by @hmellor in [#47384]\r\n* add paged attention tests support for XPU (#47163) by @kaixuanliu in [#47163]\r\n* Move `value` padding into the attention interfaces that need it (#47451) by @hmellor in [#47451]\r\n* Simplify function dispatch for linear attention (#47450) by @Cyrilvallez in [#47450]\r\n* Optimize flash attention max seqlen computation in vision attention (#47170) by @ShareLer in [#47170]\r\n* Fix `BlockMask` crash in CTRL flex-attention generation (#46854) by @jiqing-feng in [#46854]\r\n* [CB] Automatically switch attention implementation to flash (#47330) by @remi-or in [#47330]\r\n* Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez in [#47359]\r\n\r\n\r\n## Vision\r\n\r\nVision improvements in this release include performance optimizations such as faster image preprocessing for vision-language models (GLM4V, MiniMaxM3-VL, and others) by eliminating redundant tensor copies, and more efficient Flash Attention variable-length paths by precomputing maximum sequence lengths once per forward pass. Several bug fixes were also applied, including correcting dtype alignment in Kosmos2/Kosmos2_5 embedding merges, fixing a position-embedding initialization fallback in Phi4Multimodal, resolving PIL resize parity in Hunyuan-VL, and patching stop-sequence handling in the image-text-to-text pipeline.\r\n\r\n\r\n* Modularize qwen-format vision processors (#47573) by @zucchini-nlp in [#47573]\r\n* Update daily CI Docker image to torch 2.13.0 / CUDA 13.0 (#47738) by @ydshieh in [#47738]\r\n* Align image feature dtype in kosmos2 and kosmos2_5 embedding merge (#47691) by @<NOT FOUND> in [#47691]\r\n* Speed up image preprocessing for vision-language models (#47453) by @labAxiaoming in [#47453]\r\n* Fix vision position-embedding init width fallback in Phi4Multimodal (#47509) by @<NOT FOUND> in [#47509]\r\n* Fix Hunyuan-VL PIL image resize parity with reference preprocessing (#47233) by @IMvision12 in [#47233]\r\n* Fix image-text-to-text stop_sequence handling (#47032) by @Sunt-ing in [#47032]\r\n* Refactor image loading in tests to use load_test_image helper (#47218) by @LevelVoid in [#47218]\r\n\r\n\r\n## Generation\r\n\r\nSeveral generation improvements and bug fixes were made, including enabling batched audio generation for Qwen2.5/3-Omni, allowing sliding window cache layers to work with speculative decoding, and fixing memory overhead from static cache persistence across `generate()` calls. Multiple model-specific bugs were also resolved, including crashes in KyutaiSpeechToText, MusicgenForCausalLM, CTRL flex-attention, and assisted decoding for EncoderDecoder cache and OlmoHybrid models.\r\n\r\n\r\n* Align OlmoHybrid to use a native cache in generate (#47604) by @Cyrilvallez in [#47604]\r\n* [generate] Stop setting the static cache as an attribute to save memory (#47731) by @Cyrilvallez in [#47731]\r\n* Add support for batched Qwen2.5/3-Omni audio generation (#47186) by @IMvision12 in [#47186]\r\n* [cache] Allow sliding window layers to be roll-backed for speculative decoding (#47447) by @Cyrilvallez in [#47447]\r\n* Fix shape mismatch in KyutaiSpeechToText `generate()` last window (#46952) by @jiqing-feng in [#46952]\r\n* Fix typo in `MusicgenForCausalLM.generate()` (#46974) by @jiqing-feng in [#46974]\r\n* Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid (#47361) by @Cyrilvallez in [#47361]\r\n\r\n\r\n## Cache\r\n\r\nSeveral cache-related bugs were fixed, including correcting NemotronH's missing `\"mlp\"` layer-type mapping, resolving recurrent-layer padding masks being skipped during chunked prefill and cache continuation for hybrid models, and fixing assisted decoding for models with `EncoderDecoderCache` and OlmoHybrid. Additional improvements include aligning OlmoHybrid to use a native cache, enabling sliding window layers to support speculative decoding rollback, and stopping the static cache from being stored as a model attribute to reduce unexpected memory overhead.\r\n\r\n* [docs] MPS graph cache (#47304) by @stevhliu in [#47304]\r\n* Fix NemotronH: Register `\"mlp\"` in the cache layer-type mappings (#47535) by @qgallouedec in [#47535]\r\n* Fix recurrent-layer padding mask being skipped on continued forwards (chunked prefill, cache continuation) (#47087) by @abcgco in [#47087]\r\n\r\n\r\n## Kernels\r\n\r\n⚠️ The `kernels` python package will very likely be a required dependency for `transformers[torch]` in the near future. This will help us deliver maximum performance to all users; kernels will only be downloaded from trusted publishers manually approved by the HF team. Please let us know of any issues you're facing beforehands so that we may solidify our integration.\r\n\r\nImproved robustness of the kernels integration by refactoring function handling to use layer repos, fixing CI EROFS fallback patches for kernel downloads via `HfApi`, resolving a positional argument collision in `causal_conv1d_fn`, and bumping the FP8 kernels version to prevent NaNs.\r\n\r\n* [conftest] Fix EROFS fallback for kernel downloads (correct interception point) (#47794) by @ydshieh in [#47794]\r\n* [conftest] Fix EROFS fallback for kernel downloads via HfApi (#47791) by @ydshieh in [#47791]\r\n* [`Kernels`] Refactor function handling (#46883) by @vasqu in [#46883]\r\n* Kernels and loaders robustification (#47334) by @IlyasMoutawwakil in [#47334]\r\n* Fix `causal_conv1d_fn` positional `activation` colliding with hub kernel's `seq_idx` (#47527) by @qgallouedec in [#47527]\r\n* [`FP8`] Bump kernels version (#47344) by @vasqu in [#47344]\r\n* [docs] FlashAttention kernel fallback (#47345) by @stevhliu in [#47345]\r\n\r\n\r\n## Quantization\r\n\r\nQuantization support was expanded with FP8 kernels for compressed-tensors models, fixes for FP8 module normalization and format-based compression detection, and a multi-device MXFP4 dequantization race condition fix. GPTQ and MXFP4 tests were also extended to cover Intel XPU devices.\r\n\r\n\r\n* extend tests/quantization/gptq/test_gptq.py::GPTQTestCUDA and tests/q… (#47166) by @sywangyi in [#47166]\r\n* Compressed tensors fp8 (#47216) by @SunMarc in [#47216]\r\n* Fix A.X-K2 fp8 modules_to_not_convert normalization for the gated-norm MLP (#47578) by @kmswin1 in [#47578]\r\n* [Quantization]: Refactor is_quantization_compressed for format-based detection (#47152) by @rigen1048 in [#47152]\r\n* Fix multi-device mxfp4 dequantization race in `_convert_moe_packed_tensors` (#47423) by @kaixuanliu in [#47423]\r\n\r\n\r\n## Audio\r\n\r\nBatched audio generation is now supported for Qwen2.5/3-Omni, and several bug fixes were applied across audio models, including a dtype mismatch in Gemma4 audio feature merging, a bfloat16 positional embedding error in AudioFlamingo3, and missing backend requirement guards for Voxtral. The VibeVoice ASR processor was also updated to make audio input optional and support multiple audios per prompt.\r\n\r\n\r\n* feat[vLLM x v5]: Make audio optional and support multiple audios in VibeVoice ASR processor (#47483) by @harshaljanjani in [#47483]\r\n* Fix Gemma4 audio feature dtype mismatch in masked_scatter (#47482) by @danielhanchen in [#47482]\r\n* [fix] fix requirements audio feature and proc (#47113) by @eustlb in [#47113]\r\n* [AudioFlamingo3] Fix bfloat16 dtype mismatch in audio encoder positional embedding (#47258) by @snkii in [#47258]\r\n\r\n\r\n## Parallelization\r\n\r\nExpanded FSDP support across 94 `ForCausalLM` model classes with auto-generated FSDP plans, added end-to-end FSDP tests including distributed checkpoint save/load and generation, and introduced a dedicated FSDP CI job. Additionally, fixed a device mismatch bug in `create_bidirectional_sliding_window_mask` under model parallelism and resolved a tensor parallel inference issue for models with tied embeddings.\r\n\r\n\r\n* skip fsdp tests when backend is mps (#47601) by @3outeille in [#47601]\r\n* Fix model parallel device mismatch in `create_bidirectional_sliding_window_mask` (#47560) by @abcgco in [#47560]\r\n* Add FSDP plans to all models (#47165) by @3outeille in [#47165]\r\n* Fix TP inference for tied embedding (#47503) by @3outeille in [#47503]\r\n* Add FSDP CI and end-to-end FSDP tests + save fsdp (#47357) by @3outeille in [#47357]\r\n\r\n\r\n## Tokenization\r\n\r\nThis release adds native support for Mistral's \"tekken\" tokenizer format via AutoTokenizer, fixes a CodeLlama tokenizer bug where leading whitespace was incorrectly dropped during decode, and patches a potential ReDoS vulnerability caused by unescaped tokenizer filenames being used as regex patterns in `from_pretrained`.\r\n\r\n\r\n* [Mistral] Add native tekken tokenizer support to AutoTokenizer (#47507) by @juliendenize in [#47507]\r\n* Fix CodeLlama tokenizer dropping leading whitespace on decode (#47488) by @SuryanshSS1011 in [#47488]\r\n* Fix potential ReDoS by escaping tokenizer filename used as regex pattern (#47498) by @hameedibrh in [#47498]\r\n\r\n\r\n## Serve\r\n\r\nImproved the `serve` chat parsing to unify streaming and non-streaming paths under a single response parser that handles tool calls, reasoning, and content, simplifying the addition of new model support. Additionally, hardened daily CI reporting by fixing GitHub API diagnostic output being captured in Slack payloads and adding rate-limit resilience to prevent report failures when paginating large job matrices.\r\n\r\n\r\n* CI: Log GitHub API diagnostics to stderr (#47635) by @tarekziade in [#47635]\r\n* Update serve chat parsing (#46267) by @SunMarc in [#46267]\r\n* ci: harden daily CI reporting against GitHub API rate limits (#47382) by @tarekziade in [#47382]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Fix cached_files silently returning stale file on read-only filesystem (EROFS) (#47852) by @ydshieh in [#47852]\r\n* Fix `PhimoeIntegrationTest` (#46539) by @ydshieh in [#46539]\r\n* make examples under doc device agnostic (#47812) by @kaixuanliu in [#47812]\r\n* cancel deterministic for XPU in gemma4 tests (#47790) by @kaixuanliu in [#47790]\r\n* Serialize post-mlinter-review after post-link to avoid PR description race (#47832) by @ydshieh in [#47832]\r\n* Use content hash for mlinter review deduplication (#47830) by @ydshieh in [#47830]\r\n* Add new args in auto-docstring (#47737) by @zucchini-nlp in [#47737]\r\n* Add check_model_inits.py (#47656) by @guarin in [#47656]\r\n* add xpu in installation guide (#47785) by @sywangyi in [#47785]\r\n* Fix mlinter review job: checkout before artifact download (#47820) by @ydshieh in [#47820]\r\n* Post mlinter findings as inline PR review comments (#47819) by @ydshieh in [#47819]\r\n* Fix ci style (#47818) by @vasqu in [#47818]\r\n* Hotfix axk2 indexer norm (#47810) by @kmswin1 in [#47810]\r\n* open fla support for XPU to benefit from the acceleration (#47799) by @kaixuanliu in [#47799]\r\n* [docs] Update BatchEncoding.to() type annotation and docstring (#47789) by @samyuktahegde in [#47789]\r\n* Fix linting (#47807) by @Cyrilvallez in [#47807]\r\n* Fix patching in some models (#47798) by @zucchini-nlp in [#47798]\r\n* Add post-mlinter-review job to post-dashboard-link workflow (#47800) by @ydshieh in [#47800]\r\n* Migrate torchao integration off deleted torchao.dtypes (#47797) by @vkuzo in [#47797]\r\n* [conftest] Also wrap snapshot_download for EROFS fallback (#47796) by @ydshieh in [#47796]\r\n* Fix MI355 CI: bump hf-workflows pin to NUM_SLICES=4 (#47792) by @Abdennacer-Badaoui in [#47792]\r\n* [Fix] Wrong type hint in get_number_of_image_patches (#47788) by @remi-or in [#47788]\r\n* Fix Dac offload tests (#47775) by @guarin in [#47775]\r\n* [Fix] Swapped height and width in KimiK25 (#47786) by @remi-or in [#47786]\r\n* Remove dangling files and folders (#47764) by @Cyrilvallez in [#47764]\r\n* Fix AI-written conversion mappings (#47755) by @Cyrilvallez in [#47755]\r\n* update mistral common version for PR 47507 (#47677) by @itazap in [#47677]\r\n* Fix spelling/grammar in model files (batch 3/3) (#47684) by @Rocketknight1 in [#47684]\r\n* Fix spelling/grammar in core library, examples, and utils (#47685) by @Rocketknight1 in [#47685]\r\n* Fix spelling/grammar in model files (batch 2/3) (#47683) by @Rocketknight1 in [#47683]\r\n* update sonicmoe versions (#47769) by @IlyasMoutawwakil in [#47769]\r\n* clean up reverse_op fixme in compressed_tensors (#47701) by @DhanushPillay in [#47701]\r\n* PR CI with torch 2.13 (#47767) by @ydshieh in [#47767]\r\n* Import utils compilation fixes (#47726) by @IlyasMoutawwakil in [#47726]\r\n* Fix DBRX MoE hidden size and expert GLU transposes (#47671) by @kaixuanliu in [#47671]\r\n* Fix multi token decode merging (#47762) by @IlyasMoutawwakil in [#47762]\r\n* Fix: Remove redundant @can_return_tuple conflicting with @capture_out… (#47733) by @guarin in [#47733]\r\n* fix processor config nested key fallback (#47628) by @YunzhuLu in [#47628]\r\n* [CI - Debug] Skip /transformers-dependent steps for CPU runner (#47759) by @ydshieh in [#47759]\r\n* [CI] Add CPU runner support to ssh-runner workflow (#47757) by @ydshieh in [#47757]\r\n* Simplify reverse weight conversion (#47725) by @Cyrilvallez in [#47725]\r\n* Remove useless linting for inv_freq (#47753) by @Cyrilvallez in [#47753]\r\n* Use explicit nn.Buffer everywhere for modular (#47722) by @Cyrilvallez in [#47722]\r\n* Remove stale and redundant _no_split_modules entries (#47645) by @guarin in [#47645]\r\n* Executorch exporter fixes (#47243) by @IlyasMoutawwakil in [#47243]\r\n* Acc fix in xpu (#47500) by @sywangyi in [#47500]\r\n* update Dockerfile for xpu torch2.13 (#47502) by @sywangyi in [#47502]\r\n* feat[vLLM]: Support text replacement offsets in the remaining old-format processors (#47614) by @harshaljanjani in [#47614]\r\n* Fix ImportError in transformers.exporters on torch < 2.8 (#47711) by @Neal006 in [#47711]\r\n* Fix failing tests for axk1 and axk2 (#47727) by @kaixuanliu in [#47727]\r\n* Fix failing tests for granite_swa and granitemoe_swa (#47723) by @kaixuanliu in [#47723]\r\n* skip invalid test cases for inkling tests (#47493) by @kaixuanliu in [#47493]\r\n* [docs] Fix BatchEncoding documentation inconsistencies (#47647) by @samyuktahegde in [#47647]\r\n* Fix spelling/grammar in model files (batch 1/3) (#47682) by @Rocketknight1 in [#47682]\r\n* Fix spelling/grammar in English docs (q → z) (#47681) by @Rocketknight1 in [#47681]\r\n* Fix spelling/grammar in English docs (h → p) (#47680) by @Rocketknight1 in [#47680]\r\n* Fix spelling/grammar in English docs (a → g) (#47679) by @Rocketknight1 in [#47679]\r\n* fix npu check (#47587) by @DhanushPillay in [#47587]\r\n* fix: correct text input validation logic in 8 multimodal processors (and → or) (#47663) by @AbdullahRasheed45 in [#47663]\r\n* Fix feature dtype mismatch in masked_scatter for seven multimodal models (#47673) by @<NOT FOUND> in [#47673]\r\n* [docs] storing and loading chat templates (#47650) by @stevhliu in [#47650]\r\n* adding amd quark config class changes (#47322) by @debasisdwivedy in [#47322]\r\n* Fix compressed tensors impl  (#47652) by @SunMarc in [#47652]\r\n* Hoist special-token lookups in wav2vec2 decode paths and drop a dead filter in wav2vec2_phoneme (#47557) by @ishan-1010 in [#47557]\r\n* silencing elastic warning by import distributed lib inside functions  (#47665) by @3outeille in [#47665]\r\n* Fix missing github_utils.py download in PR CI dashboard workflow (#47668) by @ydshieh in [#47668]\r\n* Exportable kimi (#47096) by @IlyasMoutawwakil in [#47096]\r\n* better guarding to handle torch compiled with USE_DISTRIBUTED=0 (#47619) by @3outeille in [#47619]\r\n* Remove gemma4 warnings (#47664) by @Cyrilvallez in [#47664]\r\n* [Chat Parsing] Type inline tool-call arguments from the calling tool's JSON Schema (#47529) by @yonigozlan in [#47529]\r\n* Improve Trainer DataLoader Controls for Streaming and Multiprocessing (#47164) by @muyihao in [#47164]\r\n* Update maintainer list (#47644) by @Rocketknight1 in [#47644]\r\n* Allow position_ids_start=2 on DataCollatorWithFlattening for RoBERTa etc. (#47525) by @tomaarsen in [#47525]\r\n* Remove Rotary warning (#47642) by @Cyrilvallez in [#47642]\r\n* Drop multimodal inputs natively in prepare_inputs_for_generation if not in prefill (#47622) by @Cyrilvallez in [#47622]\r\n* Simplify all Rotary modules (#47598) by @Cyrilvallez in [#47598]\r\n* [docs] response_template when serving (#47626) by @stevhliu in [#47626]\r\n* [docs] MTP support (#47301) by @stevhliu in [#47301]\r\n* Fix GPT-2 c_proj depth scaling initialization (#47459) by @DavidJohnQuinlan in [#47459]\r\n* Fix fp8_linear compilability  (#47623) by @IlyasMoutawwakil in [#47623]\r\n* byebye torch 2.4 (#47609) by @ydshieh in [#47609]\r\n* Vectorize NoRepeatNGramLogitsProcessor and remove its host sync (#47571) by @hameedibrh in [#47571]\r\n* Fix CUDA Graph breaking host to device copy from scalar tensor allocation (#47547) by @hmellor in [#47547]\r\n* Fix some processors (#47608) by @zucchini-nlp in [#47608]\r\n* CI: Add serge review relay workflow and review rules (#47610) by @tarekziade in [#47610]\r\n* [docs] Exporters (#47374) by @stevhliu in [#47374]\r\n* CI: use a single function for GH calls (#47474) by @tarekziade in [#47474]\r\n* Remove redundant guarding for distributed (#47570) by @3outeille in [#47570]\r\n* Fix modular for mamba packages (#47494) by @Cyrilvallez in [#47494]\r\n* Remove deprecated conversion in Kimi (#47581) by @zucchini-nlp in [#47581]\r\n* CI: use transformers-ci daily workflow with OTEL (#47360) by @tarekziade in [#47360]\r\n* Fix slow tensor path in _check_special_mm_tokens (#47580) by @guan404ming in [#47580]\r\n* CI: let's run integration failure cron at 10pm (#47537) by @tarekziade in [#47537]\r\n* Better and more extensive tests for RoPE (#46912) by @zucchini-nlp in [#46912]\r\n* Fix mamba2 family decode and simplify all reshape ops (#47569) by @Cyrilvallez in [#47569]\r\n* [DiffusionGemma] Cast the decoder padding mask to bool (#47295) by @kashif in [#47295]\r\n* General maintenance  (#47517) by @zucchini-nlp in [#47517]\r\n* fixed the benchmark script with DistributedConfig  (#47568) by @tarekziade in [#47568]\r\n* Fix failing tests for mimo_v2_flash (#47284) by @kaixuanliu in [#47284]\r\n* Make tokenization_mistral_common importable without mistral_common installed (#47397) by @juliendenize in [#47397]\r\n* Use --flake-runs=1 in check_bad_commit.py for PR comment CI (#47522) by @ydshieh in [#47522]\r\n* Deprecate the old response_schema (#47320) by @Rocketknight1 in [#47320]\r\n* fix: add pickle support to _LazyConfigMapping for spawn multiprocessing (#46026) by @kfojcik-intel in [#46026]\r\n* Delete old deprecations (#47518) by @zucchini-nlp in [#47518]\r\n* Run only @slow tests in PR comment CI (#47521) by @ydshieh in [#47521]\r\n* Fix incorrect type hint (#47519) by @hmellor in [#47519]\r\n* Deprecate CB config in gen configuration (#47291) by @remi-or in [#47291]\r\n* [Offloading] [Bugfix] Fix fully offloaded model saving (#47336) by @kylesayrs in [#47336]\r\n* CI: fix torchaudio pinning +proper break in rnnt (#47422) by @tarekziade in [#47422]\r\n* Fix Qwen2.5-Omni Token2Wav DiT rotary embedding layout (interleaved cos/sin) (#47403) by @HenryVarro666 in [#47403]\r\n* CI: add reproduce mode to serge verify caller (#47492) by @tarekziade in [#47492]\r\n* [fix][whisper]: fix max_new_tokens handling (#46795) by @eustlb in [#46795]\r\n* Isolate MLA KV expansion to make it easier to bypass (#47460) by @hmellor in [#47460]\r\n* Fix loss alignment and Trainer token counting for encoder decoder models (#46903) by @OmkumarSolanki in [#46903]\r\n* Fix the HunyuanVL's torchvision backend (#47499) by @Mi-Jiazhi in [#47499]\r\n* [DiffusionGemma] Support gradient checkpointing (#46572) by @kashif in [#46572]\r\n* Fix failing tests for zaya (#47268) by @kaixuanliu in [#47268]\r\n* tipsv2_dpt: fix failing tests for XPU (#47292) by @kaixuanliu in [#47292]\r\n* Fix some failed test cases related with XPU Expectations (#47173) by @kaixuanliu in [#47173]\r\n* fix: guard DTensor import in sharding_utils.py for PyTorch < 2.5 (#47481) by @<NOT FOUND> in [#47481]\r\n* CPU can incur a slow path on non-contiguous magnitudes (#47351) by @vbayanag in [#47351]\r\n* Route chat management calls to the service root (#47138) (#47303) by @dhruv7477 in [#47303]\r\n* fix: liger unnecessarily materializes logits in VRAM during eval, causing OOM (#45273) by @excepshenal in [#45273]\r\n* Fix gradient inflation when combining label smoothing with gradient accumulation (#47261) by @Incheonkirin in [#47261]\r\n* Fix MoE expert decompression for non-32-divisible bit widths (#47315) by @KKothuri in [#47315]\r\n* [Qwen3ASR] Add hotword parsing, and fix language parsing and training. (#47111) by @ebezzam in [#47111]\r\n* Remove deprecated training args and `is_fast` property (#46917) by @cyyever in [#46917]\r\n* Consistent output shape from `get_image_features` (#46405) by @zucchini-nlp in [#46405]\r\n* fix failed test cases for qwen3_omni_moe model (#47449) by @kaixuanliu in [#47449]\r\n* Fix double-shifted training loss in GitForCausalLM (#47395) by @<NOT FOUND> in [#47395]\r\n* Fix CohereASR training-loss double-shift (same as Moonshine fix #46784) (#46895) by @sharmax-vikas in [#46895]\r\n* Warn when `group_by_length` is silently ignored for iterable datasets (#47379) by @qgallouedec in [#47379]\r\n* Update bug report list (#46607) by @molbap in [#46607]\r\n* fix: remove unreachable return in special token builder (#47420) by @hai1222 in [#47420]\r\n* Add Harry to slow CI (#47454) by @vasqu in [#47454]\r\n* BLT: vectorize patch length processing (#47385) by @sj0618 in [#47385]\r\n* Fix `TrackioCallback` fails to log evaluation metrics after training ends (#46935) by @lewtun in [#46935]\r\n* [Kimi] add integration tests (#47383) by @zucchini-nlp in [#47383]\r\n* Add distributed runtime utils and DistributedMixin  (#47352) by @3outeille in [#47352]\r\n* Fix Cosmos 3 Edge Patch packing order (#47399) by @atharvajoshi10 in [#47399]\r\n* Fix yarn `mscale_all_dim` for DeepSeek v2 and Mistral 4 (#47435) by @hmellor in [#47435]\r\n* Hoist special-token lookups out of per-token loops in six slow tokenizers (#47425) by @ishan-1010 in [#47425]\r\n* fix typos and variable naming in quicktour.md (#47418) by @yashasvi-srivastava21 in [#47418]\r\n* Normalize multimodal input keys in AnyToAnyPipeline (#47074) by @Sunt-ing in [#47074]\r\n* Fix Aria checkpoint key conversion mapping (#47151) by @sywangyi in [#47151]\r\n* extend tests/models/qwen3_next/test_modeling_qwen3_next.py::Qwen3Next… (#47184) by @sywangyi in [#47184]\r\n* CI: add serge verify (GPU) caller workflow (#47381) by @tarekziade in [#47381]\r\n* Fix orthogonal_ init for low-precision dtypes (bf16/fp16) (#47252) by @janbernloehr in [#47252]\r\n* [docs] Fix decode examples and expected output in fast_tokenizers (#47369) by @samyuktahegde in [#47369]\r\n* [serge] Fix 20 integration tests for model `whisper` failing with `output_mismatch` (list output differs (10), other (6) (#47150) by @sergereview[bot] in [#47150]\r\n* [Mistral] Move MistralConverter into integrations/mistral/ package (#46603) by @juliendenize in [#46603]\r\n* Fix GLM video frame padding for temporal patches (#47141) by @labAxiaoming in [#47141]\r\n* [docs] Inkling (#47350) by @stevhliu in [#47350]\r\n* Fix Daily CI reporting issues (#47364) by @tarekziade in [#47364]\r\n* [`peft`] Support key_mapping with PEFT models (#46766) by @tomaarsen in [#46766]\r\n* Fix model tests for tipsv2 (#47356) by @kaixuanliu in [#47356]\r\n* Fix TimesFM 2.5 window_size AttributeError (#47363) by @kashif in [#47363]\r\n* fix: allow num_labels property to return None when id2label is unset (#47069) by @SebTardif in [#47069]\r\n* [Tests] Fix slow video tensor creation from list of numpy arrays in SmolVLM (#44731) by @Defalt-Meh in [#44731]\r\n* Update dev version on main (#47366) by @vasqu in [#47366]\r\n* Fix deepgemm on multiple devices (#47323) by @IlyasMoutawwakil in [#47323]\r\n* ci: Handle empty GitHub token in CI run lookup (#47362) by @tarekziade in [#47362]\r\n* Fix inkling feature extractor (#47349) by @ArthurZucker in [#47349]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @ydshieh\r\n    * Fix cached_files silently returning stale file on read-only filesystem (EROFS) (#47852)\r\n    * Fix `PhimoeIntegrationTest` (#46539)\r\n    * Serialize post-mlinter-review after post-link to avoid PR description race (#47832)\r\n    * Use content hash for mlinter review deduplication (#47830)\r\n    * Fix mlinter review job: checkout before artifact download (#47820)\r\n    * Post mlinter findings as inline PR review comments (#47819)\r\n    * Add post-mlinter-review job to post-dashboard-link workflow (#47800)\r\n    * [conftest] Also wrap snapshot_download for EROFS fallback (#47796)\r\n    * [conftest] Fix EROFS fallback for kernel downloads (correct interception point) (#47794)\r\n    * [conftest] Fix EROFS fallback for kernel downloads via HfApi (#47791)\r\n    * PR CI with torch 2.13 (#47767)\r\n    * [CI - Debug] Skip /transformers-dependent steps for CPU runner (#47759)\r\n    * [CI] Add CPU runner support to ssh-runner workflow (#47757)\r\n    * Update daily CI Docker image to torch 2.13.0 / CUDA 13.0 (#47738)\r\n    * Fix missing github_utils.py download in PR CI dashboard workflow (#47668)\r\n    * byebye torch 2.4 (#47609)\r\n    * Use --flake-runs=1 in check_bad_commit.py for PR comment CI (#47522)\r\n    * Run only @slow tests in PR comment CI (#47521)\r\n* @kaixuanliu\r\n    * make examples under doc device agnostic (#47812)\r\n    * cancel deterministic for XPU in gemma4 tests (#47790)\r\n    * open fla support for XPU to benefit from the acceleration (#47799)\r\n    * Fix DBRX MoE hidden size and expert GLU transposes (#47671)\r\n    * Fix failing tests for axk1 and axk2 (#47727)\r\n    * Fix failing tests for granite_swa and granitemoe_swa (#47723)\r\n    * skip invalid test cases for inkling tests (#47493)\r\n    * Fix failing tests for mimo_v2_flash (#47284)\r\n    * Fix failing tests for zaya (#47268)\r\n    * tipsv2_dpt: fix failing tests for XPU (#47292)\r\n    * Fix some failed test cases related with XPU Expectations (#47173)\r\n    * add paged attention tests support for XPU (#47163)\r\n    * Fix multi-device mxfp4 dequantization race in `_convert_moe_packed_tensors` (#47423)\r\n    * fix failed test cases for qwen3_omni_moe model (#47449)\r\n    * Fix model tests for tipsv2 (#47356)\r\n* @vasqu\r\n    * Fix ci style (#47818)\r\n    * 🚨 [`Kernels`] Refactor all linear attn models & native kernels fallback (#47630)\r\n    * [`Kernels`] Refactor function handling (#46883)\r\n    * Add AXK2 from SKT (#47528)\r\n    * Add Harry to slow CI (#47454)\r\n    * Update dev version on main (#47366)\r\n    * [`FP8`] Bump kernels version (#47344)\r\n* @kmswin1\r\n    * Hotfix axk2 indexer norm (#47810)\r\n    * Fix A.X-K2 fp8 modules_to_not_convert normalization for the gated-norm MLP (#47578)\r\n    * add_axk1 (#46867)\r\n* @remi-or\r\n    * [Fix] Fix multi-head latent attention (MLA) (#47761)\r\n    * [Fix] Wrong type hint in get_number_of_image_patches (#47788)\r\n    * [Fix] Swapped height and width in KimiK25 (#47786)\r\n    * Deprecate CB config in gen configuration (#47291)\r\n    * [CB] Automatically switch attention implementation to flash (#47330)\r\n* @juliendenize\r\n    * [Mistral] Add native tekken tokenizer support to AutoTokenizer (#47507)\r\n    * Make tokenization_mistral_common importable without mistral_common installed (#47397)\r\n    * [Mistral] Move MistralConverter into integrations/mistral/ package (#46603)\r\n* @IMvision12\r\n    * Add support for batched Qwen2.5/3-Omni audio generation (#47186)\r\n    * Fix Hunyuan-VL PIL image resize parity with reference preprocessing (#47233)\r\n* @jiqing-feng\r\n    * 🚨 Enable SDPA (and other attention backends) for T5 and propagate to the T5 family (#47014)\r\n    * Fix shape mismatch in KyutaiSpeechToText `generate()` last window (#46952)\r\n    * Fix typo in `MusicgenForCausalLM.generate()` (#46974)\r\n    * Fix `BlockMask` crash in CTRL flex-attention generation (#46854)\r\n* @tarekziade\r\n    * CI: Log GitHub API diagnostics to stderr (#47635)\r\n    * CI: Add serge review relay workflow and review rules (#47610)\r\n    * CI: use a single function for GH calls (#47474)\r\n    * CI: use transformers-ci daily workflow with OTEL (#47360)\r\n    * CI: let's run integration failure cron at 10pm (#47537)\r\n    * fixed the benchmark script with DistributedConfig  (#47568)\r\n    * CI: fix torchaudio pinning +proper break in rnnt (#47422)\r\n    * CI: add reproduce mode to serge verify caller (#47492)\r\n    * CI: add serge verify (GPU) caller workflow (#47381)\r\n    * ci: harden daily CI reporting against GitHub API rate limits (#47382)\r\n    * Fix Daily CI reporting issues (#47364)\r\n    * ci: Handle empty GitHub token in CI run lookup (#47362)\r\n* @daviswer\r\n    * Add Granite-swa and Granitemoe-swa model support (#47179)\r\n* @ShareLer\r\n    * Optimize flash attention max seqlen computation in vision attention (#47170)\r\n* @atharvajoshi10\r\n    * Fix Cosmos 3 Edge Patch packing order (#47399)\r\n    * Add Cosmos3 Edge model support (#47181)\r\n","publishedAt":"2026-08-10T10:28:13.000Z","fetchedAt":"2026-08-10T12:22:08.382Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.15.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/3d8e548e-f84f-4269-8bd0-a12722d7ab01","r2Key":"releases/3b3bd4fcc5a84e04d06a1843391c2595b9d5e000fcffb02e6acfa116b3ed210a.png","r2Url":"https://media.releases.sh/releases/3b3bd4fcc5a84e04d06a1843391c2595b9d5e000fcffb02e6acfa116b3ed210a.png"},{"type":"image","url":"https://github.com/user-attachments/assets/2c2b87f0-466a-413a-a4be-25ceae49c9a5","r2Key":"releases/1b7ede12ad94cbb53b97004a9c92c5359ed5f7119f2865ac624ab96e54805cc4.png","r2Url":"https://media.releases.sh/releases/1b7ede12ad94cbb53b97004a9c92c5359ed5f7119f2865ac624ab96e54805cc4.png"},{"type":"image","url":"https://github.com/user-attachments/assets/a20665bb-43ee-4af0-bb6f-80da495ea4f3","r2Key":"releases/1d646435d1a9e40097eed4cc9a4362074e69a8ecbb5cd263e272703f53a0526a.png","r2Url":"https://media.releases.sh/releases/1d646435d1a9e40097eed4cc9a4362074e69a8ecbb5cd263e272703f53a0526a.png"},{"type":"image","url":"https://github.com/user-attachments/assets/16e41b90-11a5-4f22-8c56-11162b9c0a5f","r2Key":"releases/644a7c38d9a3c00c595a29fee44a36c9a72e27c7bf73d7d361289decba7cf662.png","r2Url":"https://media.releases.sh/releases/644a7c38d9a3c00c595a29fee44a36c9a72e27c7bf73d7d361289decba7cf662.png"}],"coverageCount":0},{"id":"rel_lci2Xsv3nffvTCkArOPG0","version":"v5.14.1","type":"feature","title":"Patch release: v5.14.1","summary":"Fixed assisted decoding for models using EncoderDecoderCache (e.g., OlmoHybrid) and a SDPA prefill issue with position_bias during StaticCache. Also bumps FP8 kernels and fixes DeepGEMM on multiple devices.","titleGenerated":"Transformers v5.14.1 fixes assisted decoding and SDPA prefill with position bias","titleShort":"Assisted decoding fixed for EncoderDecoderCache models; SDPA prefill hardened","breaking":"none","importance":2,"content":"# Patch release v5.14.1\r\n\r\nThis patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. \r\nIt contains the following commits:\r\n\r\n- Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez\r\n- Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid (#47361) by @Cyrilvallez\r\n- [FP8] Bump kernels version (#47344) by @vasqu \r\n- Fix deepgemm on multiple devices (#47323) by @IlyasMoutawwakil ","publishedAt":"2026-07-16T09:41:36.000Z","fetchedAt":"2026-07-16T10:11:46.043Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.14.1","media":[],"coverageCount":0},{"id":"rel_YMP7trDgZ97mZ2oo0Vz-u","version":"v5.14.0","type":"feature","title":"Release v5.14.0","summary":"Added Inkling, a 975B-parameter multimodal model, and TIPSv2. GPTNeoX now remaps embed_out to lm_head and GPTBigCode has attention backend changes for vLLM compatibility. Multi-Token Prediction decoding support, SDPA prefill with FlashAttention for StaticCache (up to 260% faster), and numerous bug fixes across MoE, cache, and generation.","titleGenerated":"Transformers v5.14.0 adds Inkling 975B model and breaks GPTNeoX/GPTBigCode weight naming","titleShort":"Inkling 975B model added; GPTNeoX weight naming changed","breaking":"major","importance":4,"content":"# Release v5.14.0\r\n\r\n## New Model additions\r\n\r\n### Inkling (fresh from Thinking Machines): 975B total, 41B active\r\n\r\n* Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp \r\n\r\n<img width=\"3840\" height=\"2160\" alt=\"image\" src=\"https://github.com/user-attachments/assets/051f819a-512f-4987-9bee-6e2fa2af3db7\" />\r\n\r\n\r\nInkling is a general-purpose multimodal model that accepts text, image and audio inputs and\r\ngenerates text outputs. It is intended for use in English and other languages, and across\r\nmultiple coding languages. The model is designed to be used by developers building AI-\r\npowered applications, including agentic and tool-use systems, coding assistants, chatbots, and\r\nretrieval-augmented generation systems, and is suitable for general-purpose conversational\r\nuse, instruction-following, and other natural language and multimodal tasks. It is released with\r\nopen weights to support research, fine-tuning and integration into third-party products by\r\ndownstream developers.\r\n\r\n\r\n\r\n\r\n### TIPSv2\r\n<img width=\"1555\" height=\"1306\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2d9f21e5-05f8-4c36-93ef-22f03c089f52\" />\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/tipsv2)\r\n* Add TIPSv2 (#46347) by @Ternura143 in [#46347](https://github.com/huggingface/transformers/pull/46347)\r\n\r\n### TIPSv2 DPT\r\n<img width=\"794\" height=\"245\" alt=\"image\" src=\"https://github.com/user-attachments/assets/09c0d4da-6c1c-4229-bf02-a512ed435e50\" />\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/tipsv2_dpt)\r\n* Add TIPSv2 (#46347) by @Ternura143 in [#46347](https://github.com/huggingface/transformers/pull/46347)\r\n\r\n\r\n\r\n## :rotating_light:  Breaking changes\r\n\r\nGPTNeoX now remaps `embed_out` to `lm_head` and GPTBigCode has `_supports_attention_backend = True` enabled for vLLM compatibility; users relying on the previous weight naming or attention backend behavior for these models should update their code accordingly.\r\n* :rotating_light: Fix GPTBigCode and GPTNeoX for the Transformers modelling backend for vLLM (#47198) by @hmellor\r\n\r\n\r\n## Kernels\r\n\r\nSeveral kernel-related fixes and improvements were made, including pinning the `kernels` dependency to a compatible version in the benchmark workflow, removing a deprecated `package_name` argument from `LocalLayerRepository`, and making the DeepGEMM Triton fallback more robust when `CUDA_HOME` is unset or misconfigured. Additionally, SDPA prefill was updated to leverage the FlashAttention kernel with `StaticCache`, yielding significant performance gains (up to 260% faster for large input sizes).\r\n\r\n\r\n* Pin kernels to compatible version in benchmark workflow (#47339) by @tarekziade in [#47339]\r\n* [Fix] Remove deprecated argument from `kernels` call (#47100) by @remi-or in [#47100]\r\n* [Fix] Make DeepGEMM triton fallback more robust (#47126) by @remi-or in [#47126]\r\n* [sdpa] Allow prefill to use FA kernel with StaticCache (#47094) by @Cyrilvallez in [#47094]\r\n\r\n\r\n## Generation\r\n\r\nGeneration improvements include adding Multi-Token Prediction (MTP) decoding support, static ensemble verification for speculative decoding to improve draft token acceptance rates, and a fix for crashes in greedy assisted generation with different tokenizers. A misleading double-negative warning message for `synced_gpus` in continuous batching mode was also corrected.\r\n\r\n\r\n* [generation] Fix misleading synced_gpus warning in continuous batching (#47158) by @Partha-Shankar in [#47158]\r\n* [generate] Add proper MTP support (#46229) by @Cyrilvallez in [#46229]\r\n* Fix crash in greedy assisted generation with different tokenizers (#46936) by @Sunt-ing in [#46936]\r\n* [Generation] Add static ensemble verification for lossy speculative decoding (#45979) by @kasakh in [#45979]\r\n\r\n\r\n## Performance\r\n\r\nFixed a Flash Attention performance regression affecting models like Qwen3-VL and resolved a MoE decode optimization bug where the grouped-to-batched matrix multiplication switch was not applied to experts residing in submodels (e.g., VLMs with a nested text config).\r\n\r\n\r\n* Fix FA performance regression (#47134) by @andreasgoulas in [#47134]\r\n* Fix MoE decode optimization for experts living in a submodel (#47107) by @IlyasMoutawwakil in [#47107]\r\n* Make doc builds faster (#47099) by @mishig25 in [#47099]\r\n\r\n\r\n## Cache\r\n\r\nCache dispatch logic was simplified by introducing explicit layer-type mappings for sliding and static layers, reducing complexity in cache routing. Additionally, fixes were made for read-only cache failures in CPU CI environments and for MPS graph cache growth during variable-length batch training on Apple Silicon.\r\n\r\n\r\n* Fix CI read-only cache failures by patching cached_files in conftest (#47043) by @ydshieh in [#47043]\r\n* trainer: clear MPS graph cache via torch_empty_cache_steps (#45818) by @anagnorisis2peripeteia in [#45818]\r\n* [cache] Simplify cache dispatch based on layer_types (#47118) by @Cyrilvallez in [#47118]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* ci: cover xet as well (runtime error) (#47338) by @tarekziade in [#47338]\r\n* [docs] TokenizersBackend fallback (#47302) by @stevhliu in [#47302]\r\n* Resolve continuous batching XPU availability checks at runtime (#47185) by @kaixuanliu in [#47185]\r\n* [Nit] Add kernels_fallback_ok kwarg to is_flash_attn_N_available (#47318) by @remi-or in [#47318]\r\n* [Nit] Add expectations for gemma4 tests on H100 (#47311) by @remi-or in [#47311]\r\n* [docs] DeepGEMM requirements (#47324) by @stevhliu in [#47324]\r\n* DeepGEMM shouldn't pad on SM90 (#47313) by @IlyasMoutawwakil in [#47313]\r\n* Fix half-precision torch.compile crash in DETR-family sine position embeddings (#47238) by @David-Wu1119 in [#47238]\r\n* Fix hardcoded paths in siglip checkpoint/vocab loading (#47178) by @XanxusCrypto in [#47178]\r\n* Update AMD CI runner groups to amd-mi300 (#47307) by @Abdennacer-Badaoui in [#47307]\r\n* Point to Gemma 4 model in Gemma4ForCausalLM docstring example (#47255) by @lefft in [#47255]\r\n* Fix Qwen Omni batched text postprocessing (#47197) by @Sunt-ing in [#47197]\r\n* Fix AqlmConfig error messages to say \"int\" instead of \"float\" (#47089) by @Sreekant13 in [#47089]\r\n* Fix check for interactive stdout in _style function (#47283) by @smart8986 in [#47283]\r\n* Fix get_json_schema crash on non-string docstring choices (#47072) by @Sreekant13 in [#47072]\r\n* Make `MODEL_IDS_TO_TOKENIZERS_BACKEND` capture all DeepSeek R1 distills (#47296) by @hmellor in [#47296]\r\n* Update doc preprocessing regex to prevent ReDoS (#47187) by @WilliamRoyNelson in [#47187]\r\n* Shard on read Dtensor aware (#46717) by @3outeille in [#46717]\r\n* Switch AMD daily CI to mi300 runners (#47259) by @Abdennacer-Badaoui in [#47259]\r\n* tests: reduce processor test memory usage by using tiny Hub checkpoints (#47213) by @ydshieh in [#47213]\r\n* Torch compile backend defaults to \"neuron\" (#47035) by @michaelbenayoun in [#47035]\r\n* Fix flash-attn Docker build broken by setuptools 83 removing pkg_resources (#47251) by @ydshieh in [#47251]\r\n* Add heterogeneous config support (per-layer configuration) (#45333) by @eladsegal in [#45333]\r\n* [fix] update integration test values (#47146) by @eustlb in [#47146]\r\n* Fix DeepSpeed SP loss aggregation and LocalLayerRepository kwargs (#47073) by @sshivampeta in [#47073]\r\n* tests only for the top 10 download models (#47244) by @3outeille in [#47244]\r\n* Fix InputTokensDetails missing cache_write_tokens for openai>=2.34.0 (#47248) by @ydshieh in [#47248]\r\n* Revert \"Trigger a scheduled run\" (#47249) by @ydshieh in [#47249]\r\n* Remove Executorch from CI until latest version is supported and fully tested on CI env (#47242) by @IlyasMoutawwakil in [#47242]\r\n* Be more defensive with `remap_legacy_layer_types` for custom models (#47245) by @hmellor in [#47245]\r\n* Fix DistributedConfig docstring for unimplemented sp_plan  (#47237) by @3outeille in [#47237]\r\n* Switch mlinter to 0.1.2 (#47172) by @tarekziade in [#47172]\r\n* Trigger a scheduled run (#47209) by @ydshieh in [#47209]\r\n* Make executorch exporter tests always use xnnpack backend (#47201) by @tarekziade in [#47201]\r\n* No agent PR descriptions (#45790) by @Rocketknight1 in [#45790]\r\n* Clarify that max_steps is required for datasets without __len__ (#47155) by @albertvillanova in [#47155]\r\n* Cleanup pipelines, stop materializing generators (#47142) by @Rocketknight1 in [#47142]\r\n* Fix device_map computation when the no_split_modules have different sizes (#47203) by @Cyrilvallez in [#47203]\r\n* Add native FSDP2 module + migration (#46707) by @3outeille in [#46707]\r\n* Fix experts implementation in two spots (#47097) by @remi-or in [#47097]\r\n* [Fix] Remove old automatic cross attn pattern from output recorders (#47117) by @remi-or in [#47117]\r\n* 🌐 [i18n-KO] Translate accelerator_selection.md to Korean (#47157) by @kkwjk2718 in [#47157]\r\n* [i18n-KO] Translate optimum.md to Korean and fix Furiosa typo (#47156) by @kkwjk2718 in [#47156]\r\n* [docs] fix curly quotes rendering to straight quotes (#47135) by @clijo in [#47135]\r\n* Fix custom code which doesn't know about the new linear layer type names (#47174) by @hmellor in [#47174]\r\n* Reject path traversal in the `transformers_weights` config field (#46890) by @LinZiyuu in [#46890]\r\n* [docs] Custom code conversion mapping (#47114) by @stevhliu in [#47114]\r\n* Add exporters min version requirements and test skip (#47161) by @IlyasMoutawwakil in [#47161]\r\n* tests: reduce processor test memory usage and use tiny test assets (#47168) by @ydshieh in [#47168]\r\n* Clarify input device placement in the Quicktour inference example (#47136) by @samyuktahegde in [#47136]\r\n* Extend continuous batching memory prediction test to XPU (#47159) by @sywangyi in [#47159]\r\n* Fix case where `_LazyAutoMapping.register` is passed a `str` key (#47148) by @hmellor in [#47148]\r\n* [docs] MoE decode switching (#47149) by @stevhliu in [#47149]\r\n* add XPU output expectations for minicpm3 tests (#47092) by @kaixuanliu in [#47092]\r\n* Diffusion gemma: fix failed test cases (#47025) by @kaixuanliu in [#47025]\r\n* add XPU Expectation for cosmos3_omni tests (#46880) by @kaixuanliu in [#46880]\r\n* Fix IndexError Bug in XLMRoberta/Camembert ForMultipleChoice by restoring the pooler (#47147) by @pariidanDKE in [#47147]\r\n* Skip caching_allocator_warmup on Neuron (no reuse pool to warm; currently OOMs) (#47029) by @dacorvo in [#47029]\r\n* [docs] continuous batching (offloading behavior, max batch tokens, block size minimum) (#46925) by @stevhliu in [#46925]\r\n* [docs] fix autolinks (#46968) by @stevhliu in [#46968]\r\n* revert #47121 (#47144) by @eustlb in [#47144]\r\n* Fix output labels for AudioFlamingo3 (and related) models (#47112) by @ebezzam in [#47112]\r\n* Fix false __len__ claims in Trainer docstrings (#47131) by @albertvillanova in [#47131]\r\n* processor tests: use tiny Hub repos to reduce CI memory (#47115) by @ydshieh in [#47115]\r\n* [serge] Fix 12 integration tests for model `dac` failing with `output_mismatch` (tensor values differ (6), other (6)) (#47121) by @sergereview[bot] in [#47121]\r\n* Fix CLI compatibility with huggingface_hub 1.22 (#47059) (#47064) by @dhruv7477 in [#47064]\r\n* we want to run the CI in the release branches (#47125) by @tarekziade in [#47125]\r\n* Small improvement (#47128) by @Cyrilvallez in [#47128]\r\n* [Model] Support `use_cache=False` for DeepSeek V4 (#46965) by @kylesayrs in [#46965]\r\n* docs-fix: IMDb dataset link in sequence classification guide (#47062) by @abhishekkapoorx in [#47062]\r\n* Fix AltCLIP text embedding resize test (#47079) by @IMvision12 in [#47079]\r\n* fix mask return-type contract regression and add correctness guard for (#47019) by @kaixuanliu in [#47019]\r\n* Fix save_pretrained with offloading and weight conversions (#47018) by @Cyrilvallez in [#47018]\r\n* Update dev (#47044) by @vasqu in [#47044]\r\n* [`Gemma4`] Update 1 integration test (#47042) by @vasqu in [#47042]\r\n\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @ArthurZucker\r\n    * v5.14.0\r\n* @tarekziade\r\n    * ci: cover xet as well (runtime error) (#47338)\r\n    * Pin kernels to compatible version in benchmark workflow (#47339)\r\n    * Switch mlinter to 0.1.2 (#47172)\r\n    * Make executorch exporter tests always use xnnpack backend (#47201)\r\n    * Remove executorch from all-latest-gpu image + add torch smoke test (#47196)\r\n    * we want to run the CI in the release branches (#47125)\r\n* @remi-or\r\n    * [Nit] Add kernels_fallback_ok kwarg to is_flash_attn_N_available (#47318)\r\n    * [Nit] Add expectations for gemma4 tests on H100 (#47311)\r\n    * [Fix] Remove deprecated argument from `kernels` call (#47100)\r\n    * [Fix] Make DeepGEMM triton fallback more robust (#47126)\r\n    * Fix experts implementation in two spots (#47097)\r\n    * [Fix] Remove old automatic cross attn pattern from output recorders (#47117)\r\n* @ydshieh\r\n    * tests: reduce processor test memory usage by using tiny Hub checkpoints (#47213)\r\n    * Fix flash-attn Docker build broken by setuptools 83 removing pkg_resources (#47251)\r\n    * Fix InputTokensDetails missing cache_write_tokens for openai>=2.34.0 (#47248)\r\n    * Revert \"Trigger a scheduled run\" (#47249)\r\n    * Fix CI read-only cache failures by patching cached_files in conftest (#47043)\r\n    * Trigger a scheduled run (#47209)\r\n    * tests: reduce processor test memory usage and use tiny test assets (#47168)\r\n    * processor tests: use tiny Hub repos to reduce CI memory (#47115)\r\n* @eladsegal\r\n    * Add heterogeneous config support (per-layer configuration) (#45333)\r\n* @eustlb\r\n    * [fix] update integration test values (#47146)\r\n    * revert #47121 (#47144)\r\n* @Ternura143\r\n    * Add TIPSv2 (#46347)\r\n","publishedAt":"2026-07-15T19:02:19.000Z","fetchedAt":"2026-07-15T22:04:25.271Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.14.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/051f819a-512f-4987-9bee-6e2fa2af3db7","r2Key":"releases/70ade92563b5d15dd5d533879a2e7ae59ab8f969018f434e90d8c2dced48e985.png","r2Url":"https://media.releases.sh/releases/70ade92563b5d15dd5d533879a2e7ae59ab8f969018f434e90d8c2dced48e985.png"},{"type":"image","url":"https://github.com/user-attachments/assets/2d9f21e5-05f8-4c36-93ef-22f03c089f52","r2Key":"releases/0fbd4cc138d556ab00592ed450dddeb46aebb8bf7489dce6290bbc3408d2b382.png","r2Url":"https://media.releases.sh/releases/0fbd4cc138d556ab00592ed450dddeb46aebb8bf7489dce6290bbc3408d2b382.png"},{"type":"image","url":"https://github.com/user-attachments/assets/09c0d4da-6c1c-4229-bf02-a512ed435e50","r2Key":"releases/603725e1214e6ca5bd3dcdcd680585fe643edd9c18246cedacae2c876a61f214.png","r2Url":"https://media.releases.sh/releases/603725e1214e6ca5bd3dcdcd680585fe643edd9c18246cedacae2c876a61f214.png"}],"coverageCount":0},{"id":"rel_S7zF2jPItcutR2sVN8Tu2","version":"v5.13.1","type":"feature","title":"Patch release v5.13.1","summary":"Fixes custom model compatibility with the latest vllm release by being more defensive with remap_legacy_layer_types and handling cases where custom code doesn't know about the new linear layer type names. Also fixed a key type assertion in _LazyAutoMapping.register.","titleGenerated":"Transformers v5.13.1 fixes custom model compatibility with new linear layer type names","titleShort":"Custom models no longer break on new linear layer type names","breaking":"none","importance":2,"content":"# Patch release v5.13.1 \r\n\r\nThis patch is focused on enabling `transformers` for the latest release of vllm! \r\n\r\n- Be more defensive with remap_legacy_layer_types for custom models (#47245) from @hmellor \r\n- Fix custom code which doesn't know about the new linear layer type names (#47174) from @hmellor \r\n- Fix case where _LazyAutoMapping.register is passed a str key (#47148) from @hmellor ","publishedAt":"2026-07-11T09:15:36.000Z","fetchedAt":"2026-07-11T13:00:04.595Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.13.1","media":[],"coverageCount":0},{"id":"rel_qXn3w6x6YDl_82Xg1NG9Y","version":"v5.13.0","type":"feature","title":"Release v5.13.0","summary":"This release adds support for seven new model architectures: Kimi 2.5–2.7 (multimodal agentic coding), MiMo-V2-Flash (256K context MoE model), Nemotron 3.5 ASR and Nemotron ASR Streaming (multilingual speech recognition with configurable latency-accuracy tradeoffs), Qwen3 ASR with forced aligner, ZAYA1 (MoE language model), VideoPrism (video understanding encoder), and RADIO (vision foundation model family).","titleGenerated":"Transformers v5.13.0 adds Kimi, MiMo, Nemotron, Qwen3 ASR, ZAYA, VideoPrism, and RADIO models","titleShort":"Seven new model architectures: Kimi 2.5–2.7, MiMo-V2-Flash, Nemotron ASR, Qwen3 ASR, ZAYA, VideoPrism, RADIO","breaking":"minor","importance":3,"content":"# Release v5.13.0\r\n\r\n\r\n## New Model additions\r\n\r\n### KimiK 2.5, 2.6, and 2.7\r\n\r\n<img width=\"1097\" height=\"400\" alt=\"image\" src=\"https://github.com/user-attachments/assets/c24d2232-a9b4-413b-a2c8-58d013b6dfbd\" />\r\n\r\nThis release includes the architecture for Kimi 2.5 which is used by 2.5-2.7:\r\n\r\nKimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in [Kimi K2.5: Visual Agentic Intelligence](https://www.kimi.com/en/blog/kimi-k2-5) and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence).\r\n\r\nKimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming languages (Rust, Go, Python) and domains spanning front-end, DevOps, and performance optimization. The model is capable of transforming simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations with deliberate aesthetic precision.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/kimi_k25)\r\n* Add new model: Kimi2-6 (#45630) by @zucchini-nlp in [#45630](https://github.com/huggingface/transformers/pull/45630)\r\n\r\n### MiMo-V2-Flash\r\n\r\n<img width=\"6900\" height=\"904\" alt=\"image\" src=\"https://github.com/user-attachments/assets/8bd8d5f0-0381-4f8c-8ada-0203e11ff494\" />\r\n\r\n**MiMo-V2-Flash** is a Mixture-of-Experts (MoE) language model developed by the Xiaomi MiMo team. Designed to establish a new balance between long-context modeling capabilities and inference efficiency, the model is built for strong performance in complex reasoning and agentic tasks. Trained on 27T tokens with native 32k sequence lengths, MiMo-V2-Flash seamlessly supports an extended **256K context window** while significantly reducing KV-cache storage compared to standard global attention models.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/mimo_v2_flash)\r\n* Add Xiaomi MiMo-V2 (#45144) by @casinca in [#45144](https://github.com/huggingface/transformers/pull/45144)\r\n\r\n### Nemotron 3.5 ASR\r\n\r\n<img width=\"1632\" height=\"735\" alt=\"image\" src=\"https://github.com/user-attachments/assets/597bbb9c-b046-4e47-b9fd-f242e0a5b04d\" />\r\n\r\nNemotron 3.5 ASR is a 600M-parameter multilingual speech recognition model from NVIDIA, built for high-quality transcription in both low-latency streaming and high-throughput batch settings, with native punctuation and capitalization. For streaming, it offers configurable chunk sizes—80ms, 160ms, 560ms, and 1120ms, letting users trade off latency against accuracy to suit their application. Its cache-aware FastConformer-RNNT architecture is central to this capability: unlike traditional buffered streaming, which repeatedly reprocesses overlapping audio windows, the model processes only each new incoming chunk while reusing cached encoder context from prior chunks. This eliminates redundant computation, significantly improves efficiency, and minimizes end-to-end delay without sacrificing accuracy, making it well suited to real-time transcription workloads.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/nemotron3_5_asr)\r\n* Add Nemotron 3.5 ASR Streaming (#46565) by @eustlb in [#46565](https://github.com/huggingface/transformers/pull/46565)\r\n\r\n### NemotronAsrStreaming\r\n\r\nNemotron ASR Streaming is a 600M-parameter English speech recognition model from NVIDIA, built for high-quality transcription in both low-latency streaming and high-throughput batch settings, with native punctuation and capitalization. For streaming, it offers configurable chunk sizes—80ms, 160ms, 560ms, and 1120ms, letting users trade off latency against accuracy to suit their application. Its cache-aware FastConformer-RNNT architecture is central to this capability: unlike traditional buffered streaming, which repeatedly reprocesses overlapping audio windows, the model processes only each new incoming chunk while reusing cached encoder context from prior chunks. This eliminates redundant computation, significantly improves efficiency, and minimizes end-to-end delay without sacrificing accuracy, making it well suited to real-time transcription workloads.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/nemotron_asr_streaming)\r\n* Add Nemotron ASR Streaming (#46332) by @eustlb in [#46332](https://github.com/huggingface/transformers/pull/46332)\r\n\r\n### Qwen3 ASR\r\n\r\n<img width=\"3646\" height=\"2036\" alt=\"image\" src=\"https://github.com/user-attachments/assets/41ed13e3-a0bf-463a-8473-bc6beb8ebd73\" />\r\n\r\nQwen3 ASR is an automatic speech recognition model from Alibaba's Qwen team that combines a Whisper-style audio encoder with a Qwen3 language model decoder for speech-to-text transcription. The model supports automatic language detection and multilingual transcription.\r\n\r\nA forced aligner model is also included. It can be used to timestamp a provided transcript and its audio. It uses the same audio encoder model with a classification head that predicts a word's length. This model can be used with the transcript from any ASR model (see the example below with Parakeet CTC).\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/qwen3_asr)\r\n* Qwen3 ASR and Forced Aligner (#43838) by @mbtariq82 in [#43838](https://github.com/huggingface/transformers/pull/43838)\r\n\r\n### ZAYA\r\n\r\n<img width=\"1200\" height=\"628\" alt=\"image\" src=\"https://github.com/user-attachments/assets/2935eba8-ab74-455c-9d44-f088636b2785\" />\r\n\r\nZAYA1 is a 760M active / 8.4B total parameter MoE language model trained by Zyphra. It combines Compressed\r\nConvolutional Attention (CCA), a nonlinear ZAYA1 router, and residual scaling.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/zaya)\r\n* [new model] Add Zyphra/ZAYA1-8B (#45862) by @JJJYmmm in [#45862](https://github.com/huggingface/transformers/pull/45862)\r\n\r\n### VideoPrism\r\n\r\nThe VideoPrism model was proposed in the paper [VideoPrism: A Foundational Visual Encoder for Video Understanding](https://huggingface.co/papers/2402.13217) by Google DeepMind ([blog post](https://research.google/blog/videoprism-a-foundational-visual-encoder-for-video-understanding/)).\r\n\r\nVideoPrism is a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. The model is pretrained on a large-scale heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding through global-local distillation of semantic video embeddings and a token shuffling scheme, enabling the model to focus primarily on the video modality while leveraging text associated with videos. VideoPrism achieves state-of-the-art performance on 31 out of 33 video understanding benchmarks across four broad task groups, from web video question answering to computer vision for science.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/videoprism)\r\n* Add Videoprism (#39895) by @MHRDYN7 in [#39895](https://github.com/huggingface/transformers/pull/39895)\r\n\r\n### RADIO\r\n\r\n[RADIO](https://huggingface.co/papers/2312.06709) (Reduce All Domains Into One) is a family of vision foundation models from NVIDIA trained by multi-teacher distillation (e.g. CLIP, DINOv2, SAM) into a single ViT backbone. It produces both an image-level `summary` embedding and dense spatial `features`, and supports variable input resolutions through a Cropped Position Embedding (CPE) patch generator.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/radio)\r\n* Add support for RADIO models (#46425) by @meatybobby in [#46425](https://github.com/huggingface/transformers/pull/46425)\r\n\r\n### MiniCPM3\r\n\r\nMiniCPM3 is the third-generation MiniCPM dense language model from OpenBMB. The 4B variant\r\n([`openbmb/MiniCPM3-4B`](https://huggingface.co/openbmb/MiniCPM3-4B)) outperforms many 7B–9B open\r\nmodels on standard benchmarks while remaining lightweight enough for on-device usage.\r\n\r\nMiniCPM3 combines several architectural ideas:\r\n\r\n- **Multi-head Latent Attention (MLA)** from DeepSeek-V2, which compresses the key/value cache\r\n  into a low-rank latent representation while still using rotary embeddings on a portion of the\r\n  query/key heads.\r\n- A standard SwiGLU MLP (no MoE).\r\n- Three scalar scaling factors that govern signal flow:\r\n  - `scale_emb` — scales input embeddings.\r\n  - `scale_depth / sqrt(num_hidden_layers)` — scales residual connections.\r\n  - `hidden_size / dim_model_base` — scales hidden states before the language model head.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/minicpm3)\r\n* Add MiniCPM3 (#41116) by @bzantium in [#41116](https://github.com/huggingface/transformers/pull/41116)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nA broad set of modeling changes have been made to standardize layer declarations, mask/cache construction, and hybrid-attention handling, making many models cleanly exportable (ONNX, `torch.export`, ExecuTorch) and fullgraph-compilable — users relying on internal modeling APIs may need to update their code accordingly.\r\n* 🚨 Modeling changes for export, compile, and hybrid-attention standardization (#46738) by @IlyasMoutawwakil\r\n\r\nAttention masking for image tokens in Gemma 3/4 models has been fixed to correctly respect sliding window boundaries in local layers, which changes model behavior and may affect reproducibility of previous results.\r\n* 🚨 [gemma 3/4] Fix bidirectional attention masking crossing sliding window boundaries (#46850) by @douglas-reid\r\n\r\nThe Expert Parallelism (EP) router contract has been corrected across many models and FP8 scale format handling has been fixed, requiring users of EP or FP8 quantization with affected models to verify their configurations and potentially update conversion mappings.\r\n* 🚨 EP: fix EP router contract for many models + honor FP8 scale format (#46818) by @IlyasMoutawwakil\r\n\r\nThe `Kernels` integration has been synced to the latest version, which includes a breaking change where model-type repositories are no longer accepted by the kernels interface — users must migrate to the updated kernel repository format as shown in the updated tests.\r\n* :rotating_light: [`Kernels`] Sync to latest version (#46039) by @vasqu\r\n\r\n\r\n## HfExporters: Native, Unified export for PyTorch / ONNX / ExecuTorch\r\n\r\n<img alt=\"thumbnail\" src=\"https://github.com/user-attachments/assets/3ba5751b-c99e-4945-b5a8-b2b29231f5df\" />\r\n\r\nA native, in-Transformers export pipeline — one base class (`HfExporter`), three subclasses for the runtimes we care about, one unified API:\r\n\r\n| Exporter | Output | Runtime |\r\n|---|---|---|\r\n| `DynamoExporter` | `ExportedProgram` | Any PyTorch runtime, AOT compilation |\r\n| `OnnxExporter` | `ONNXProgram` | Any ONNX runtime (ORT, TensorRT, OpenVINO, …) |\r\n| `ExecutorchExporter` | `ExecutorchProgramManager` | Mobile and edge (ExecuTorch) |\r\n\r\nSame call shape across all three. Dynamic shapes by default. Generation-style models split automatically into prefill + decode (+ vision/audio sub-encoders for VLMs).\r\n\r\n```python\r\nfrom transformers import AutoModelForMaskedLM, AutoTokenizer\r\nfrom transformers.exporters import OnnxExporter, OnnxConfig\r\n\r\nmodel_id = \"hf-internal-testing/tiny-random-BertForMaskedLM\"\r\ntokenizer = AutoTokenizer.from_pretrained(model_id)\r\nmodel = AutoModelForMaskedLM.from_pretrained(model_id).eval()\r\ninputs = tokenizer([\"Hello, my dog is cute\"] * 2, return_tensors=\"pt\")\r\nonnx_program = OnnxExporter().export(model, inputs, config=OnnxConfig(dynamic=True))\r\n\r\nnew_input = tokenizer(\"Hello, my cat is so adorable!\", return_tensors=\"pt\")\r\ntorch.testing.assert_close(\r\n    onnx_program.call_reference(**new_input)[0],   # numpy reference\r\n    onnx_program(**new_input)[0],                  # onnxruntime\r\n    rtol=1e-4, atol=1e-4,\r\n)\r\n```\r\n\r\nSwap one line for another runtime — `DynamoExporter()` / `DynamoConfig` or `ExecutorchExporter()` / `ExecutorchConfig(backend=...)`.\r\n\r\nFor generative models the prefill/decode split is captured automatically:\r\n\r\n```python\r\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\r\nfrom transformers.exporters import OnnxExporter, OnnxConfig\r\n\r\nmodel_id = \"hf-internal-testing/tiny-random-LlamaForCausalLM\"\r\ntokenizer = AutoTokenizer.from_pretrained(model_id)\r\nmodel = AutoModelForCausalLM.from_pretrained(model_id).eval()\r\ninputs = tokenizer([\"Hello, my dog is cute\"] * 2, return_tensors=\"pt\")\r\n\r\nartifacts = OnnxExporter().export_for_generation(model, inputs, config=OnnxConfig(dynamic=True))\r\n# {\"prefill\": ONNXProgram, \"decode\": ONNXProgram}\r\n# For VLMs: also vision_encoder, audio_encoder, multi_modal_projector, language_model, lm_head\r\n```\r\n\r\n\r\n## Kernels\r\n\r\n**Kernels:** Fixed a silent SDPA math-kernel fallback for GQA models with `head_dim > 256` (e.g., Gemma4) that caused O(S²) memory materialization, and resolved a regression where `use_kernels=True` failed to apply kernel mappings. Additional improvements include lazy loading of the default kernel mapping to prevent import failures with incompatible kernel versions, ROCm routing to AITER Triton kernels for AMD GPUs, GB10/SM121 Hub-kernel support for Qwen3.6 Gated DeltaNet, and expanded documentation for the kernel API.\r\n\r\n\r\n* Fix silent SDPA math-kernel fallback for GQA when key/value head_dim > 256 or differ (#46960) by @Butterfingrz in [#46960]\r\n* [docs] AITER kernels (#46871) by @stevhliu in [#46871]\r\n* Documentation for the kernel API (#46754) by @michaelbenayoun in [#46754]\r\n* update kernels-community/aiter-rope version (#46810) by @Abdennacer-Badaoui in [#46810]\r\n* Add GB10/SM121 Hub-kernel path for Qwen3.6 Gated DeltaNet (#46423) by @AzeezIsh in [#46423]\r\n* [`Kernels`] Trigger proper kernelization on `use_kernels=True` (#46755) by @vasqu in [#46755]\r\n* Lazily build the default kernel mapping to decouple `kernels` from normal transformers usage (#46681) by @jiqing-feng in [#46681]\r\n* Add some AITER kernel routing for ROCm (#46268) by @Abdennacer-Badaoui in [#46268]\r\n* fix: position ids does not exist in upstream rotary kernel (#46619) by @NanoCode012 in [#46619]\r\n* docs(zh): add Chinese translation of kernels.md (#46621) by @shoushinya123 in [#46621]\r\n\r\n\r\n## Generation\r\n\r\nSeveral generation bugs were fixed, including Mamba2 chunked-prefill and speculative decoding for hybrid models (Zamba2, Nemotron-H, Bamba, FalconH1, GraniteMoeHybrid), beam search for Mamba models, prompt lookup decoding crashes with no EOS token, and incorrect stateful model handling for LFM2. Additional improvements include reduced unnecessary generation warnings, a fix for continuous batching output mutation, and a new option to keep input tensors on CPU during generation to avoid retracing on Neuron/TPU devices.\r\n\r\n\r\n* Fix Mamba2 chunked-prefill / speculative decoding for Zamba2, Nemotron-H, Bamba, FalconH1 and GraniteMoeHybrid (#46741) by @Sunt-ing in [#46741]\r\n* Remove some unnecessary generate warnings (#46955) by @Cyrilvallez in [#46955]\r\n* Reject assisted generation for LFM2 and LFM2-MoE (set _is_stateful) (#46937) by @Sunt-ing in [#46937]\r\n* Fix beam search for mamba models (#46819) by @Cyrilvallez in [#46819]\r\n* Fix prompt lookup decoding crash when no EOS token is configured (#46790) by @Sunt-ing in [#46790]\r\n* [Continuous Batching] Snapshot generation outputs without mutating request state (#46670) by @Incheonkirin in [#46670]\r\n* [docs] keep generation tensors on cpu (#46675) by @stevhliu in [#46675]\r\n* feat(generation): allow user to keep input tensors on cpu (#46590) by @dacorvo in [#46590]\r\n\r\n\r\n## Attention\r\n\r\nSeveral attention-related bugs were fixed in this release, including silent SDPA math-kernel fallbacks for GQA with large head dimensions, broken Flash Attention with `StaticCache`, incorrect causal masking in Xcodec2, a cross-attention reshape regression in Blip2, and eager GQA support in Evolla. Accelerate hook handling was also corrected for models using linear attention to prevent silently wrong results during offloading.\r\n\r\n\r\n* Fix accelerate hooks for all models using linear attention (#46978) by @Cyrilvallez in [#46978]\r\n* Fix Xcodec2 attention to be non-causal. (#46963) by @ebezzam in [#46963]\r\n* Fix flash attention with StaticCache (#46914) by @Cyrilvallez in [#46914]\r\n* Fix Evolla eager attention for the GQA text decoder (#46860) by @jiqing-feng in [#46860]\r\n* [docs] metal flash attention (#46349) by @stevhliu in [#46349]\r\n* [`Blip2`] Fix cross attention reshape (#46695) by @vasqu in [#46695]\r\n\r\n\r\n## Cache\r\n\r\nCache APIs were improved by consolidating redundant getters into a cleaner `get_max_length` method and updating documentation accordingly. Several bug fixes were also applied, including correcting mask generation beyond sliding windows, fixing a dimension issue in cumulative length tracking, resolving device mismatches in offloaded cache for hybrid models, and fixing crashes when loading trust_remote_code models from symlinked local caches.\r\n\r\n\r\n* [docs] update cache apis (#46892) by @stevhliu in [#46892]\r\n* Rework some old cache getters/properties (#46862) by @Cyrilvallez in [#46862]\r\n* Fix expanded dim in the cache's cumulative length (#46856) by @Cyrilvallez in [#46856]\r\n* Fix mask when generating beyond sliding window (#46839) by @zucchini-nlp in [#46839]\r\n* Fix offloaded cache device mismatch on hybrid models (#46748) by @Sunt-ing in [#46748]\r\n* Fix dynamic module symlinked cache on trust_remote_code models (#46618) by @ldkhang1201 in [#46618]\r\n\r\n\r\n## Serve\r\n\r\nSeveral fixes and improvements were made to the Serve functionality, including lazy imports to prevent CLI crashes when the optional `serve` extra is not installed, a fix for dropped attributes during serialization of subclassed Pydantic models, and added documentation for the kernel API.\r\n\r\n\r\n* fix(cli/serve): import serve handlers lazily so the CLI works without the `serve` extra (#46473) by @<NOT FOUND> in [#46473]\r\n* [Fix] Serve drops some attributes at serialization (#46680) by @remi-or in [#46680]\r\n* Reduce per_page from 100 to 50 in GitHub API calls to avoid server errors (#46678) by @ydshieh in [#46678]\r\n\r\n\r\n## Quantization\r\n\r\nFixed dtype casting bugs in Gemma4's vision and audio multimodal embedders when using BitsAndBytes quantization, where inputs were incorrectly cast to integer storage dtypes (`uint8`/`int8`) instead of the actual compute dtype. Also corrected FP8 quantization to round block scales before quantizing weights, ensuring dequantization produces correct values for `ue8m0` (DeepSeek-V4 style) format.\r\n\r\n\r\n* [Gemma4] Fix dtype casting for quantized vision/audio embedders (#46933) by @sharmax-vikas in [#46933]\r\n* Fix dtype casting for quantized multimodal embedders (#46904) by @praful-srinivasan-027 in [#46904]\r\n* Round the ue8m0 FP8 scale before quantizing so dequant matches the stored inverse (#46763) by @Incheonkirin in [#46763]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Update workflow callers to use `transformers-ci` (#47040) by @ydshieh in [#47040]\r\n* Add HunYuan VL model (#46417) by @Mi-Jiazhi in [#46417]\r\n* Add tiny_model_id support to ProcessorTesterMixin for memory-sensitive tests (#47005) by @ydshieh in [#47005]\r\n* chore(linter): add TRF018 modeling rule (#46259) by @tarekziade in [#46259]\r\n* [PoC] HF exporters (#41992) by @IlyasMoutawwakil in [#41992]\r\n* TST Skip PEFT tests if PEFT version is too low (#47027) by @BenjaminBossan in [#47027]\r\n* CI Add PEFT integration tests (#47021) by @BenjaminBossan in [#47021]\r\n* [glm-mode-dsa] Indexer uses interleaved rope (#46842) by @pcuenca in [#46842]\r\n* Use standard arg names in Mllama (#46977) by @zucchini-nlp in [#46977]\r\n* Bump min peft 0.19.1 remove weight conversion duplicate code (#46442) by @BenjaminBossan in [#46442]\r\n* Raise a loud error for missing prefix (#46980) by @Rocketknight1 in [#46980]\r\n* Fix typo in Qwen3 ASR no_split_module (#47002) by @ebezzam in [#47002]\r\n* only in the original repo (#46982) by @tarekziade in [#46982]\r\n* Fix typos in Gemma 4 Assistant documentation (#46975) by @RaunaqDavidNath in [#46975]\r\n* the CI status should be a comment (#46976) by @tarekziade in [#46976]\r\n* QwenVL model conversion (#46881) by @zucchini-nlp in [#46881]\r\n* Remove default dtype in FusedRMSNormGated modules (#46953) by @Cyrilvallez in [#46953]\r\n* FIX PEFT test changed error type (#46959) by @BenjaminBossan in [#46959]\r\n* Fix path traversal via vocab-file arguments in tokenizer_config.json (#46279) by @LinZiyuu in [#46279]\r\n* docs(conditional_detr): fix num_queries default in docstring (100 -> 300) (#46939) by @Kropiunig in [#46939]\r\n* Use common floats_list method for feature extractor tests. (#46956) by @ebezzam in [#46956]\r\n* Fix RT-DETR indexing error when num_feature_levels exceeds backbone o… (#46833) by @c1prk in [#46833]\r\n* Fix Florence2 training-loss double-shift (same pattern as Moonshine #… (#46898) by @sharmax-vikas in [#46898]\r\n* [Olmo3] different RoPE per layer type (#46911) by @zucchini-nlp in [#46911]\r\n* Use inspect.getsource instead of open() for source-reading in _can_set_*_implementation (#46207) by @rasmi in [#46207]\r\n* Don't pin the gated delta net norm to `cuda:0` with a hardcoded device (#46817) by @Sunt-ing in [#46817]\r\n* Fix auto-mappings registration for remote code & fixes a few custom code issues (#46876) by @Cyrilvallez in [#46876]\r\n* Fix broken internal documentation links (#46945) by @sezer-muhammed in [#46945]\r\n* Insert a Grafana badge in the PR (#46774) by @tarekziade in [#46774]\r\n* [NemotronAsrStreaming] fix pipeline (#46870) by @eustlb in [#46870]\r\n* [NemotronAsrStreaming] processor without modular (#46865) by @eustlb in [#46865]\r\n* [`Dia`] Fix docs (#46923) by @vasqu in [#46923]\r\n* [Docs] Fix full disk offloading docs (#46905) by @kylesayrs in [#46905]\r\n* [CB] Changes to increase max_batch_tokens (#46712) by @remi-or in [#46712]\r\n* Redirect to diffusers pipe in docs for experimental features (#46875) by @zucchini-nlp in [#46875]\r\n* Install in docker (#46910) by @ydshieh in [#46910]\r\n* [CI] Use pre-computed `_OLD_MODELS` in `test_new_models_require_torchvision_backend` (#46882) by @ydshieh in [#46882]\r\n* call transformers-ci in a nightly run (#46811) by @tarekziade in [#46811]\r\n* [docs] full disk offloading (#46893) by @stevhliu in [#46893]\r\n* TST Run fast PEFT tests in normal CI (#45679) by @BenjaminBossan in [#45679]\r\n* nemotron_asr_streaming: set _supports_flex_attn to False (#46878) by @kaixuanliu in [#46878]\r\n* Add native masked MSE loss for Sapiens2ForPoseEstimation (#46764) by @Sainava in [#46764]\r\n* blip 2 fix (#46816) by @itazap in [#46816]\r\n* Use meshgrid for brevity (#46861) by @zucchini-nlp in [#46861]\r\n* Add xcodec2 model (#44178) by @ebezzam in [#44178]\r\n* Prevent auto-class from being modified for all models (#46844) by @zucchini-nlp in [#46844]\r\n* Add Spanish translation of the torch.compile page (#46852) by @delcenjo in [#46852]\r\n* docs: Update NeMo AutoModel doc examples (#46857) by @adil-a in [#46857]\r\n* [docs] distributed training (#44420) by @stevhliu in [#44420]\r\n* [docs] require trust_remote_code for custom_generate (#46677) by @stevhliu in [#46677]\r\n* add distributed config (#46705) by @3outeille in [#46705]\r\n* [Offloading] [Bugfix] Fix disk offloading of models with explicit tensor dtypes (#46849) by @kylesayrs in [#46849]\r\n* Streamable chat parsing (#45847) by @Rocketknight1 in [#45847]\r\n* Fix BitNet packed-weight unpacking dtype (`F.linear` dtype mismatch) (#46808) by @jiqing-feng in [#46808]\r\n* Fix typos in code (#46579) by @cyyever in [#46579]\r\n* Fix Moonshine training-loss double-shift (train against labels, not labels[..., 1:]) (#46784) by @Incheonkirin in [#46784]\r\n* [CB] Fix issues with FA read / writes (#46765) by @remi-or in [#46765]\r\n* Switch decorator order (#46853) by @Cyrilvallez in [#46853]\r\n* docs(trainer): add JIT checkpointing to trainer recipes (#46826) by @efazal in [#46826]\r\n* Import diffusion_gemma in models init (#46841) by @boringcrypto in [#46841]\r\n* [skills] help your agent get started (#45732) by @stevhliu in [#45732]\r\n* Fix use_cache with seq_len > 1 ( #46032) (#46084) by @Ramshankar07 in [#46084]\r\n* [Offloading] Support full disk offloading (#46749) by @kylesayrs in [#46749]\r\n* fix: raise `ValueError` for empty conversation in `apply_chat_template` (#46753) by @sharmax-vikas in [#46753]\r\n* Fix VideoPrismForVideoClassification returning last_hidden_state as h… (#46830) by @sharmax-vikas in [#46830]\r\n* Avoid NumPy 2.0 `__array__` copy-keyword deprecation in `create_mm_token_type_ids` (#46827) by @qgallouedec in [#46827]\r\n* docs: update apple silicon doc with safetensors `0.8.0` benefits (#46744) by @McPatate in [#46744]\r\n* [`CB`] Add FA2 to the fast path (#46729) by @vasqu in [#46729]\r\n* Fix flex_attention block mask creation when `get_seq_length` returns a tensor (#46802) by @jiqing-feng in [#46802]\r\n* Fix left-padding token selection in `BioGptForSequenceClassification` (#46782) by @Sunt-ing in [#46782]\r\n* Fix broken internal links in model documentation (#46807) by @ShamSaleem in [#46807]\r\n* DiffusionGemma: mask layout and CI (#46654) by @zucchini-nlp in [#46654]\r\n* Use cached added-token dicts in per-token decode loops (#46535) by @ishan-1010 in [#46535]\r\n* fix another flaky test (#46767) by @zucchini-nlp in [#46767]\r\n* Fix secondary rate limit when downloading artifacts in slack report (#46796) by @ydshieh in [#46796]\r\n* docs: move SmolLM3 to Text models category in _toctree.yml (#46770) by @yyouretoast in [#46770]\r\n* Fix several bugs in `cache_implementation=static` (#46446) by @dacorvo in [#46446]\r\n* [CI] Fix artifact download path in self-comment-ci workflow (#46769) by @ydshieh in [#46769]\r\n* fixes per head minimaxm3 (#46719) by @ArthurZucker in [#46719]\r\n* [`CI`] Fix some failures introduced by myself :grimacing:  (#46751) by @vasqu in [#46751]\r\n* Fix regression in ProcessorMixin._load_tokenizer_from_pretrained for tokenizers at root (#46592) by @<NOT FOUND> in [#46592]\r\n* fix(aria): use math.ceil in get_number_of_image_patches to match actual patch count (#46732) by @arnavkewalram in [#46732]\r\n* Return logits from semantic segmentation post-process (#46163) by @guarin in [#46163]\r\n* Fall back to the for-loop grouped_mm on CPU (#46743) by @Sunt-ing in [#46743]\r\n* Kernelize refactor (#46520) by @michaelbenayoun in [#46520]\r\n* ci: add comment explaining why secrets are not inherited in security gate (#46750) by @ydshieh in [#46750]\r\n* ci: trigger PR CI on ci-* branches (#46746) by @ydshieh in [#46746]\r\n* finegrained v3 (#46742) by @IlyasMoutawwakil in [#46742]\r\n* Improve AutoImageProcessor error for unavailable backends (#46727) by @sisaman in [#46727]\r\n* skip decorators must appear after @parameterized.expand in pytest (#46737) by @rasmi in [#46737]\r\n* [RecurrentGemma] Support attn_implementation dispatch (#46320) by @YangKai0616 in [#46320]\r\n* [docs] clarify initialization module usage (#46698) by @stevhliu in [#46698]\r\n* feat: bump safetensors to `0.8.0` (#46523) by @McPatate in [#46523]\r\n* ci: disable CircleCI by replacing config with no-op (#46721) by @ydshieh in [#46721]\r\n* [CB] Fix offloading (#46587) by @remi-or in [#46587]\r\n* [`Templates`] Update members (#46720) by @vasqu in [#46720]\r\n* feat[vLLM x v5]: Expose max_source_positions on VibeVoiceAsrConfig (#46472) by @harshaljanjani in [#46472]\r\n* Laguna: support per-element output gating (#46690) by @joerowell in [#46690]\r\n* ci: grant pull-requests:write to the security gate caller (#46715) by @ydshieh in [#46715]\r\n* Multi-gpu loading when the whole backbone is tied (#46625) by @zucchini-nlp in [#46625]\r\n* Delete docstring if same as in auto-doc (#46284) by @zucchini-nlp in [#46284]\r\n* Update GLM-5.2 docs (#46703) by @Dovis01 in [#46703]\r\n* add conversion scripts for EUPE (#46691) by @molbap in [#46691]\r\n* [docs] compile level and batch/scheduling limits (#46676) by @stevhliu in [#46676]\r\n* [blip_2] Support attn_implementation dispatch (#46401) by @YangKai0616 in [#46401]\r\n* [CTRL] Support attn_implementation dispatch (#46073) by @YangKai0616 in [#46073]\r\n* Lfm2: also thread `seq_idx` through ShortConv.slow_forward (non-fast-path) (#46633) by @ChangyiYang in [#46633]\r\n* feat(pipelines): accept numpy arrays and tensors in ImageClassificationPipeline (#39607) (#46573) by @kamran-nizamani in [#46573]\r\n* Smovlm: pad videos up to max frames (#46662) by @zucchini-nlp in [#46662]\r\n* mistral common backend fix (#46667) by @itazap in [#46667]\r\n* [pr template] update (#46606) by @stevhliu in [#46606]\r\n* Fix AttributeError in auto_factory when model_class lacks config_class (#46669) by @atharv1945 in [#46669]\r\n* [CB] Slice logits inside the model (#46660) by @remi-or in [#46660]\r\n* ci: add NO_COLOR=1 to suppress ANSI color codes in CI output (#46659) by @ydshieh in [#46659]\r\n* Fix dynamic RoPE not resetting inv_freq when layer_type is None (#46624) by @Incheonkirin in [#46624]\r\n* Better processing tests (#46374) by @zucchini-nlp in [#46374]\r\n* ci: add merge_group trigger to pr-ci-caller.yml (#46668) by @ydshieh in [#46668]\r\n* skip invalid quant_cache test for nemotron_h (#46368) by @kaixuanliu in [#46368]\r\n* Revert \"Disable PR CI workflow for PRs from forked repo. during the weekend\" (#46652) by @ydshieh in [#46652]\r\n* [CB] Fix seqlens and use TypedDict (#46593) by @remi-or in [#46593]\r\n* Disable PR CI workflow for PRs from forked repo. during the weekend (#46609) by @ydshieh in [#46609]\r\n* Update post release (#46608) by @vasqu in [#46608]\r\n* Fix `peft` lower bound (#46605) by @hmellor in [#46605]\r\n* Fix docstring formatting issues causing Sphinx autodoc warnings (#46596) by @kurtmckee in [#46596]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @ydshieh\r\n    * Update workflow callers to use `transformers-ci` (#47040)\r\n    * Add tiny_model_id support to ProcessorTesterMixin for memory-sensitive tests (#47005)\r\n    * Install in docker (#46910)\r\n    * [CI] Use pre-computed `_OLD_MODELS` in `test_new_models_require_torchvision_backend` (#46882)\r\n    * Fix secondary rate limit when downloading artifacts in slack report (#46796)\r\n    * [CI] Fix artifact download path in self-comment-ci workflow (#46769)\r\n    * ci: add comment explaining why secrets are not inherited in security gate (#46750)\r\n    * ci: trigger PR CI on ci-* branches (#46746)\r\n    * ci: disable CircleCI by replacing config with no-op (#46721)\r\n    * ci: grant pull-requests:write to the security gate caller (#46715)\r\n    * Reduce per_page from 100 to 50 in GitHub API calls to avoid server errors (#46678)\r\n    * ci: add NO_COLOR=1 to suppress ANSI color codes in CI output (#46659)\r\n    * ci: add merge_group trigger to pr-ci-caller.yml (#46668)\r\n    * Revert \"Disable PR CI workflow for PRs from forked repo. during the weekend\" (#46652)\r\n    * Disable PR CI workflow for PRs from forked repo. during the weekend (#46609)\r\n* @Mi-Jiazhi\r\n    * Add HunYuan VL model (#46417)\r\n* @tarekziade\r\n    * chore(linter): add TRF018 modeling rule (#46259)\r\n    * only in the original repo (#46982)\r\n    * the CI status should be a comment (#46976)\r\n    * Insert a Grafana badge in the PR (#46774)\r\n    * call transformers-ci in a nightly run (#46811)\r\n* @casinca\r\n    * Add Xiaomi MiMo-V2 (#45144)\r\n* @JJJYmmm\r\n    * [new model] Add Zyphra/ZAYA1-8B (#45862)\r\n* @ebezzam\r\n    * Fix typo in Qwen3 ASR no_split_module (#47002)\r\n    * Fix Xcodec2 attention to be non-causal. (#46963)\r\n    * Use common floats_list method for feature extractor tests. (#46956)\r\n    * Add xcodec2 model (#44178)\r\n* @meatybobby\r\n    * Add support for RADIO models (#46425)\r\n* @douglas-reid\r\n    * 🚨 [gemma 3/4] Fix bidirectional attention masking crossing sliding window boundaries (#46850)\r\n* @Sunt-ing\r\n    * Fix Mamba2 chunked-prefill / speculative decoding for Zamba2, Nemotron-H, Bamba, FalconH1 and GraniteMoeHybrid (#46741)\r\n    * Reject assisted generation for LFM2 and LFM2-MoE (set _is_stateful) (#46937)\r\n    * Don't pin the gated delta net norm to `cuda:0` with a hardcoded device (#46817)\r\n    * Fix prompt lookup decoding crash when no EOS token is configured (#46790)\r\n    * Fix left-padding token selection in `BioGptForSequenceClassification` (#46782)\r\n    * Fix offloaded cache device mismatch on hybrid models (#46748)\r\n    * Fall back to the for-loop grouped_mm on CPU (#46743)\r\n* @eustlb\r\n    * Add Nemotron 3.5 ASR Streaming (#46565)\r\n    * [NemotronAsrStreaming] fix pipeline (#46870)\r\n    * [NemotronAsrStreaming] processor without modular (#46865)\r\n    * Add Nemotron ASR Streaming (#46332)\r\n    * [fix] enable base64 str audio in load_audio (#46694)\r\n* @vasqu\r\n    * [`Dia`] Fix docs (#46923)\r\n    * [`CB`] Add FA2 to the fast path (#46729)\r\n    * [`Kernels`] Trigger proper kernelization on `use_kernels=True` (#46755)\r\n    * [`CI`] Fix some failures introduced by myself :grimacing:  (#46751)\r\n    * :rotating_light: [`Kernels`] Sync to latest version (#46039)\r\n    * [`Templates`] Update members (#46720)\r\n    * [`Blip2`] Fix cross attention reshape (#46695)\r\n    * Update post release (#46608)\r\n* @mbtariq82\r\n    * Qwen3 ASR and Forced Aligner (#43838)\r\n* @remi-or\r\n    * [CB] Changes to increase max_batch_tokens (#46712)\r\n    * [CB] Fix issues with FA read / writes (#46765)\r\n    * [CB] Fix offloading (#46587)\r\n    * [Fix] Serve drops some attributes at serialization (#46680)\r\n    * [CB] Slice logits inside the model (#46660)\r\n    * [CB] Fix seqlens and use TypedDict (#46593)\r\n* @jiqing-feng\r\n    * Fix BitNet packed-weight unpacking dtype (`F.linear` dtype mismatch) (#46808)\r\n    * Fix Evolla eager attention for the GQA text decoder (#46860)\r\n    * Fix flex_attention block mask creation when `get_seq_length` returns a tensor (#46802)\r\n    * Lazily build the default kernel mapping to decouple `kernels` from normal transformers usage (#46681)\r\n* @bzantium\r\n    * Add MiniCPM3 (#41116)\r\n* @MHRDYN7\r\n    * Add Videoprism (#39895)\r\n* @YangKai0616\r\n    * [RecurrentGemma] Support attn_implementation dispatch (#46320)\r\n    * [blip_2] Support attn_implementation dispatch (#46401)\r\n    * [CTRL] Support attn_implementation dispatch (#46073)","publishedAt":"2026-07-03T16:06:27.000Z","fetchedAt":"2026-07-03T19:06:21.638Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.13.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/c24d2232-a9b4-413b-a2c8-58d013b6dfbd","r2Key":"releases/6dc181efeabd8497b3f0a454e8beb9ac767dbc71ee2b80a68aed454a070d5642.png","r2Url":"https://media.releases.sh/releases/6dc181efeabd8497b3f0a454e8beb9ac767dbc71ee2b80a68aed454a070d5642.png"},{"type":"image","url":"https://github.com/user-attachments/assets/8bd8d5f0-0381-4f8c-8ada-0203e11ff494","r2Key":"releases/870c22be97af64cd80afe311ab1eccb589d70631ab60f7cdbffbf69cbd4beea3.png","r2Url":"https://media.releases.sh/releases/870c22be97af64cd80afe311ab1eccb589d70631ab60f7cdbffbf69cbd4beea3.png"},{"type":"image","url":"https://github.com/user-attachments/assets/597bbb9c-b046-4e47-b9fd-f242e0a5b04d","r2Key":"releases/28e0e32c5accd365cba8ead4ee136cc364c10ff268d2647ccd194316374893e4.png","r2Url":"https://media.releases.sh/releases/28e0e32c5accd365cba8ead4ee136cc364c10ff268d2647ccd194316374893e4.png"},{"type":"image","url":"https://github.com/user-attachments/assets/41ed13e3-a0bf-463a-8473-bc6beb8ebd73","r2Key":"releases/7c245f781bfcad2dd9473dfead020cdfa806bfbd7a7886e5b933cc29e0483d74.png","r2Url":"https://media.releases.sh/releases/7c245f781bfcad2dd9473dfead020cdfa806bfbd7a7886e5b933cc29e0483d74.png"},{"type":"image","url":"https://github.com/user-attachments/assets/2935eba8-ab74-455c-9d44-f088636b2785","r2Key":"releases/34f0dbd7392645ba46374f0c3a14ba3ee6cfa104b75bd4151bccd97bdfea4114.png","r2Url":"https://media.releases.sh/releases/34f0dbd7392645ba46374f0c3a14ba3ee6cfa104b75bd4151bccd97bdfea4114.png"},{"type":"image","url":"https://github.com/user-attachments/assets/3ba5751b-c99e-4945-b5a8-b2b29231f5df"}],"coverageCount":0},{"id":"rel_O2QB0gl5x5yOZjwM2dbqN","version":"v5.12.1","type":"feature","title":"Patch release v5.12.1","summary":"Fixed mistral tokenizer resolution when `mistral-common` is installed and updated the lower bound for PEFT. This is similar to v5.10.3 minus fixes already in the main release.","titleGenerated":"Transformers v5.12.1 fixes mistral tokenizer resolution","titleShort":"Mistral tokenizer resolution fixed","breaking":"unknown","importance":null,"content":"# Patch release v5.12.1\r\nUpdated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when `mistral-common` is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - vLLM will first target 5.10.3 :hugs: \r\n\r\n* Fix `peft` lower bound #46605 by @hmellor (#46605)\r\n* mistral common backend fix #46667 by @itazap (#46667)\r\n\r\n\r\n**Full Changelog**: https://github.com/huggingface/transformers/compare/v5.12.0...v5.12.1","publishedAt":"2026-06-15T17:29:59.000Z","fetchedAt":"2026-06-15T19:03:34.958Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.12.1","media":[],"coverageCount":0},{"id":"rel_OfRySh6X1PDGNt60ukJJL","version":"v5.10.3","type":"feature","title":"Patch release v5.10.3","summary":"This patch release fixes several regressions introduced by previous changes, including issues with {image/video/audio}_token_ids in ProcessorMixin, InternVL models, and offsets in processing. It also addresses a regression in the Mistral common backend and updates the `peft` lower bound.","titleGenerated":"Transformers v5.10.3 fixes several regressions and processing issues","titleShort":"Several regressions and processing issues fixed","breaking":"unknown","importance":null,"content":"# Patch release v5.10.3\r\nA few fixes needed for vLLM to sync with transformers :hugs: \r\n\r\n* [fix] regression introduced by #45534 #46456 by @eustlb (#46456)\r\n* Fix {image/video/audio}_token_ids in ProcessorMixin #46500 by @hmellor (#46500)\r\n* Fix InternVL models #46524 by @hmellor (#46524)\r\n* Fix the offsets in processing #46525 by @zucchini-nlp (#46525)\r\n* Fix `peft` lower bound #46605 by @hmellor (#46605)\r\n* mistral common backend fix #46667 by @itazap (#46667)\r\n\r\n\r\n**Full Changelog**: https://github.com/huggingface/transformers/compare/v5.10.2...v5.10.3","publishedAt":"2026-06-15T17:29:39.000Z","fetchedAt":"2026-06-15T19:03:34.958Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.10.3","media":[],"coverageCount":0},{"id":"rel_9ekbqtdxhUjzd74LbsJAI","version":"v5.12.0","type":"feature","title":"Release v5.12.0","summary":"This release introduces the MiniMax-M3-VL vision-language model, the PP-OCRv6 OCR system, and the Parakeet-RNNT model for speech processing. Several bug fixes and improvements were also made, including changes to CI, stop string matching, and model documentation.","titleGenerated":"Transformers v5.12.0 adds MiniMax-M3-VL, PP-OCRv6, and Parakeet-RNNT models","titleShort":"New models: MiniMax-M3-VL, PP-OCRv6, Parakeet-RNNT","breaking":"unknown","importance":null,"content":"# Release v5.12.0\r\n\r\n\r\n## New Model additions\r\n\r\n### MiniMax-M3-VL\r\n\r\n<img width=\"886\" height=\"583\" alt=\"image\" src=\"https://github.com/user-attachments/assets/ae9dd96f-6877-4531-a06b-a756686f24e5\" />\r\n\r\nMiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone. It uses a mixed dense/sparse Mixture-of-Experts decoder with SwiGLU-OAI gated experts and a lightning indexer for block-sparse attention. The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/minimax_m3_vl)\r\n* Add minimax m3vl (#46600) by @ArthurZucker in [#46600](https://github.com/huggingface/transformers/pull/46600)\r\n\r\n\r\n### PP-OCRv6: update documentation and slow tests (#46576)\r\n\r\n<img width=\"3840\" height=\"1494\" alt=\"image\" src=\"https://github.com/user-attachments/assets/e62284ec-78bf-49cb-8aa2-deccc665372f\" />\r\n\r\nThe official weights for PP-OCRv6 are out: PP-OCRv6 is a lightweight OCR system that combines architectural innovation with data-centric optimization. It redesigns the backbone, detection neck, and recognition neck around a unified MetaFormer-style building block with structural reparameterization. Three model tiers (medium, small, tiny) share the same block primitives, covering deployment scenarios from server to edge.\r\n\r\n* PP-OCRv6: update documentation and slow tests (#46576) by @ zhang-prog\r\n\r\n\r\n### Add Parakeet-RNNT (#46331)\r\n\r\nParakeetForRNNT: a Fast Conformer Encoder + an RNN-T (RNN Transducer) decoder\r\n\r\n- RNN-T Decoder: Standard neural transducer:\r\n    - LSTM prediction network maintains language context across token predictions.\r\n       - Joint network combines encoder and decoder outputs.\r\n       - Greedy transducer decoding for inference: a blank emission advances the encoder frame by one, a non-blank emission stays on the same frame.\r\n\r\n* Add Parakeet-RNNT (#46331) by @eustlb\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* [CI] don't export OTELs within the tests (#46602) by @tarekziade in [#46602]\r\n* [CI] capture checkers output in OTEL (#46601) by @tarekziade in [#46601]\r\n* Lfm2: thread `seq_idx` through ShortConv for packed/varlen inputs (#46588) by @ChangyiYang in [#46588]\r\n* put output_hidden_states into filter_output_hidden_states (#46422) by @molbap in [#46422]\r\n* a11 for checkers (#46599) by @tarekziade in [#46599]\r\n* Fix stop string matching for byte-fragment tokens (#46530) by @Incheonkirin in [#46530]\r\n* [DiffusionGemma] better docs and links (#46569) by @gante in [#46569]\r\n* Require `trust_remote_code` to run a local-directory `custom_generate` (#46483) by @LinZiyuu in [#46483]\r\n* Fix torchaudio version not tied to torch version in docker file  (#46594) by @ydshieh in [#46594]\r\n* [CI] Enable PR CI for all fork PRs via security gate (#46591) by @ydshieh in [#46591]\r\n* [CB] [Minor] Add parameter to tune default compile level (#46533) by @remi-or in [#46533]\r\n* Make DiffusionGemma trainable (#46568) by @kashif in [#46568]\r\n* docs: 🌐 add Turkish translation for README file (#46312) by @onuralpszr in [#46312]\r\n* fix-trainer-tests (#46541) by @SunMarc in [#46541]\r\n* Remove unnecessary expand_as in get_placeholder_mask across VLMs (#44907) by @syncdoth in [#44907]\r\n* [CI] Catch all shell/process execution issues in security gate via Bandit JSON report (#46560) by @ydshieh in [#46560]\r\n* Honor a concrete dtype in AutoModel for composite checkpoints (#46514) by @qflen in [#46514]\r\n* [CI] Implement real security check in PR CI security gate (#46557) by @ydshieh in [#46557]\r\n* [CI] Add 60s delay in security gate for flow observation (#46555) by @ydshieh in [#46555]\r\n* [TBC] [CI] Auto-approve PR CI for fork PRs via security gate (#46553) by @ydshieh in [#46553]\r\n* [CI] fix and make less flaky (#46543) by @zucchini-nlp in [#46543]\r\n* Fix hf_hub_download not placing file in current dir for url_to_local_path (#46545) by @ydshieh in [#46545]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @ArthurZucker\r\n    * Add minimax m3vl (#46600)\r\n* @eustlb\r\n    * Add Parakeet-RNNT (#46331)\r\n","publishedAt":"2026-06-12T14:39:40.000Z","fetchedAt":"2026-06-12T16:02:56.655Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.12.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/ae9dd96f-6877-4531-a06b-a756686f24e5","r2Key":"releases/2bbdbaa1355f77cb3ce9e695e4e4c7cbc1f16f24ef9a6247985c075dd09e35c0.png","r2Url":"https://media.releases.sh/releases/2bbdbaa1355f77cb3ce9e695e4e4c7cbc1f16f24ef9a6247985c075dd09e35c0.png"},{"type":"image","url":"https://github.com/user-attachments/assets/e62284ec-78bf-49cb-8aa2-deccc665372f","r2Key":"releases/a2a1b8a3605da8eda3ce1f197989fcc92e447f7b6743e04d76d992cfc06676ba.png","r2Url":"https://media.releases.sh/releases/a2a1b8a3605da8eda3ce1f197989fcc92e447f7b6743e04d76d992cfc06676ba.png"}],"coverageCount":0},{"id":"rel__DJqumOclHKZv-FMJKM2v","version":"v5.11.0","type":"feature","title":"Release v5.11.0","summary":"New models DiffusionGemma and DeepSeek-V3.2 have been added, featuring optimizations for inference speed and efficient long-context handling. The Kernels API was extended for module fusion and parameter transformation, with added support for fp8/fp4 Triton kernels. Model parallel beam search bugs in Qwen2-VL model families were fixed.","titleGenerated":"Transformers v5.11.0 adds DiffusionGemma and DeepSeek-V3.2 models","titleShort":"DiffusionGemma and DeepSeek-V3.2 models added","breaking":"unknown","importance":null,"content":"# Release v5.11.0\r\n\r\n\r\n## New Model additions\r\n\r\n### DiffusionGemma\r\n\r\n<img width=\"1240\" height=\"700\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5081e449-6374-4076-bd96-d295c8334ca4\" />\r\n\r\nDiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed. During inference, DiffusionGemma leverages multi-canvas sampling, where rather than generating one token at a time, the model iteratively denoises a full block of tokens using a diffusion sampler. This block-autoregressive approach facilitates text generation at higher speeds compared to traditional sequential generation methods.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/diffusion_gemma)\r\n* GPU go brr (#46540) by @gante in [#46540](https://github.com/huggingface/transformers/pull/46540)\r\n\r\n\r\n\r\n### DeepSeek-V3.2\r\n\r\n<img width=\"1135\" height=\"671\" alt=\"image\" src=\"https://github.com/user-attachments/assets/24c9694d-eeae-402c-9a98-f7a3971dd9d0\" />\r\n\r\nDeepSeek-V3.2-Exp is an experimental model from DeepSeek-AI that introduces DeepSeek Sparse Attention (DSA), a trainable, fine-grained sparse attention mechanism designed to improve training and inference efficiency in long-context scenarios. Built on top of DeepSeek-V3.1-Terminus with a 685B-parameter Mixture-of-Experts backbone, it reduces the quadratic cost of attention over long sequences by attending only to a selected subset of past tokens while maintaining virtually identical benchmark performance. The work was extended in DeepSeek-V3.2 which pairs DSA with scalable reinforcement learning and achieves gold-medal level results on competition math and competitive programming benchmarks.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/deepseek_v32) | [Paper](https://huggingface.co/papers/2512.02556)\r\n* Add deepseek 3.2 exp (#41251) by @ArthurZucker in [#41251](https://github.com/huggingface/transformers/pull/41251)\r\n\r\n\r\n## Kernels\r\n\r\nThe `KernelConfig` API was extended to support n-to-1 module fusion and parameter transformation, simplifying how custom kernels are integrated with Transformers modules. Additional fixes include resolving a dtype mismatch in the Mamba2 CUDA kernel path for NemotronH/Zamba2, adding fine-grained fp8/fp4 Triton kernel support, and correcting the FalconMamba fast-path warning to recommend `pip install kernels` instead of `mamba-ssm`.\r\n\r\n\r\n* Extended & simplified n-to-1 kernel fusion via KernelConfig (#46339) by @michaelbenayoun in [#46339]\r\n* Triton finegrained fp8/fp4 (#46407) by @IlyasMoutawwakil in [#46407]\r\n* Fix dtype mismatch in NemotronH/Zamba2 Mamba2 CUDA-kernel path (`out_proj`) (#46487) by @yuekaizhang in [#46487]\r\n* fix(falcon_mamba): recommend `pip install kernels` in fast-path warning (#46343) by @Anai-Guo in [#46343]\r\n\r\n\r\n## Parallelization\r\n\r\nFixed model parallel beam search bugs in the Qwen2-VL, Qwen2.5-VL, and Qwen3-VL MoE model families, and added documentation for tensor parallelism support with continuous batching.\r\n\r\n\r\n* [docs] tp for continuous batching (#46019) by @stevhliu in [#46019]\r\n* revisit history parallel beam search tests to avoid unnecessary fix (#46495) by @kaixuanliu in [#46495]\r\n* fix qwen series VL model's model parallel bug (#46316) by @kaixuanliu in [#46316]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Fix the offsets in processing (#46525) by @zucchini-nlp in [#46525]\r\n* Fix buggy action sha pin (#46534) by @ydshieh in [#46534]\r\n* Fix trailing comma bug in DataCollatorForLanguageModeling example (#46527) by @JemmaUZH in [#46527]\r\n* Fix missing Gemma4Processor._compute_audio_num_tokens (#46416) by @csantosbh in [#46416]\r\n* Fix InternVL models (#46524) by @hmellor in [#46524]\r\n* fix(afmoe): reduce tokens in test_compile_static_cache to avoid flaky bfloat16 drift (#46521) by @ydshieh in [#46521]\r\n* [CB] Add a \"max_requests_per_batch\" parameter (#46434) by @remi-or in [#46434]\r\n* revamp cv docs and fix rf-detr (#46219) by @merveenoyan in [#46219]\r\n* Update hub metadata (#46379) by @zucchini-nlp in [#46379]\r\n* extend DeepseekV4FlashIntegrationTest to non-cuda device (#46517) by @sywangyi in [#46517]\r\n* [docs] deepgemm (#46361) by @stevhliu in [#46361]\r\n* [fix] regression introduced by #45534 (#46456) by @eustlb in [#46456]\r\n* Use torchvision's native LANCZOS interpolation instead of PIL fallback (#46496) by @NicolasHug in [#46496]\r\n* Add debugging info in `pr-ci-caller.yml` (#46505) by @ydshieh in [#46505]\r\n* Fix tests: 'Cohere2MoeModel' object has no attribute 'hf_device_map' (#46337) by @kaixuanliu in [#46337]\r\n* Bump the actions group across 1 directory with 19 updates (#46414) by @dependabot[bot] in [#46414]\r\n* Log some information in `.github/workflows/pr-ci-post-dashboard-link.yml` (#46499) by @ydshieh in [#46499]\r\n* feat(quantizers): support non-weight param names in TorchAo safetensors loading (#46325) by @agesf in [#46325]\r\n* docs: fix typo in make_list_of_images docstring (#46469) by @ramkumar27072006 in [#46469]\r\n* add XPU expectation for deepseek_ocr2 model tests (#46492) by @kaixuanliu in [#46492]\r\n* Fix sapiens2 tests: add XPU device expectations (#46488) by @kaixuanliu in [#46488]\r\n* Add vLLM smoke test to CI (#46383) by @hmellor in [#46383]\r\n* extend deepseek v4 test to xpu (#46366) by @sywangyi in [#46366]\r\n* Added cosmos3 model (#46146) by @MaciejBalaNV in [#46146]\r\n* fbgemm_fp8:Keep the current device aligned with the input tensor (#46403) by @kaixuanliu in [#46403]\r\n* [Modular] Add `no_inherit_decorators` and fixup wrong RoPE related inheritances  (#46440) by @Bissmella in [#46440]\r\n* skip deepgemm test except cuda (#46090) by @jiqing-feng in [#46090]\r\n* Fix/video classification pipeline video processor (#46256) by @J3r3myPerera in [#46256]\r\n* ci: less flaky test_assisted_decoding_matches_greedy_search_1_same (#46445) by @ydshieh in [#46445]\r\n* Fix flip_back graph break (#46344) by @guarin in [#46344]\r\n* Add the other processors to auto-mappings (#46046) by @zucchini-nlp in [#46046]\r\n* fix: compatibility with torch<=2.7 (#46393) by @andylin-hao in [#46393]\r\n* fix: remove dynamic per-actor Slack ID lookup in ssh-runner workflow (#46327) by @ydshieh in [#46327]\r\n* [docs] Romanian translation of `pipeline_tutorial.md`, `pipeline_gradio.md`, `pipeline_webserver.md` and `add_new_pipeline.md`. (#46388) by @filipinescu in [#46388]\r\n* [docs] gemma4 typos (#46351) by @stevhliu in [#46351]\r\n* [docs] padding-free training (#46333) by @stevhliu in [#46333]\r\n* fix[vLLM x v5]: Default untied embeddings in AudioFlamingo3 and VibeVoice (#46400) by @harshaljanjani in [#46400]\r\n* Fix deepspeed docker (#46108) by @SunMarc in [#46108]\r\n* Fix conversion for clip models (#46406) by @zucchini-nlp in [#46406]\r\n* ci: mention code quality failure in CI dashboard comment (#46415) by @ydshieh in [#46415]\r\n* Fix noisy logging from image_processing module aliases issue - 46298 (#46350) by @skshmjn in [#46350]\r\n* Raise tqdm minimum to 4.60 to match tqdm.contrib.logging import (#46397) by @n0gu-furiosa in [#46397]\r\n* fix(gemma4_unified): conversion script and config bugs (#46398) by @douglas-reid in [#46398]\r\n* [docs] remove sparsity from compressed-tensors (#46387) by @stevhliu in [#46387]\r\n* [CB] Fix crashes when fork is not possible (#46251) by @remi-or in [#46251]\r\n* Improve CI dashboard comment: rename and deduplicate (#46412) by @ydshieh in [#46412]\r\n* Fix missing f-string prefixes in error messages (#46354) by @joaopedroassad in [#46354]\r\n* Add workflow to post CI Grafana dashboard link to PR (#46410) by @ydshieh in [#46410]\r\n* [docs] Romanian translation of `fast_tokenizers.md`, `custom_tokenizers.md`, `tokenizer_summary.md`, `image_processors.md` and `video_processors.md`. (#46356) by @filipinescu in [#46356]\r\n* Clean up new models after release (#46092) by @zucchini-nlp in [#46092]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @ArthurZucker\r\n    * Add deepseek 3.2 exp (#41251)\r\n* @gante\r\n    * GPU go brr (#46540)\r\n* @merveenoyan\r\n    * revamp cv docs and fix rf-detr (#46219)\r\n* @sgerrard\r\n    * Quantization for small models (#46449)\r\n* @MaciejBalaNV\r\n    * Added cosmos3 model (#46146)\r\n* @J3r3myPerera\r\n    * Fix/video classification pipeline video processor (#46256)\r\n* @filipinescu\r\n    * [docs] Romanian translation of `pipeline_tutorial.md`, `pipeline_gradio.md`, `pipeline_webserver.md` and `add_new_pipeline.md`. (#46388)\r\n    * [docs] Romanian translation of `fast_tokenizers.md`, `custom_tokenizers.md`, `tokenizer_summary.md`, `image_processors.md` and `video_processors.md`. (#46356)","publishedAt":"2026-06-10T16:32:44.000Z","fetchedAt":"2026-06-10T19:03:14.820Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.11.0","media":[{"type":"image","url":"https://github.com/user-attachments/assets/5081e449-6374-4076-bd96-d295c8334ca4","r2Key":"releases/c9ccc725b99e6f64362fd0a76aba5732768f33f7b04c4920e9e5383c2df90580.png","r2Url":"https://media.releases.sh/releases/c9ccc725b99e6f64362fd0a76aba5732768f33f7b04c4920e9e5383c2df90580.png"},{"type":"image","url":"https://github.com/user-attachments/assets/24c9694d-eeae-402c-9a98-f7a3971dd9d0","r2Key":"releases/db98db5130d964d9e55d017c3ec27c4123c8ba50d1881d76c16860eef273cb89.png","r2Url":"https://media.releases.sh/releases/db98db5130d964d9e55d017c3ec27c4123c8ba50d1881d76c16860eef273cb89.png"}],"coverageCount":0},{"id":"rel_jVriQTYwWbtOv8DbohX_Q","version":"v5.10.2","type":"feature","title":"Patch release v5.10.2","summary":"Fixed a conversion bug for CLIP models that affected downstream models like SAM3.","titleGenerated":"Transformers v5.10.2 fixes CLIP model conversion bug affecting SAM3","titleShort":"CLIP model conversion bug fixed","breaking":"unknown","importance":null,"content":"# Patch release v5.10.2\r\nThere was a big bug in the model conversion of models related to clip, this affected models like sam3 and others. Please make sure to update :pray: \r\n\r\n* Fix conversion for clip models by @zucchini-nlp (#46406)\r\n\r\n\r\n**Full Changelog**: https://github.com/huggingface/transformers/compare/v5.10.1...v5.10.2","publishedAt":"2026-06-04T18:43:06.000Z","fetchedAt":"2026-06-04T23:02:56.210Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.10.2","media":[],"coverageCount":0},{"id":"rel_yiTu-0WK6auFbR0DQMM3m","version":"v5.10.1","type":"feature","title":"Release v5.10.1","summary":"Added Gemma4 12B Unified, an encoder-free multimodal model that projects raw vision and audio inputs directly into language model space; Sapiens2, a vision transformer family for human-centric tasks; DeepSeek-OCR-2 for document understanding; and Mellum, a code-focused mixture-of-experts model. Fixed numerous model parallelism bugs across tensor and expert parallelism, beam search under parallel settings, and loss over-counting; also fixed encoder-decoder cache initialization regression and BitsAndBytes quantization tensor-dropping bug.","titleGenerated":"Transformers v5.10.1 adds Gemma4 Unified, Sapiens2, and model parallelism fixes","titleShort":"Gemma4 Unified; Sapiens2; model parallelism hardened","breaking":"unknown","importance":null,"content":"# Release v5.10.1\r\nv5.10.0 was yanked as we publish on a corrupted branch. Sorry everyone, this happens when we rush a release!!! \r\n\r\n## New Model additions\r\n\r\n### Gemma4 unified+ Gemma4 MTP\r\n<img width=\"2000\" height=\"400\" alt=\"image\" src=\"https://github.com/user-attachments/assets/5e3ee940-f78d-4343-ac7a-889930800aa6\" />\r\n\r\nGemma 4 12B Unified is an **encoder-free** multimodal model with pretrained and instruction-tuned variants. Unlike [standard Gemma 4](./gemma4), which uses dedicated encoder towers, Gemma 4 12B Unified projects raw inputs directly into the language model's embedding space through lightweight linear pipelines. This results in a simpler architecture while maintaining strong multimodal performance.\r\n\r\nKey differences from standard Gemma 4:\r\n- **No Vision Tower**: Raw pixel patches are projected directly into LM space via a `Dense + LayerNorm` pipeline with factorized 2D positional embeddings, replacing the vision encoder.\r\n- **No Audio Tower**: Raw 16 kHz waveform samples are chunked into fixed-length frames and projected through a simple `RMSNorm → Linear` pipeline, replacing the mel spectrogram + Conformer encoder.\r\n- **Shared Multimodal Pipeline**: Both vision and audio use the same `Gemma4UnifiedMultimodalEmbedder` (RMSNorm → Linear) for the final projection to text hidden space.\r\n\r\nYou can find the original Gemma 4 12B Unified checkpoints under the [Gemma 4](https://huggingface.co/collections/google/gemma-4) release.\r\n\r\n* who needs encoders? (#46385) by @douglas-reid @sgerrard @vasqu @molbap\r\n\r\n### Sapiens2\r\n\r\nSapiens2 is a family of high-resolution vision transformers pretrained on ~1 billion curated human images, designed for human-centric computer vision tasks including pose estimation, body-part segmentation, surface normal estimation, and pointmap estimation. The models scale from 0.4B to 5B parameters and train at native 1K resolution, with hierarchical 4K variants for extended spatial reasoning. Sapiens2 achieves substantial improvements over its predecessor with +4 mAP in pose estimation, +24.3 mIoU in body-part segmentation, and 45.6% error reduction in normal estimation.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/sapiens2) | [Paper](https://huggingface.co/papers/2604.21681)\r\n* Add Sapiens2 Model (#45919) by @guarin in [#45919](https://github.com/huggingface/transformers/pull/45919)\r\n\r\n### DeepSeek-OCR-2\r\n\r\nDeepSeek-OCR-2 is an OCR-specialized vision-language model built on a distinctive architecture that combines a SAM ViT-B vision encoder with a Qwen2 hybrid attention encoder, connected through an MLP projector to a DeepSeek-V2 Mixture-of-Experts (MoE) language model. The model features a hybrid attention mechanism that applies bidirectional attention over image tokens and causal attention over query tokens, enabling efficient and accurate document understanding. It supports both plain OCR tasks and grounding capabilities with coordinate-aware output for document conversion to markdown format.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/deepseek_ocr2)\r\n* Add Deepseek-OCR-2 model (#45075) by @thisisiron in [#45075](https://github.com/huggingface/transformers/pull/45075)\r\n\r\n### Mellum\r\n\r\nMellum is a code-focused Mixture-of-Experts language model developed by JetBrains. It is derived from the Qwen3-MoE architecture with per-layer-type RoPE and interleaved sliding window attention. The model has 12B total parameters with 2.5B active parameters per token, using 64 routed experts with 8 activated per token across 28 layers.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/mellum)\r\n* feat: Add support for JetBrains' `Mellum` v2 code generation model (#46112) by @shadeMe in [#46112](https://github.com/huggingface/transformers/pull/46112)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nThe Gemma4 vision pooler now casts inputs to float32 before scaling to prevent float16 overflow (inf saturation) with large checkpoints, which may cause minor numerical differences in outputs for users running Gemma-4 vision models in float16.\r\n* 🚨 Fix float16 overflow in Gemma4 vision pooler (#46277) by @Bluear7878\r\n\r\nAudio Language Models (ALMs) now have a dedicated base model class without a language modeling head, aligning them with the design of Vision Language Models (VLMs); users relying on the previous model class structure should update their code to use the new base model class where appropriate.\r\n* 🚨 [ALM] Add base model without head (#45534) by @eustlb\r\n\r\n\r\n\r\n## Parallelization\r\n\r\nThis release includes numerous bug fixes for model parallelism across multiple models (Gemma4, AltCLIP, ChineseClip, Blip-2, Whisper, Ovis2, Moshi) and parallel execution strategies, including fixes for tensor parallelism (TP), expert parallelism (EP), beam search under model parallel settings, and loss over-counting under TP/EP configurations. The continuous batching manager was also reworked for clearer control flow and improved TP race condition handling, and FSDP initialization via `from_pretrained` was introduced.\r\n\r\n\r\n* Fix dsv4 dequant + tp/ep (#46378) by @IlyasMoutawwakil in [#46378]\r\n* [CB] [Major] Rework manager to have clearer control flow + handle TP (#46070) by @remi-or in [#46070]\r\n* fix series of bugs for model parallel beam search (#46280) by @kaixuanliu in [#46280]\r\n* Fix model parallel issue for altclip model and ChineseClip model (#45487) by @kaixuanliu in [#45487]\r\n* Model parallel fix (#46230) by @kaixuanliu in [#46230]\r\n* [`Revert`] FSDP+Dtensor refactor related changes (#46246) by @vasqu in [#46246]\r\n* Fix model parallel bugs for Gemma4 (#45817) by @kaixuanliu in [#45817]\r\n* init FSDP through from_pretrained (#46102) by @3outeille in [#46102]\r\n* fix model parallel device mismatch issue in `create_bidirectional_mask` (#46221) by @kaixuanliu in [#46221]\r\n* Trainer.compute_loss: fix loss over-counting under TP and EP-as-TP (#45994) by @AmineDiro in [#45994]\r\n* Fix caching allocator warmup byte estimation for EP model loading (#46149) by @sywangyi in [#46149]\r\n\r\n\r\n## Cache\r\n\r\nFixed a regression in encoder-decoder cache initialization where the decoder config was incorrectly applied to the cross-attention cache, and resolved a `RuntimeError` caused by buffer size limits when warming up the cache on MPS devices. Additional test infrastructure improvements were made to support read-only cache environments used in CI.\r\n\r\n\r\n* fix: cache warmup `RuntimeError` on mps (#46239) by @McPatate in [#46239]\r\n* Make more tests work with read-only cache (#46299) by @ydshieh in [#46299]\r\n* Update a test to avoid writing to the default xet cache (#46250) by @ydshieh in [#46250]\r\n* Fix a regression in encoder-decoder generation cache initialization (#46111) by @kaixuanliu in [#46111]\r\n\r\n\r\n## Quantization\r\n\r\nAdded support for DeepGEMM BF16, mixed FP8/FP4, and MegaMoE quantization via a grouped linear refactor, while fixing two bugs: an FP8 MoE reverse substring issue affecting DSv4 initialization, and a BitsAndBytes 4-bit/8-bit quantization bug that silently dropped chunked tensors from one-to-many weight converters.\r\n\r\n\r\n* DeepGEMM BF16 + mixed FP8/FP4 + MegaMoE + refactor (#45634) by @IlyasMoutawwakil in [#45634]\r\n* Fix fp8 moe reverse substring (#46265) by @ArthurZucker in [#46265]\r\n* Fix bnb 4bit/8bit quantization drop chunked tensors bug (#46210) by @kaixuanliu in [#46210]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Fix wrong changes produced by style/repo. check bot (#46371) by @ydshieh in [#46371]\r\n* Fix path traversal when saving Bark voice preset embeddings (#46237) by @LinZiyuu in [#46237]\r\n* Pass library_name/version to Hub calls via a shared HfApi (#46318) by @Wauplin in [#46318]\r\n* docs: update ACL Anthology URL in CITATION.cff (#46352) by @irfaan101 in [#46352]\r\n* [docs] contributing (#45465) by @stevhliu in [#45465]\r\n* [docs] Romanian translation of `contributing.md`, `modular_transformers.md`, `multimodal_processing.md`, `add_vision_processing_components.md`, `add_audio_processing_components.md`, `modeling_rules.md`, `model_output_tracing.md`, `auto_docstring.md`, `testing.md`, `pr_checks.md` and `add_new_model.md` . (#46345) by @filipinescu in [#46345]\r\n* [docs] xpu continuous batching (#46334) by @stevhliu in [#46334]\r\n* Fix incorrect attribute mapping relationships in GLM MoE DSA Config (#46338) by @Dovis01 in [#46338]\r\n* Fix grammar typos in Whisper documentation (#46336) by @calliec-1223 in [#46336]\r\n* [docs] update num_items_in_batch for causal LMs (#46335) by @stevhliu in [#46335]\r\n* Update compressed tensors minimum version (#46342) by @SunMarc in [#46342]\r\n* Fix _is_package_available reporting available without a version (#46125) by @blipbyte in [#46125]\r\n* remove sec (#46346) by @ydshieh in [#46346]\r\n* fix: include transitive relative imports when loading from local directory (#46022) by @trducng in [#46022]\r\n* perf(feature_extraction_sequence): skip re-splitting already-batched numpy arrays in pad() (#46329) by @Anai-Guo in [#46329]\r\n* [Zamba] Support attn_implementation dispatch (#46317) by @YangKai0616 in [#46317]\r\n* Fix TestAppRoutes test failures caused by deprecated asyncio.get_event_loop() on Python 3.10+ (#46340) by @ydshieh in [#46340]\r\n* [Qwen3VL] Fix video token placeholder: use self.video_token instead of hardcoded \"<|placeholder|>\" (#46296) by @kpal002 in [#46296]\r\n* chore(linter): fixes for rule 16 (#46023) by @tarekziade in [#46023]\r\n* [docs] Romanian translation of  `weightconverter.md`,  `models.md`,  `custom_models.md`,  `monkey_patching.md`,  `fusion_mapping.md`, `how_to_hack_models.md`, `model_sharing.md` and `serialization.md`. (#46309) by @filipinescu in [#46309]\r\n* Normalize CUDA OOM errors when comparing commit failures in check_bad_commit (#46322) by @ydshieh in [#46322]\r\n* Fix unhandled exception noise from background safetensors conversion thread (#45752) by @dhruv7477 in [#45752]\r\n* Add Expectations for pipeline token classification tests (#46151) by @kaixuanliu in [#46151]\r\n* [docs] fix auto-add release dates (#46283) by @zucchini-nlp in [#46283]\r\n* Separate pip command syntax for notebook and CLI tabs in Quickstart (#46243) by @pvelayudhan in [#46243]\r\n* Romanian translation of README.md, index.md, installation.md, _config.py and quicktour.md. (#46166) by @filipinescu in [#46166]\r\n* Fall back to flat kwarg when modality dict is passed without it (#46195) by @Ace3Z in [#46195]\r\n* Fix load_adapter OOM caused by full-model warmup sizing (#46145) by @Yooniel in [#46145]\r\n* Replace assert with raise ImportError for optuna/ray dependency checks (#46263) by @SebTardif in [#46263]\r\n* chore(linter): respect TRF017 modeling rule (#46260) by @tarekziade in [#46260]\r\n* Delete dead code in qwen-vl series (#45827) by @zucchini-nlp in [#45827]\r\n* qa: fix ty caching and align CI with local run (#46278) by @tarekziade in [#46278]\r\n* Guard DeviceMesh import in continuous batching (#46205) by @danyalahmed1995 in [#46205]\r\n* Processor compatibility with vLLM  (#46258) by @zucchini-nlp in [#46258]\r\n* Fix PR CI workflow cancellation condition (#46276) by @ydshieh in [#46276]\r\n* [fix] toctree (#46106) by @stevhliu in [#46106]\r\n* add more generic support for distributed trainer tests (#46109) by @kaixuanliu in [#46109]\r\n* add XPU Expectations for florence2 and lfm2_vl model test (#46275) by @kaixuanliu in [#46275]\r\n* Fix `StaticCache` building an empty layer list when `num_kv_shared_layers == 0` (#46235) by @tengomucho in [#46235]\r\n* Fix inverted assertion in remove_handler (#46227) by @SebTardif in [#46227]\r\n* [ShieldGemma2] Support attn_implementation dispatch (#46069) by @YangKai0616 in [#46069]\r\n* [Gemma4] Replace one-hot matmul with F.embedding in position embeddings (#46176) by @Sriniketh24 in [#46176]\r\n* fix: kosmos2.5: properly expand embeddings table (#45835) by @nunq in [#45835]\r\n* find pytest launch error in torch 2.13.0.dev20260526 (#46252) by @sywangyi in [#46252]\r\n* [Test][Kosmos2.5] Add XPU expectations for integration tests (#46135) by @YangKai0616 in [#46135]\r\n* Support FA2 flash_attn_with_kvcache for XPU continuous batching (#46028) by @YangKai0616 in [#46028]\r\n* [`Configs`] Fix layer type validation to include its mlp counterpart (#46220) by @vasqu in [#46220]\r\n* Fix `num_items_in_batch` over-counting for causal LM losses (#46204) by @qgallouedec in [#46204]\r\n* RF-DETR doc fixes (#46244) by @merveenoyan in [#46244]\r\n* Use `main` instead of commit SHA for now (#46241) by @ydshieh in [#46241]\r\n* Enable push event (to main) for PR CI workflow (#46240) by @ydshieh in [#46240]\r\n* fix(hrm_text): Add XPU Expectations for tests (#46214) by @kaixuanliu in [#46214]\r\n* [deepseek_v4] keep hc_head / sinks / position_bias in fp32 (#46198) by @ArthurZucker in [#46198]\r\n* Fix FSDP2 and distributed checkpointing imports for older PyTorch versions (#46141) by @ryota-komatsu in [#46141]\r\n* Fix Gemma4 Array Mask Indexing (#46203) by @petecao in [#46203]\r\n* utils: handle flash_attn missing from importlib packages_distributions without crashing (#45524) by @SAY-5 in [#45524]\r\n* [AMD CI] revert AMD mi325 hf-workflows ref from SHA back to @main (#46213) by @Abdennacer-Badaoui in [#46213]\r\n* [GLM-4.6V] Update with GLM-GA Processor (#46184) by @zRzRzRzRzRzRzR in [#46184]\r\n* update xpu expectation for falcon mamba (#46086) by @sywangyi in [#46086]\r\n* chore: enable Dependabot weekly GitHub Actions bumps (#46157) by @hf-dependantbot-rollout[bot] in [#46157]\r\n* Fix Gemma4 use_bidirectional_attention=\"all\" mask behavior (#46079) by @oliverholworthy in [#46079]\r\n* Fix loading with only 1 device or distributed config (#46197) by @Cyrilvallez in [#46197]\r\n* Fix TypeError on list-typed ignore_keys_at_rope_validation in RoPE config (#46142) by @Charly21r in [#46142]\r\n* Support XPU autocast dtype fallback for FlashAttention (#46199) by @YangKai0616 in [#46199]\r\n* Fix path traversal when saving named chat templates (#46191) by @LinZiyuu in [#46191]\r\n* Fix is_last off-by-one in MaskGenerationPipeline for partial batches (#46136) by @J3r3myPerera in [#46136]\r\n* Fix wrong variable in check_model_type isinstance check (#46080) by @SebTardif in [#46080]\r\n* Enable passing kwargs through RoFormer models (#46171) by @ir2718 in [#46171]\r\n* Update cohere2_moe tp_plan (#46189) by @Cyrilvallez in [#46189]\r\n* Update release tool (#46193) by @Cyrilvallez in [#46193]\r\n* [loading] Fix base_model_prefix issues in conversions (#46067) by @Cyrilvallez in [#46067]\r\n* Bump dev version (#46188) by @Cyrilvallez in [#46188]\r\n* Update self-comment-ci (#46137) by @guarin in [#46137]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @filipinescu\r\n    * [docs] Romanian translation of `contributing.md`, `modular_transformers.md`, `multimodal_processing.md`, `add_vision_processing_components.md`, `add_audio_processing_components.md`, `modeling_rules.md`, `model_output_tracing.md`, `auto_docstring.md`, `testing.md`, `pr_checks.md` and `add_new_model.md` . (#46345)\r\n    * [docs] Romanian translation of  `weightconverter.md`,  `models.md`,  `custom_models.md`,  `monkey_patching.md`,  `fusion_mapping.md`, `how_to_hack_models.md`, `model_sharing.md` and `serialization.md`. (#46309)\r\n    * Romanian translation of README.md, index.md, installation.md, _config.py and quicktour.md. (#46166)\r\n* @remi-or\r\n    * [CB] [Major] Rework manager to have clearer control flow + handle TP (#46070)\r\n* @thisisiron\r\n    * Add Deepseek-OCR-2 model (#45075)\r\n* @kaixuanliu\r\n    * Add Expectations for pipeline token classification tests (#46151)\r\n    * fix series of bugs for model parallel beam search (#46280)\r\n    * add more generic support for distributed trainer tests (#46109)\r\n    * add XPU Expectations for florence2 and lfm2_vl model test (#46275)\r\n    * Fix model parallel issue for altclip model and ChineseClip model (#45487)\r\n    * Model parallel fix (#46230)\r\n    * fix(hrm_text): Add XPU Expectations for tests (#46214)\r\n    * Fix model parallel bugs for Gemma4 (#45817)\r\n    * Fix bnb 4bit/8bit quantization drop chunked tensors bug (#46210)\r\n    * fix model parallel device mismatch issue in `create_bidirectional_mask` (#46221)\r\n    * Fix a regression in encoder-decoder generation cache initialization (#46111)\r\n* @shadeMe\r\n    * feat: Add support for JetBrains' `Mellum` v2 code generation model (#46112)\r\n* @vasqu\r\n    * [`Revert`] FSDP+Dtensor refactor related changes (#46246)\r\n    * [`Configs`] Fix layer type validation to include its mlp counterpart (#46220)\r\n* @zRzRzRzRzRzRzR\r\n    * [GLM-4.6V] Update with GLM-GA Processor (#46184)\r\n* @eustlb\r\n    * 🚨 [ALM] Add base model without head (#45534)\r\n","publishedAt":"2026-06-03T15:37:41.000Z","fetchedAt":"2026-06-03T17:03:20.892Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.10.1","media":[{"type":"image","url":"https://github.com/user-attachments/assets/5e3ee940-f78d-4343-ac7a-889930800aa6","r2Key":"releases/51f28dfd1e6998b5282dafde7977844b798ec8af1c111248504800fbcaa42ba8.png","r2Url":"https://media.releases.sh/releases/51f28dfd1e6998b5282dafde7977844b798ec8af1c111248504800fbcaa42ba8.png"}],"coverageCount":0},{"id":"rel_onVV6ONIOpM0JzX0O6gVK","version":"v5.9.0","type":"feature","title":"Release v5.9.0","summary":"Added support for Cohere2Moe (a Mixture-of-Experts model with sliding window and full attention), HRM-Text (hierarchical reasoning model with two transformer stacks), and Parakeet tdt speech model. SAM3, EdgeTAM, and SAM3-Lite-Text now expect full text embeddings instead of pooler outputs, requiring input updates. Fixed generation issues including inputs_embeds handling for Gemma4, an AttributeError in RAG's generate() caused by missing config fields, memory leaks from lru decorators in vision models, and improved audio/vision encoder compilability.","titleGenerated":"Transformers v5.9.0 adds Cohere2Moe, HRM-Text, and Parakeet models; fixes generation bugs","titleShort":"Three new models added; SAM3 text embeddings API changes; generation bugs fixed","breaking":"unknown","importance":null,"content":"# Release v5.9.0\r\n\r\n\r\n## New Model additions\r\n\r\n### Cohere2Moe\r\n\r\nCommand A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers. The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/cohere2_moe)\r\n* Add new cohere2_moe model (#46115) by @Cyrilvallez in [#46115](https://github.com/huggingface/transformers/pull/46115)\r\n\r\n### Parakeet tdt (#44171)\r\n\r\n* Parakeet tdt (#44171) by @lmaksym\r\n\r\n### HRM-Text\r\n\r\nHRM-Text is an improved autoregressive language-modeling variant of the Hierarchical Reasoning Model (HRM) that uses a hierarchical recurrent forward pass with two transformer stacks - one for slow, abstract planning (H) and one for fast, detailed computation (L) - reused inside a nested recurrence. It features PrefixLM attention where instruction tokens attend bidirectionally while response tokens attend causally, per-head sigmoid output gates, and parameterless RMSNorm. The model is designed as a base language model without instruction tuning or chat templates.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/hrm_text) | [Paper](https://huggingface.co/papers/2506.21734)\r\n* Add hrm text (#46025) by @abcd1927 in [#46025](https://github.com/huggingface/transformers/pull/46025)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nThe `text_embeds` input for SAM3, EdgeTAM, and SAM3-Lite-Text models now expects full text embeddings instead of just pooler outputs, aligning with other models in the library — users must update their inputs accordingly.\r\n* 🚨Fix memory leaks caused by lru decorators in vision models (#45922) by @yonigozlan\r\n\r\n\r\n\r\n## Audio\r\n\r\nAudio support was expanded with the addition of AudioFlamingoNext model checkpoints and improved compilability of audio/vision encoders via standalone pure functions. Additional improvements include better error messaging when loading audio from video files and new documentation for audio/video processors.\r\n\r\n\r\n* user friendly error when loading audio from video (#45221) by @eustlb in [#45221]\r\n* [docs] adding audio/video processors (#45795) by @stevhliu in [#45795]\r\n* Support Audio Flamingo Next checkpoints (#44830) by @lashahub in [#44830]\r\n* Extract dynamic vision/audio tensors into standalone pure functions (#45396) by @IlyasMoutawwakil in [#45396]\r\n\r\n\r\n## Generation\r\n\r\nFixed generation issues including `inputs_embeds` and `per_layer_inputs` handling for Gemma4, an `AttributeError` in RAG's `generate()` caused by missing config fields, and flaky VLM generation tests by blocking special image tokens during sampling.\r\n\r\n\r\n* Fix Gemma4 generation from inputs_embeds and per_layer_inputs (#46049) by @Cyrilvallez in [#46049]\r\n* Fix AttributeError in RAG generate() for missing config fields (#46035) by @Sriniketh24 in [#46035]\r\n* Block image_start/end_token_id in generation test sampling (#45914) by @Rocketknight1 in [#45914]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* Remove mask visualization tool from `masking_utils.py` (#46066) by @Cyrilvallez in [#46066]\r\n* fix: owned_by field in GET /v1/models returns list instead of string (#46006) by @nileshpatil6 in [#46006]\r\n* [CB] Remove OpenTelemetry (#45984) by @remi-or in [#45984]\r\n* docs(readme): use canonical `huggingface.co` domain in prose links (#46042) by @kiwigitops in [#46042]\r\n* Fix remaining RAG doc examples that crash on current transformers (#46044) by @Sriniketh24 in [#46044]\r\n* Init the actual tensor, not a copy (#46030) by @Rocketknight1 in [#46030]\r\n* docs: sync legacy ACL anthology URLs and update metrics across i18n READMEs (#46027) by @irfaan101 in [#46027]\r\n* [MultimodalLM] add language_model to the get/set_input_embeddings logic (#46029) by @eustlb in [#46029]\r\n* [`HRM Text`] Add integration tests (#46033) by @vasqu in [#46033]\r\n* hy_v3: add XPU expectations (#45858) by @kaixuanliu in [#45858]\r\n* exaone4_5: add XPU expectations (#45890) by @kaixuanliu in [#45890]\r\n* hyperclovax: add XPU Expectations for CI test (#45926) by @kaixuanliu in [#45926]\r\n* chore(ci): remove dead env vars from circleci-failure-summary-comment.yml (#45972) by @XciD in [#45972]\r\n* [CB] [Major] Add tensor paralellism (#45821) by @remi-or in [#45821]\r\n* docs: update models architecture count and sync ACL anthology URLs (#46001) by @irfaan101 in [#46001]\r\n* bugfix(ci): avoid E2BIG in pr_slow_ci_suggestion  (#45983) by @tarekziade in [#45983]\r\n* RFDetr - use correct Roboflow org for release (#45946) by @sbucaille in [#45946]\r\n* docs: Fix formatting issues in weightconverter.md (#45988) by @ArjunSrivastava1 in [#45988]\r\n* Fix colqwen2 test (#45981) by @IlyasMoutawwakil in [#45981]\r\n* Fix M-RoPE device mismatch in Qwen3VL family under FSDP2 CPU offload (#45861) by @jamesbraza in [#45861]\r\n* [docs] chat template prefill (#45947) by @stevhliu in [#45947]\r\n* [docs] decode fast path (#45899) by @stevhliu in [#45899]\r\n* fix: restore `_attn_implementation `and fix request offset in `generate_batch()` (#45943) by @sergiopaniego in [#45943]\r\n* Expose `per_layer_inputs` for every Gemma4 variants (#45927) by @Cyrilvallez in [#45927]\r\n* chore: update benchmark_v2.yml (#45966) by @hf-security-analysis[bot] in [#45966]\r\n* fix(ci): set persist-credentials: false on actions/checkout and close remaining template injection findings (#45964) by @XciD in [#45964]\r\n* chore(ci): set default workflow permissions to contents: read (#45961) by @XciD in [#45961]\r\n* fix(ci): remove template injection on pull_request_target workflows (#45956) by @XciD in [#45956]\r\n* chore(ci): pin all GitHub Actions and reusable workflows by SHA (#45955) by @XciD in [#45955]\r\n* [docs] ALMModelTest (#45900) by @stevhliu in [#45900]\r\n* Enhance apply_chat_template to support custom field prefilling (reasoning_content, thinking, etc.) (#45896) by @Mamiglia in [#45896]\r\n* BUGFIX: Support hubert models that don't have conv_pos_batch_norm configured (#45921) by @igordertigor in [#45921]\r\n* Revert 45777 (#45942) by @Rocketknight1 in [#45942]\r\n* pass the otel secrets (#45933) by @tarekziade in [#45933]\r\n* Add initial torch_tpu backend support (#45918) by @tengomucho in [#45918]\r\n* [CB] Hide activation footprint by using the CUDA graph pool (#45911) by @remi-or in [#45911]\r\n* Require input_ids for repetition penalty (#45389) by @ruben-aghayan in [#45389]\r\n* Fix undefined 'input' variable (#45895) by @fullyz in [#45895]\r\n* Fix post processing RF-DETR (#46041) by @yonigozlan (direct commit on v5.9.0)\r\n* [loading] Free up tensors faster inside ConversionOps (#46110) by @Cyrilvallez (direct commit on v5.9.0)\r\n* Add new cohere2_moe model (#46115) by @Cyrilvallez (direct commit on v5.9.0)\r\n* Fix cohere2 tp_plan for release by @Cyrilvallez (direct commit on v5.9.0)\r\n* Release v5.9.0 by @Cyrilvallez (direct commit on v5.9.0)\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @lmaksym\r\n    * Parakeet tdt (#44171)\r\n* @eustlb\r\n    * user friendly error when loading audio from video (#45221)\r\n    * [MultimodalLM] add language_model to the get/set_input_embeddings logic (#46029)\r\n* @remi-or\r\n    * [CB] Remove OpenTelemetry (#45984)\r\n    * [CB] [Major] Add tensor paralellism (#45821)\r\n    * [CB] Hide activation footprint by using the CUDA graph pool (#45911)\r\n* @abcd1927\r\n    * Add hrm text (#46025)","publishedAt":"2026-05-20T14:12:54.000Z","fetchedAt":"2026-05-20T18:03:31.314Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.9.0","media":[],"coverageCount":0},{"id":"rel_HOSTAN_TNLzhm71qgv64o","version":"v5.8.1","type":"feature","title":"Patch release v5.8.1","summary":"Fixed Deepseek V4 integration issues including CSA mask collapse and WeightConverter regex incorrectly matching shared_experts as experts. Also added fatal_error to ContinuousBatchingManager for serving operations.","titleGenerated":"Transformers v5.8.1 fixes Deepseek V4 integration","titleShort":"Deepseek V4 integration fixed","breaking":"unknown","importance":null,"content":"# Patch release v5.8.1 \r\nThis release is mainly to fix the Deepseek V4 integration!!! \r\n\r\n<img width=\"714\" height=\"774\" alt=\"image\" src=\"https://github.com/user-attachments/assets/0d85e891-a0ff-436e-a9d4-b6633096f2b5\" />\r\n\r\n\r\n* [fix] Add fatal_error to ContinuousBatchingManager so the serving... by @qgallouedec, @remi-or\r\n* Fix WeightConverter regex incorrectly matching shared_experts as experts by @silencelamb, @claude\r\n* Fix deepseek v4 by @ArthurZucker (#45892)\r\n* Deepseek v4 csa mask collapse by @ArthurZucker, @Sawyer117 (#45928)","publishedAt":"2026-05-13T03:21:23.000Z","fetchedAt":"2026-05-13T05:02:58.223Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.8.1","media":[],"coverageCount":0},{"id":"rel_oP71J5Z3r1csDka7MGv8h","version":"v5.8.0","type":"feature","title":"Release 5.8.0","summary":"# Release v5.8.0\r\n\r\n\r\n## New Model additions\r\n\r\n### DeepSeek-V4\r\n\r\n<img width=\"6604\" height=\"3574\" alt=\"image\" src=\"https://github.com/user-attachment...","titleGenerated":null,"titleShort":null,"breaking":"unknown","importance":null,"content":"# Release v5.8.0\r\n\r\n\r\n## New Model additions\r\n\r\n### DeepSeek-V4\r\n\r\n<img width=\"6604\" height=\"3574\" alt=\"image\" src=\"https://github.com/user-attachments/assets/4c0fdb29-f770-463c-a97b-d24438896a4c\" />\r\n\r\nDeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3. The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table. This implementation covers DeepSeek-V4-Flash, DeepSeek-V4-Pro, and their -Base pretrained variants, which share the same architecture but differ in width, depth, expert count and weights.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/deepseek_v4) | [Paper](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash/blob/main/DeepSeek_V4.pdf)\r\n* Add DeepSeek V4 (#45643) by @ArthurZucker in [#45643](https://github.com/huggingface/transformers/pull/45643)\r\n\r\n### Gemma 4 Assistant\r\n\r\n<img width=\"2000\" height=\"400\" alt=\"image\" src=\"https://github.com/user-attachments/assets/02c79b0b-a172-4495-b09d-a6a4b625ee66\" />\r\n\r\nGemma 4 Assistant is a small, text-only model that enables speculative decoding for Gemma 4 models using the Multi-Token Prediction (MTP) method and associated candidate generator. The model shares the same Gemma4TextModel backbone as other Gemma 4 models but uses KV sharing throughout the entire model, allowing it to reuse the KV cache populated by the target model and skip the pre-fill phase entirely. This architecture includes cross-attention to make the most of the target model's context, allowing the assistant to accurately predict more drafted tokens per drafting round.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/gemma4_assistant)\r\n* First model (#45788) by @SindhuRaghuram97 in [#45788](https://github.com/huggingface/transformers/pull/45788)\r\n\r\n### GraniteSpeechPlus\r\n\r\n<img width=\"1310\" height=\"930\" alt=\"image\" src=\"https://github.com/user-attachments/assets/94fc3730-742c-4b9e-ab6a-ed2e5c75d0bf\" />\r\n\r\nGranite Speech Plus is a variant of Granite Speech that enhances the projector by consuming the concatenation of the encoder's final hidden states with an arbitrary subset of its intermediate hidden states along the feature dimension. It is a multimodal speech-to-text model that can transcribe audio, provide speaker annotation and word level timestamps by responding to text prompts. The model inherits the same architecture components as Granite Speech including the speech encoder, query transformer projector, language model, and optional LoRA adapter.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/granite_speech_plus)\r\n* Support for a new Granite-Speech-Plus model (#45695) by @zvik in [#45695](https://github.com/huggingface/transformers/pull/45695)\r\n\r\n### Granite4Vision\r\n\r\nGranite Vision 4.1 is a vision-language model from IBM Research designed for enterprise-grade document data extraction. It specializes in chart extraction (Chart2CSV, Chart2Summary, Chart2Code), table extraction (JSON, HTML, OTSL), and semantic key-value pair extraction. The model builds on LLaVA-NeXT with architectural innovations including SigLIP2 Vision Encoder, Window Q-Former Projectors, and DeepStack Feature Injection with 8 vision-to-LLM injection points.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/granite4_vision)\r\n* Add Granite 4.1 Vision (granite4_vision) (#45597) by @artem-spector in [#45597](https://github.com/huggingface/transformers/pull/45597)\r\n\r\n### EXAONE-4.5\r\n\r\n<img width=\"3840\" height=\"2160\" alt=\"image\" src=\"https://github.com/user-attachments/assets/55eb732d-f9da-4f97-8226-2cd3f6476ca0\" />\r\n\r\nEXAONE 4.5 is the first open-weight vision language model developed by LG AI Research, integrating a dedicated visual encoder into the existing EXAONE 4.0 framework to expand multimodal capabilities. The model features 33 billion parameters in total, including 1.2 billion parameters from the vision encoder, and achieves competitive performance in general benchmarks while outperforming similar-sized models in document understanding and Korean contextual reasoning. It builds on EXAONE 4.0 with key enhancements including an expanded vocabulary of 153,600 tokens, support for up to 256K token context windows, and a Multi-Token Prediction (MTP) mechanism.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/exaone4_5) | [Paper](https://huggingface.co/papers/2604.08644) | [Blog Post](https://www.lgresearch.ai/blog/view?seq=641)\r\n* Add EXAONE 4.5 implementations (#45471) by @nuxlear in [#45471](https://github.com/huggingface/transformers/pull/45471)\r\n\r\n### PP-FormulaNet\r\n\r\nPP-FormulaNet-L and PP-FormulaNet_plus-L are lightweight models designed for table structure recognition, focusing on accurately recognizing table structures in documents and natural scenes. The models are part of the SLANet series and can be used for image-to-text tasks, specifically for detecting and processing mathematical formulas and table structures from images.\r\n\r\n**Links:** [Documentation](https://huggingface.co/docs/transformers/main/en/model_doc/pp_formulanet)\r\n* [Model] Add PP-FormulaNet Model Support (#45626) by @zhang-prog in [#45626](https://github.com/huggingface/transformers/pull/45626)\r\n\r\n\r\n\r\n## Breaking changes\r\n\r\nApex integration has been removed from the library (including RMSNorm usage in T5 and related models), so users relying on Apex for mixed precision or fused ops should migrate to PyTorch's native equivalents instead.\r\n* 🚨 Get rid of most Apex references (#45723) by @Rocketknight1\r\n\r\n\r\n\r\n## Tokenization\r\n\r\nFixed tokenizer mapping issues for DeepSeek R1 distilled (Qwen2) and DeepSeek OCR models, and resolved a significant performance regression in `PreTrainedTokenizer.convert_ids_to_tokens` where `skip_special_tokens=True` was rebuilding the special token set on every iteration, resulting in a ~300x speedup for that code path.\r\n\r\n\r\n* deepseek r1 distilled tokenizer fix for qwen2 mapping (#45741) by @itazap in [#45741]\r\n* DeepSeek OCR specifies an incorrect tokenizer class on the Hub (#45739) by @hmellor in [#45739]\r\n* PythonBackend slow tokenizer convert_ids_to_tokens fix (#45728) by @i3hz in [#45728]\r\n\r\n\r\n## Bugfixes and improvements\r\n\r\n* fix: correct spelling in continuous_api docstring (#45749) by @Dhruv908615 in [#45749]\r\n* Fix link to modular transformers documentation (#45746) by @SangbumChoi in [#45746]\r\n* Gemma4: fix failed test cases (#45568) by @kaixuanliu in [#45568]\r\n* Fix CI: Allow more artifacts to be download in CI (#45785) by @ydshieh in [#45785]\r\n* Add `concurrency` to `PR CI` workflow file (`pr-ci-caller.yml`) (#45786) by @ydshieh in [#45786]\r\n* Reorder decorators for autodoc and dataclass (#45702) by @zucchini-nlp in [#45702]\r\n* Unwrap `text_config` in `AutoModelFor*.from_config` (#45770) by @jamesbraza in [#45770]\r\n* fix: Added Mps support in float fallback backends list  (#45687) by @rigen1048 in [#45687]\r\n* Github Actions PR CI (caller) (#45476) by @ydshieh in [#45476]\r\n* make sure we call check_auto in CI (#45775) by @tarekziade in [#45775]\r\n* Fix auto mapping script (#45774) by @Cyrilvallez in [#45774]\r\n* [MINISTRAL3] Fix conversion script yarn's apply_scale support. (#45744) by @juliendenize in [#45744]\r\n* [nemotron_h] respect _no_reinit flag on dt_bias and out_proj.weight (#45591) by @vai-minzhou in [#45591]\r\n* fix(utils): Resolve backbone utils test regressions (#45594) by @harshaljanjani in [#45594]\r\n* [CB] Better overall script and decode bucketting (#45653) by @remi-or in [#45653]\r\n* [docs] model testing (#45152) by @stevhliu in [#45152]\r\n* update dev (#45726) by @vasqu in [#45726]\r\n* Doc translate to Persian(farsi)  (#45664) by @zeoses in [#45664]\r\n* [`OAI Privacy Filter`] Add integration test (#45725) by @vasqu in [#45725]\r\n* Speedup Qwen2VLImageProcessor (#45719) by @lgeiger in [#45719]\r\n* Remove dead beam-search dummies from dummy_pt_objects.py (#45722) by @jw9603 in [#45722]\r\n* chore(typing): add ty type checking for 10 utility files (#45703) by @moonbogi in [#45703]\r\n* Llama3 video fix (#45040) by @sywangyi in [#45040]\r\n* Fix custom-module copies inheriting read-only permissions (#45686) by @nurpax in [#45686]\r\n* Python code in model docs (#45608) by @zucchini-nlp in [#45608]\r\n* fix failed test cases for blt model (#45596) by @kaixuanliu in [#45596]\r\n* chore(typing): add ty type checking for 3 pipeline files (#45667) by @moonbogi in [#45667]\r\n\r\n## Significant community contributions\r\n\r\nThe following contributors have made significant changes to the library over the last release:\r\n\r\n* @artem-spector\r\n    * Add Granite 4.1 Vision (granite4_vision) (#45597)\r\n* @SindhuRaghuram97\r\n    * First model (#45788)\r\n* @nuxlear\r\n    * Add EXAONE 4.5 implementations (#45471)\r\n* @ArthurZucker\r\n    * Add DeepSeek V4 (#45643)\r\n* @remi-or\r\n    * [CB] Better overall script and decode bucketting (#45653)\r\n* @zhang-prog\r\n    * [Model] Add PP-FormulaNet Model Support (#45626)\r\n* @zvik\r\n    * Support for a new Granite-Speech-Plus model (#45695)","publishedAt":"2026-05-05T16:52:21.000Z","fetchedAt":"2026-05-05T17:01:15.701Z","url":"https://github.com/huggingface/transformers/releases/tag/v5.8.0","media":[],"coverageCount":0}],"pagination":{"nextCursor":"2026-05-05T16:52:21.000Z|2026-05-05T17:01:15.701Z|rel_oP71J5Z3r1csDka7MGv8h","limit":20},"summaries":{"rolling":{"windowDays":90,"summary":"Transformers shipped v5.0 as a major overhaul after five years, overhauling tokenization APIs and introducing dynamic weight loading with quantization support, while simultaneously accelerating a wave of multimodal and specialized model integrations. Gemma 4 arrived with vision capabilities handling variable image sizes via spatial 2D RoPE, VidEoMT landed as a lightweight video segmentation encoder achieving 5-10x speedups, and the library absorbed a steady stream of domain-specific architectures—from speech (VibeVoice ASR, VoxtralRealtime) and document understanding (PP-DocLayoutV3, UVDoc) to mixture-of-experts variants (EXAONE-MoE, GLM-5) and multilingual models (EuroBERT). Concurrent v5 RCs prioritized MoE performance optimizations using batched expert implementations and resolved tokenizer class enforcement issues by preferring the TokenizersBackend, while the v4 line stabilized with targeted fixes for model loading and generation methods.","releaseCount":12,"generatedAt":"2026-04-07T17:27:18.483Z"},"monthly":[{"year":2026,"month":3,"summary":"March shipped new model support across vision, audio, and language domains. VidEoMT brought lightweight video segmentation running at 160 FPS through query propagation across frames, while EuroBERT added multilingual encoding with an 8192-token context window. The month also integrated PaddlePaddle models, Mistral 4, and Jina Embeddings v3 alongside specialized models for document layout, speech recognition, and time-series forecasting.","releaseCount":2,"generatedAt":"2026-04-07T17:27:20.129Z"}]}}