A new fused LM head, on by default in SFT, DPO, KTO, GRPO, RLOO and Distillation, projects hidden states through the LM head in Triton-kernel tiles to avoid building the full logits tensor, raising max trainable sequence length up to 6.9x and cutting peak memory 52% to 82%. The full-logits scoring path, use_liger_kernel chunked path and _forward_redirection are removed, so a PEFT adapter on lm_head now raises and use_liger_kernel=True is deprecated for v2.0.0 removal.
Fine-tuning
Libraries for efficient model fine-tuning and alignment
npx @buildinternet/releases get fine-tuningGRPO, RLOO, and Distillation on the default non-vLLM generation path now stop completions at every eos id the model declares, fixing silent training on tokens generated past the end of the turn that showed up as completions/clipped_ratio near 1. Also fixes DFT loss on batches with no trainable tokens, precompute_ref_log_probs under FSDP with no reference model, completion offsets where the tokenized prompt and prompt+completion diverge, and guards the requests and urllib3 imports in the vLLM client.
Fixes an issue that prevented encoder-decoder models from working when using Transformers 5.18.0 or later.
The trl.losses module introduced in v1.13 is removed, with DPO, KTO, and GRPO now streaming their own per-token log-probs through _ChunkedLogProbFunction, which also unlocks DPO with precomputed reference log-probs that did not exist before. A fused Triton logprob and entropy kernel now ships in-tree at trl.kernels and is on by default on CUDA, ROCm, and XPU, and six experimental trainers were removed from trl.experimental.
Riemannian LoRA, KaSA, Super-Tuning, ShadowPEFT added
Breaking (minor)Four new fine-tuning methods ship: a Riemannian-preconditioned LoRA optimizer, the KaSA LoRA variant, Super-Tuning sparsity allocation, and ShadowPEFT. Also fixed a LoRA scaling applied twice in the SVD path of add_weighted_adapter, sped up MoE target parameter computation, and reduced Orthogonal Subspace Fine-tuning memory 22% and training time 46%.
PPOTrainer, PPOConfig, and the value-head modeling modules are removed. The default chunked_nll lm_head projection now runs on bf16 tensor cores instead of fp32 SIMT, cutting a 23.37 ms fwd+bwd chunk to 3.86 ms (−6×) and peak memory from 5.99 GB to 3.03 GB, with end-to-end tokens/s gains of 1.2–1.7× across models. Vendored fused linear losses (DPO, KTO, GRPO, JSD) into trl.losses, added a long-context guide for 1M-token training, raised dependency floors (peft, deepspeed, vLLM), and added QLoRA test coverage.
DistillationTrainer and DistillationConfig move from trl.experimental.distillation to the top-level trl package, with the old path emitting a FutureWarning until removal in v2.0.0. AsyncGRPOTrainer gains an experimental loop-owning path for training external agents like opencode, plus VLM support in DistillationTrainer and new DiffusionGemma SFT example.
Fixes a crash when training with Liger kernel on pre-Ampere (Fermi, Kepler, Maxwell, Pascal) GPUs and corrects DAPO, CISPO, and VESPO loss normalization when steps_per_generation differs from gradient_accumulation_steps. Also fixes vLLM server-mode communicator initialization, queue wait time metric in AsyncGRPOTrainer, and prepare_deepspeed crash with CPU offload optimizer.
GRPO and RLOO trainers now support iterable datasets via a new repeat_iterable_dataset generator, and environments can own the data (train_dataset becomes optional). Message-level rollouts land in AsyncGRPO, and the DistillationTrainer received a major refactor including a fix for num_items_in_batch that counted the wrong tokens and NaN'd on prompt-only datasets.
KTO trainer graduates from experimental to the top-level trl package with the same API as DPO/GRPO/SFT, and the experimental import path still works with a FutureWarning. Environment-owned rewards let agentic RL environments define their own reward via a reserved get_reward() method, and multi-environment support allows a single training run to handle multiple environments with environment-specific tool schemas. GRPO now supports both static and adaptive entropy regularization to encourage exploration and prevent policy collapse.
Fixed a hang in GRPO + vLLM colocate + PEFT on non-NVLink hardware and corrected dataset fingerprinting in DPO/SFT tokenization. Also integrated the new response parsing API, added a prompt-learning guard for PEFT with Liger in GRPO, and fixed activation offload storage deduplication.
The default SFT loss_type is now "chunked_nll", delivering ~30% less peak VRAM on average with neutral or slightly faster wall-clock time. Also introduces experimental GMPO trainer, transformers continuous batching, AsyncGRPO weight sync with vLLM 0.22+, and paddding-free AsyncGRPO.
The release introduces a new experimental A2POTrainer for optimal advantage regression and grants KTO trainer support for vision-language models. The AsyncRolloutWorker now runs in a separate process to avoid GIL contention and potential NCCL watchdog timeouts, along with fixes for aiohttp retries and all-NaN reward columns. Gold distillation trainer now aligns tokens via byte offsets, and SDFT/SDPO leverage the vLLM server for live teacher logprobs. Other features include bidirectional masked importance sampling for IcePop, support for NemotronH and Nemotron 3 Ultra, additional training chat templates, and decoupled self-distillation trainers.
Trainer telemetry is now gated on an explicit class-name allowlist, restricting which trainer classes can send telemetry.