Adds minimal AMD ROCm GPU support with a ROCm flash-attn varlen build, triton layer norm, and an AMD Dockerfile, plus ARM64 and multi-arch CUDA Docker images including sm_121 for DGX Spark GB10. Also adds harrier-oss-v1 model support and fixes unbounded buffer allocation in PaddedBatch.from_pb, DebertaV2 large-batch runs, max_position_embeddings handling in NomicBertConfig, and single predict request metrics.
Inference
Optimized inference servers for text and embeddings
npx @buildinternet/releases get inferenceWhat's Changed
- Use
rust-toolchain.tomlbeforerustuponDockerfile-{cuda,cuda-all}by @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/842 - fix(backend): replace bare except with Exception in device check by @llukito in…
What's Changed
- Fix auto-truncate false setting by @vrdn-23 in https://github.com/huggingface/text-embeddings-inference/pull/836
- Set
pad_token_idas nullable & add support forrope_parametersby @alvarobartt in…
What's Changed
- misc(gha): expose action cache url and runtime as secrets by @mfuntowicz in https://github.com/huggingface/text-generation-inference/pull/2964
- feat: support max_image_fetch_size to limit by @drbh in…
What's Changed
Bug Fixes
- Fix error code for empty requests by @vrdn-23 in https://github.com/huggingface/text-embeddings-inference/pull/727
- Fix the infinite loop when
max_input_lengthis bigger thanmax-batch-tokensby @kozistr in…
What's Changed
- Add missing backslash by @philsupertramp in https://github.com/huggingface/text-generation-inference/pull/3311
- Revert "feat: bump flake including transformers and huggingface_hub versions" by @drbh in…
🔧 Fixed Intel MKL Support
Since Text Embeddings Inference (TEI) v1.7.0, Intel MKL support had been broken due to changes in the candle dependency. Neither static-linking nor dynamic-linking worked correctly, which caused models using Intel MKL on CPU to fail with…
Today, Google releases EmbeddingGemma, a state-of-the-art multilingual embedding model perfect for…
What's Changed
- [gaudi] Refine rope memory, do not need to keep sin/cos cache per layer by @sywangyi in https://github.com/huggingface/text-generation-inference/pull/3274
- Gaudi: add CI by @baptistecolle in…
Notable Changes
- Qwen3 support for 0.6B, 4B and 8B on CPU, MPS, and FlashQwen3 on CUDA and Intel HPUs -…
Noticeable Changes
Qwen3 was not working fine on CPU / MPS when sending batched requests on FP16 precision, due to the FP32 minimum value downcast (now manually set to FP16 minimum value instead) leading to null values, as well as a missing to_dtype call on the…
Noticeable Changes
Qwen3 support included for Intel HPU, and fixed for CPU / Metal / CUDA.
What's Changed
- Default to Qwen3 in
README.mdanddocs/examples by @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/641 - Fix Qwen3 by…
Fix for Neuron models exported with batch_size 1.
What's Changed
- [gaudi] gemma3 text and vlm model intial support. need to add sliding window … by @sywangyi in https://github.com/huggingface/text-generation-inference/pull/3270
- Neuron backend fix by @dacorvo in…
Neuron backend update.
What's Changed
- Remove useless packages by @yuanwu2017 in https://github.com/huggingface/text-generation-inference/pull/3253
- Bump neuron SDK version by @dacorvo in https://github.com/huggingface/text-generation-inference/pull/3260
- Perf opt by…
Notable change
- Added support for Qwen3 embeddigns
What's Changed
- Adding suggestions to fixing missing ONNX files. by @Narsil in https://github.com/huggingface/text-embeddings-inference/pull/624
- Add
Qwen3Modelby @alvarobartt in…
What's Changed
- [Docs] Update quick tour by @NielsRogge in https://github.com/huggingface/text-embeddings-inference/pull/574
- Update
README.mdandsupported_models.mdby @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/572 - Back with…
Gaudi improvements.
What's Changed
- upgrade to new vllm extension ops(fix issue in exponential bucketing) by @sywangyi in https://github.com/huggingface/text-generation-inference/pull/3239
- Nix: switch to hf-nix by @danieldk in…
This release updates TGI to Torch 2.7 and CUDA 12.8.
What's Changed
- change HPU warmup logic: seq length should be with exponential growth by @kaixuanliu in https://github.com/huggingface/text-generation-inference/pull/3217
- adjust the
round_up_seqlogic to align…