Adds minimal AMD ROCm GPU support with a ROCm flash-attn varlen build, triton layer norm, and an AMD Dockerfile, plus ARM64 and multi-arch CUDA Docker images including sm_121 for DGX Spark GB10. Also adds harrier-oss-v1 model support and fixes unbounded buffer allocation in PaddedBatch.from_pb, DebertaV2 large-batch runs, max_position_embeddings handling in NomicBertConfig, and single predict request metrics.
Text Embeddings Inference
npx @buildinternet/releases get text-embeddings-inferenceWhat's Changed
- Use
rust-toolchain.tomlbeforerustuponDockerfile-{cuda,cuda-all}by @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/842 - fix(backend): replace bare except with Exception in device check by @llukito in…
What's Changed
- Fix auto-truncate false setting by @vrdn-23 in https://github.com/huggingface/text-embeddings-inference/pull/836
- Set
pad_token_idas nullable & add support forrope_parametersby @alvarobartt in…
What's Changed
🚨 Fix
- Fix support for containers w/ CUDA 13.0+ by @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/831
When releasing ghcr.io/huggingface/text-embeddings-inference:cuda-1.9 with CUDA 12.9 and
cuda-compat-12-9there…
What's changed?
🚨 Breaking changes
- Default
HiddenAct::Geluto GeLU + tanh in favour of GeLU erf by…
What's Changed
Bug Fixes
- Fix error code for empty requests by @vrdn-23 in https://github.com/huggingface/text-embeddings-inference/pull/727
- Fix the infinite loop when
max_input_lengthis bigger thanmax-batch-tokensby @kozistr in…
🔧 Fixed Intel MKL Support
Since Text Embeddings Inference (TEI) v1.7.0, Intel MKL support had been broken due to changes in the candle dependency. Neither static-linking nor dynamic-linking worked correctly, which caused models using Intel MKL on CPU to fail with…
Today, Google releases EmbeddingGemma, a state-of-the-art multilingual embedding model perfect for…
Notable Changes
- Qwen3 support for 0.6B, 4B and 8B on CPU, MPS, and FlashQwen3 on CUDA and Intel HPUs -…
Noticeable Changes
Qwen3 was not working fine on CPU / MPS when sending batched requests on FP16 precision, due to the FP32 minimum value downcast (now manually set to FP16 minimum value instead) leading to null values, as well as a missing to_dtype call on the…
Noticeable Changes
Qwen3 support included for Intel HPU, and fixed for CPU / Metal / CUDA.
What's Changed
- Default to Qwen3 in
README.mdanddocs/examples by @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/641 - Fix Qwen3 by…
Notable change
- Added support for Qwen3 embeddigns
What's Changed
- Adding suggestions to fixing missing ONNX files. by @Narsil in https://github.com/huggingface/text-embeddings-inference/pull/624
- Add
Qwen3Modelby @alvarobartt in…
What's Changed
- [Docs] Update quick tour by @NielsRogge in https://github.com/huggingface/text-embeddings-inference/pull/574
- Update
README.mdandsupported_models.mdby @alvarobartt in https://github.com/huggingface/text-embeddings-inference/pull/572 - Back with…
What's Changed
- Enable intel devices CPU/XPU/HPU for python backend by @yuanwu2017 in https://github.com/huggingface/text-embeddings-inference/pull/245
- add reranker model support for python backend by @kaixuanliu in…
What's Changed
- feat: support multiple backends at the same time by @OlivierDehaene in https://github.com/huggingface/text-embeddings-inference/pull/440
- feat: GTE classification head by @kozistr in https://github.com/huggingface/text-embeddings-inference/pull/441 *…
What's Changed
- Download
model.onnx_databy @kozistr in https://github.com/huggingface/text-embeddings-inference/pull/343 - Rename 'Sentence Transformers' to 'sentence-transformers' in docstrings by @Wauplin in…
Notable Changes
- ONNX runtime for CPU deployments: greatly improve CPU deployment throughput
- Add
/similarityroute
What's Changed
- tokenizer max limit on input size by @ErikKaum in https://github.com/huggingface/text-embeddings-inference/pull/324
- docs:…
Notable Changes
- Cuda support for the Qwen2 model architecture
What's Changed
- feat(candle): support Qwen2 on Cuda by @OlivierDehaene in https://github.com/huggingface/text-embeddings-inference/pull/316
- fix(candle): fix last token pooling
Full Changelog:…


