Releases Index

Inference

Optimized inference servers for text and embeddings

$npx @buildinternet/releases get inference
Mon
Wed
Fri
OctNovDecJanFebMarAprMayJunJulAugSepOct
Less
More
Releases1Avg Interval—Avg Cadence4/mo

Adds minimal AMD ROCm GPU support with a ROCm flash-attn varlen build, triton layer norm, and an AMD Dockerfile, plus ARM64 and multi-arch CUDA Docker images including sm_121 for DGX Spark GB10. Also adds harrier-oss-v1 model support and fixes unbounded buffer allocation in PaddedBatch.from_pb, DebertaV2 large-batch runs, max_position_embeddings handling in NomicBertConfig, and single predict request metrics.

Read more →

🔧 Fixed Intel MKL Support

Since Text Embeddings Inference (TEI) v1.7.0, Intel MKL support had been broken due to changes in the candle dependency. Neither static-linking nor dynamic-linking worked correctly, which caused models using Intel MKL on CPU to fail with…

Read more →

Noticeable Changes

Qwen3 was not working fine on CPU / MPS when sending batched requests on FP16 precision, due to the FP32 minimum value downcast (now manually set to FP16 minimum value instead) leading to null values, as well as a missing to_dtype call on the…

Read more →
Latest
Sep 15, 2026