releases.sh
Ollama

Ollama

$npx @buildinternet/releases get ollama
Recently Shipped23 releases · updated Aug 31, 2026

v0.33.0 shipped native Claude Desktop integration, letting users toggle individual local models on or off from the menu bar.

Claude Desktop gained first-class local model management.

  • The Apps view lists integrations with copyable setup commands
  • Local models appear alongside cloud options inside Claude; cloud-only models show when signed in
  • The proxy no longer interrupts in-flight requests when the model catalog refreshes (v0.33.2)

MLX inference expanded on Apple Silicon.

  • Added Qwen3.8 Flash Next support and structured output to the MLX runner (v0.33.1)
  • Fixed Metal GPU timeouts when loading models from slow storage

Prefill caching became more reliable.

  • Cancelled prefills now keep every checkpoint they crossed; retries resume where they stopped
  • Resumed prefills no longer emit restore points that don't cover what they claim — previously forcing 46k-token requests to reprocess from zero on recurrent models
  • A model metadata cache (v0.32.15) cut per-request overhead further

The macOS app stabilized. v0.33.2 restored system dark mode and fixed a bug that launched a second instance instead of handing off to the already-running app.

Still relevant from earlier releases: ollama launch codex-app runs OpenAI's Codex desktop app locally with a built-in browser and git support; /api/show caching cut integration latency ~6.7x; and the v0.30.x series completed the direct llama.cpp backend migration.

AI-generated summaries may contain mistakes.

Sources

Latest releases

See all releases

Activity

168 releases36h interval20/mo
Mon
Wed
Fri
SepOctNovDecJanFebMarAprMayJunJulAug
Less
More