Ollama
$
npx @buildinternet/releases get ollamaRecently Shipped23 releases · updated Aug 31, 2026
v0.33.0 shipped native Claude Desktop integration, letting users toggle individual local models on or off from the menu bar.
Claude Desktop gained first-class local model management.
- The Apps view lists integrations with copyable setup commands
- Local models appear alongside cloud options inside Claude; cloud-only models show when signed in
- The proxy no longer interrupts in-flight requests when the model catalog refreshes (
v0.33.2)
MLX inference expanded on Apple Silicon.
- Added Qwen3.8 Flash Next support and structured output to the MLX runner (
v0.33.1) - Fixed Metal GPU timeouts when loading models from slow storage
Prefill caching became more reliable.
- Cancelled prefills now keep every checkpoint they crossed; retries resume where they stopped
- Resumed prefills no longer emit restore points that don't cover what they claim — previously forcing 46k-token requests to reprocess from zero on recurrent models
- A model metadata cache (
v0.32.15) cut per-request overhead further
The macOS app stabilized. v0.33.2 restored system dark mode and fixed a bug that launched a second instance instead of handing off to the already-running app.
Still relevant from earlier releases: ollama launch codex-app runs OpenAI's Codex desktop app locally with a built-in browser and git support; /api/show caching cut integration latency ~6.7x; and the v0.30.x series completed the direct llama.cpp backend migration.
AI-generated summaries may contain mistakes.
Sources
Latest releases
See all releasesActivity
168 releases36h interval20/mo
Mon
Wed
Fri
SepOctNovDecJanFebMarAprMayJunJulAug
LessMore