Skip to content
AI News HubLIVE
Public articles 19Collected articles 20Trust 84Refresh 120 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-09-01ID ollama-blogStatus Enabled

Official local AI model runtime blog; confirm reuse terms before full body display.

Latest public articles

Ollama's transparent pricing

Ollama's Pro, Max, and Team plans now use industry-standard per-token pricing with usage included on every plan.

Ollama BlogIn-site articleOllama's transparent pricing

Claude Desktop support with Ollama

Claude Desktop can now be configured to work with Ollama as a third-party gateway provider, making it possible to use open models in Claude.

Ollama BlogIn-site articleClaude Desktop support with Ollama

NVIDIA Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is now available on Ollama. It's a 30 billion parameter (3B active) open model built for agents that stay running, gathering context, calling tools, and working through multi-step tasks on your own hardware.

Ollama BlogIn-site articleNVIDIA Nemotron 3.5 Lightning

Muse Glimmer from Meta Superintelligence Labs is now available

Meta's Muse Glimmer, the first open model released by Meta Superintelligence Labs, is now available. Muse Glimmer is a 30B multimodal model released under the Apache 2.0 license, designed for local coding agents, and accelerated by Ollama's MLX engine with new native DFlash and image input support.

Ollama BlogIn-site articleMuse Glimmer from Meta Superintelligence Labs is now available

Ollama: all aboard open models

Serving 8.9 million developers, Ollama has raised $88M from Benchmark, Theory Ventures, 8VC, Y Combinator, and many incredible angel investors.

Ollama BlogIn-site articleOllama: all aboard open models

Faster Gemma 4 on MLX with multi-token prediction

Gemma 4 is now significantly faster in Ollama 0.31 on Apple Silicon via multi-token prediction (MTP), powered by MLX. Performance improves up to 90% on coding-agent benchmarks.

Ollama BlogIn-site articleFaster Gemma 4 on MLX with multi-token prediction

Ollama's highest performance on Apple Silicon yet with MLX

Ollama's MLX engine has been updated to deliver its highest performance on Apple Silicon yet. By leaning more heavily on Apple's unified memory and the Metal-backed MLX framework, models output higher quality responses, respond faster, and use less memory. The update includes support for NVFP4 format, up to 20% faster output, and a snapshot system for agent workflows.

Ollama BlogIn-site articleOllama's highest performance on Apple Silicon yet with MLX

Improved performance and model support with GGUF

Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp, augmenting MLX on Apple silicon and supporting more models on wider hardware.

Ollama BlogIn-site articleImproved performance and model support with GGUF

NVIDIA Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is a 550 billion parameter (55B active) open model designed for long-running agentic workflows, with 1M token context and NVFP4 optimization, leading in agentic benchmarks and cost efficiency.

Ollama BlogIn-site articleNVIDIA Nemotron 3 Ultra

Ollama is now powered by MLX on Apple Silicon in preview

Ollama announces a preview release powered by Apple's MLX framework, delivering significant performance improvements on Apple Silicon, including NVFP4 support and enhanced caching.

Ollama BlogIn-site articleOllama is now powered by MLX on Apple Silicon in preview

Subagents and web search in Claude Code

Ollama now supports subagents and web search in Claude Code. No MCP servers or API keys required. Subagents can run tasks in parallel, keeping context clean. Web search is built-in via the Anthropic compatibility layer.

Ollama BlogIn-site articleSubagents and web search in Claude Code

OpenClaw: A Local AI Assistant for Coding

OpenClaw is a personal AI assistant that connects your messaging apps to local AI coding agents, all running on your own device for privacy.

Ollama BlogIn-site articleOpenClaw: A Local AI Assistant for Coding

ollama launch

Ollama introduces `ollama launch`, a new command that sets up and runs coding tools like Claude Code, OpenCode, and Codex with local or cloud models, without needing environment variables or config files.

Ollama BlogIn-site articleollama launch

Claude Code with Anthropic API compatibility

Ollama v0.14.0+ now supports the Anthropic Messages API, enabling tools like Claude Code to work with open-source models. Run locally or connect to cloud models via ollama.com.

Ollama BlogIn-site articleClaude Code with Anthropic API compatibility

OpenAI Codex with Ollama

Open models can be used with OpenAI's Codex CLI through Ollama. Codex can read, modify, and execute code in your working directory using models such as gpt-oss:20b, gpt-oss:120b, or other open-weight alternatives.

Ollama BlogIn-site articleOpenAI Codex with Ollama

OpenAI gpt-oss-safeguard

Ollama partners with OpenAI and ROOST to launch the gpt-oss-safeguard reasoning models for safety classification. Available in 20B and 120B sizes under Apache 2.0 license, these models support custom policies, interpretable reasoning, and configurable effort.

Ollama BlogIn-site articleOpenAI gpt-oss-safeguard

MiniMax M2

MiniMax M2 is now available on Ollama's cloud. It is a model built for coding and agentic workflows, with 10 billion activated parameters (230B total). It ranks #1 among open-source models in composite intelligence benchmarks and excels at multi-file edits, agentic tool use, and long-horizon task execution.

Ollama BlogIn-site articleMiniMax M2

All sources