Skip to content
AI News HubLIVE
Original source2 min read

Ollama is now powered by MLX on Apple Silicon in preview

Summary

Ollama announces a preview release powered by Apple's MLX framework, delivering significant performance improvements on Apple Silicon, including NVFP4 support and enhanced caching.

Ollama is now powered by MLX on Apple Silicon in preview
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

Ollama is now powered by MLX on Apple Silicon in preview · Ollama Blog

Models Docs Pricing

Sign in Download

Models Download Docs Pricing

Sign in

Ollama is now powered by MLX on Apple Silicon in preview

March 30, 2026

Today, we’re previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple’s machine learning framework.

This unlocks new performance to accelerate your most demanding work on macOS:

Personal assistants like OpenClaw

Coding agents like Claude Code, OpenCode, or Codex

Accelerate coding agents like Pi or Claude Code

OpenClaw now responds much faster

Fastest performance on Apple silicon, powered by MLX

Ollama on Apple silicon is now built on top of Apple’s machine learning framework, MLX, to take advantage of its unified memory architecture.

This results in a large speedup of Ollama on all Apple Silicon devices. On Apple’s M5, M5 Pro and M5 Max chips, Ollama leverages the new GPU Neural Accelerators to accelerate both time to first token (TTFT) and generation speed (tokens per second).

Prefill performance

0 500 1000 1500 2000 tokens/s

1810 Ollama 0.19

1154 Ollama 0.18

Decode performance

0 40 80 120 160 tokens/s

112 Ollama 0.19

58 Ollama 0.18

Testing was conducted on March 29, 2026, using Alibaba’s Qwen3.5-35B-A3B model quantized to NVFP4 and Ollama’s previous implementation quantized to Q4_K_M using Ollama 0.18. Ollama 0.19 will see even higher performance (1851 token/s prefill and 134 token/s decode when running with int4 quantization).

NVFP4 support: higher quality responses and production parity

Ollama now leverages NVIDIA’s NVFP4 format to maintain model accuracy while reducing memory bandwidth and storage requirements for inference workloads.

As more inference providers scale inference using NVFP4 format, this allows Ollama users to share the same results as they would in a production environment.

It further opens up Ollama to have the ability to run models optimized by NVIDIA’s model optimizer. Other precisions will be made available based on the design and usage intent from Ollama’s research and hardware partners.

Improved caching for more responsiveness

Ollama’s cache has been upgraded to make coding and agentic tasks more efficient.

Lower memory utilization: Ollama will now reuse its cache across conversations, meaning less memory utilization and more cache hits when branching when using a shared system prompt with tools like Claude Code.

Intelligent checkpoints: Ollama will now store snapshots of its cache at intelligent locations in the prompt, resulting in less prompt processing and faster responses.

Smarter eviction: shared prefixes survive longer even when older branches are dropped.

Get started

Download Ollama 0.19

This preview release of Ollama accelerates the new Qwen3.5-35B-A3B model, with sampling parameters tuned for coding tasks.

Please make sure you have a Mac with more than 32GB of unified memory.

Claude Code:

ollama launch claude --model qwen3.5:35b-a3b-coding-nvfp4

OpenClaw:

ollama launch openclaw --model qwen3.5:35b-a3b-coding-nvfp4

Chat with the model:

ollama run qwen3.5:35b-a3b-coding-nvfp4

Future models

We are actively working to support future models. For users with custom models fine-tuned on supported architectures, we will introduce an easier way to import models into Ollama. In the meantime, we will expand the list of supported architectures.

Acknowledgments

Thank you to:

The MLX contributor team who built an incredible acceleration framework

NVIDIA contributors to NVFP4 quantization, NVFP4 model optimizer, MLX CUDA support, Ollama optimizations and testing

The GGML & llama.cpp team who built a thriving local framework and community

The Alibaba Qwen team for open-sourcing excellent models and their collaboration

© 2026 Ollama

Download Blog Docs GitHub Discord X (Twitter) Contact Privacy Terms

Blog

Download

Docs

GitHub

Discord

X (Twitter)

Meetups

Privacy

Terms

© 2026 Ollama Inc.

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • Ollama preview leverages MLX framework for fastest performance on Apple Silicon.
  • Supports NVFP4 quantization for higher quality and production parity.
  • Upgraded caching reduces memory, adds intelligent checkpoints and smarter eviction.
  • Optimized for Qwen3.5-35B-A3B model, suitable for coding tasks.

Highlights and analysis are generated automatically and may contain errors. Check the original source.