2026-05-06 00:00 UTCOriginal source2 min readUpdated: 2026-06-27 00:25 UTC

SpecMD: A Comprehensive Study on Speculative Expert Prefetching

SpecMD is a standardized framework developed by Apple researchers to benchmark and evaluate expert caching policies for Mixture-of-Experts (MoE) models. The study reveals that MoE expert access patterns do not follow temporal locality, leading to the proposal of a new eviction policy called Least-Stale, which reduces collision misses by up to 85× compared to LRU and achieves 88% hit rates with 34.7% reduction in time-to-first-token on OLMoE.

SourceApple Machine Learning Research

Article intelligence

EngineersAdvanced

Key points

SpecMD provides a standardized benchmarking framework for MoE expert caching policies across different hardware configurations.
The study finds that MoE expert access patterns are inconsistent with temporal locality assumptions like LRU and LFU.
The proposed Least-Stale eviction policy exploits predictable expert access patterns to reduce collision misses by up to 85×.
With just 5% VRAM cache capacity (0.6GB), SpecMD achieves over 88% hit rates and 34.7% TTFT reduction on OLMoE.

Why it matters

This matters because specMD provides a standardized benchmarking framework for MoE expert caching policies across different hardware configurations.

Technical impact

May affect model selection, inference cost, product capability, and evaluation benchmarks.

This panel is AI-generated and reviewed for accuracy.

SpecMD: A Comprehensive Study on Speculative Expert Prefetching - Apple Machine Learning Research

Machine Learning Research

Open MenuClose Menu

Overview

Research Highlights

Publications

Events

Work with us

research area Methods and Algorithms, research area Tools, Platforms, Frameworksconference ICML

content type paperpublished May 2026

SpecMD: A Comprehensive Study on Speculative Expert Prefetching

AuthorsDuc Hoang, Ajay Jaiswal, Mohammad Samragh Razlighi, Minsik Cho

View publication

Copy Bibtex

Mixture-of-Experts (MoE) models enable sparse expert activation, meaning that only a subset of the model’s parameters is used during each inference. However, to translate this sparsity into practical performance, an expert caching mechanism is required. Previous works have proposed hardware-centric caching policies, but how these various caching policies interact with each other and different hardware specification remains poorly understood. To address this gap, we develop SpecMD, a standardized framework for benchmarking ad-hoc cache policies on various hardware configurations. Using SpecMD, we perform an exhaustive benchmarking of several MoE caching strategies, reproducing and extending prior approaches in controlled settings with realistic constraints. Our experiments reveal that MoE expert access is not consistent with temporal locality assumptions (e.g LRU, LFU). Motivated by this observation, we propose Least-Stale, a novel eviction policy that exploits MoE’s predictable expert access patterns to reduce collision misses by up to 85× over LRU. With such gains, we achieve over 88% hit rates with up to 34.7% Time-to-first-token (TTFT) reduction on OLMoE at only 5% or 0.6GB of VRAM cache capacity.