AI News HubLIVE
Original source1 min read

Qwen3.8-2.4T-A95B now available on Modal

Qwen3.8-2.4T-A95B by Alibaba, with a 1M token context window, is now available via Modal Auto Endpoints.

All posts

Back

News

August 12, 2026 •4 minute read

Qwen3.8-2.4T-A95B now available on Modal

Harmya Bhatt

Member of Technical Staff

@racerfunction

Gilford Ting

Member of Technical Staff

@gilfordting

David Wang

Member of Technical Staff

@_dcw02

Richard Gong

Member of Technical Staff

@_gongy

Will Hu

Member of Technical Staff

@_williamhu

Greta Workman

Product Marketing

@gretaworkman

Qwen3.8-2.4T-A95B just launched as an open weights model, and it’s now available on Modal.

Over Qwen 3.7, the latest model sees substantial improvement across coding, work, research, and long-horizon tasks.

We worked with Qwen ahead of the drop to bring day zero support to Modal Auto Endpoints, backed by SGLang and a custom DFlash speculator tuned to Qwen3.8’s shape.

Try it out now as a Shared Endpoint.

Speeding up inference with custom DFlash speculation

Just getting the model running is one thing, making it fast is another. For this, we once again turn to a DFlash speculator model because—say it with us now—Speculation is all you need.

A speculator only earns its keep when the target accepts the tokens it drafts, and acceptance comes down to whether the drafter has seen sequences like the ones it's predicting. Because of Max’s improvements in coding, research, and work (things that tend to use more tool calls) we leaned into that in our training data to increase accepts.

Try it now

Qwen3.8-2.4T-A95B (text only) is available for the next month as an OpenAI compatible Shared Endpoint with token-based pricing.