vLLM Recipes
vLLM Recipes — Deploy any model on any hardware with vLLM Loading... Latest recipesnewest 8View all 153 → MiniMax-H3 MiniMaxAI 64Bbf16omni Open-weight general-purpose multimodal generation model — jointly generates 24 F…
vLLM Recipes — Deploy any model on any hardware with vLLM Loading... Latest recipesnewest 8View all 153 → MiniMax-H3 MiniMaxAI 64Bbf16omni Open-weight general-purpose multimodal generation model — jointly generates 24 FPS video with native stereo audio from text, image, video, and audio references, served via vLLM-Omni DeepSeek-V4-Flash deepseek-ai 284B/13Bfp81M ctxtext DeepSeek V4 MoE model with hybrid CSA+HCA attention, manifold-constrained hyper-connections, and three-tier reasoning (Non-think / Think High / Think Max). Inkling-Small thinkingmachines 276B/12Bnvfp41M ctxmultimodal Natively multimodal 276B-parameter MoE from Thinking Machines Lab — 12B active parameters, text/image/audio in, text out, and up to 1M context. Macaron-V1-Coding-Venti mindlab-research 743B/39Bbf161M ctxtext Macaron-V1-Coding-Venti — MindLab's coding-specialist checkpoint: the Macaron-V1-Venti L2 Coding LoRA merged into the GLM-5.2 BF16 base. Same MoE architecture and launch as GLM-5.2 (~743B total, 39B active), no runtime adapter. Kimi-K3 moonshotai 2.8T/16 experts/token + shared (of 896 routed)mxfp41M ctxmultimodaltext Pre-release 2.8T-parameter native multimodal MoE with Kimi Delta Attention, Gated MLA, Attention Residuals, and a 1M-token context window Laguna-S-2.1 poolside 118B/8Bbf16262K ctxtext Poolside's 118B total / 8B activated MoE coding model with mixed sliding-window + global attention, native interleaved reasoning, and 256K context — the larger sibling of Laguna XS-2.1, tuned for agentic coding. PaddleOCR-VL-1.6 PaddlePaddle 0.9Bbf16131K ctxmultimodal PaddleOCR-VL-1.6 (0.9B) — region-aware data optimization + progressive post-training; new SOTA 96.33% on OmniDocBench v1.6, drop-in replacement for 1.5 MOSS-Transcribe-Diarize OpenMOSS-Team 0.9Bbf16multimodal OpenMOSS's 0.9B end-to-end multi-speaker long-audio transcription model with timestamps and speaker labels, served through vLLM's OpenAI-compatible /v1/audio/transcriptions API. Browse by provider Arcee AI arcee-ai 1 recipe→ Ernie (Baidu) baidu 3 recipes→ Boson AI bosonai 1 recipe→ Seed (ByteDance) ByteDance-Seed 1 recipe→ DeepSeek deepseek-ai 9 recipes→ Fish Audio fishaudio 1 recipe→ Google Google 7 recipes→ inclusionAI inclusionAI 5 recipes→ InternLM internlm 2 recipes→ JetBrains JetBrains 2 recipes→ Jina AI jinaai 2 recipes→ Liquid AI LiquidAI 10 recipes→ LongCat (Meituan) meituan-longcat 1 recipe→ Meta meta-llama 3 recipes→ Microsoft microsoft 1 recipe→ MindLab Research mindlab-research 1 recipe→ MiniMax MiniMaxAI 6 recipes→ Mistral AI mistralai 7 recipes→ Moonshot AI moonshotai 7 recipes→ NVIDIA nvidia 11 recipes→ OpenAI openai 2 recipes→ MiniCPM (OpenBMB) openbmb 3 recipes→ InternVL (OpenGVLab) OpenGVLab 1 recipe→ OpenMOSS OpenMOSS-Team 6 recipes→ PaddlePaddle PaddlePaddle 3 recipes→ Preferred Networks pfnet 2 recipes→ Poolside poolside 4 recipes→ Qwen Qwen 23 recipes→ Stability AI stabilityai 2 recipes→ StepFun stepfun-ai 2 recipes→ Hunyuan (Tencent) tencent 4 recipes→ Thinking Machines Lab thinkingmachines 2 recipes→ Wan (Alibaba) Wan-AI 1 recipe→ MiMo (Xiaomi) XiaomiMiMo 3 recipes→ GLM (Z-AI) zai-org 14 recipes→