本文にスキップ
AI News HubLIVE
サイト内リライト6 分で読了

翻訳待ち:Agentic kernels in production

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Model performance Agentic kernels in production Baseten's agentic kernel optimization framework cuts latency by 42.3% on Qwen-Image and by 15.2% on FLUX.2. Authors Brian Li Faraz Shahsavan Pankaj Gupta Last updated Augu…

ソースBaseten Blog
翻訳待ち:Agentic kernels in production
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Model performance Agentic kernels in production Baseten's agentic kernel optimization framework cuts latency by 42.3% on Qwen-Image and by 15.2% on FLUX.2. Authors Brian Li Faraz Shahsavan Pankaj Gupta Last updated August 28, 2026 Share TL;DR We’ve built an agentic kernel development framework that identifies model-level optimization opportunities, generates improved kernels, and validates them in our serving stack. On our current models, we’ve improved end-to-end latency by 42.3% on Qwen-Image, 15.2% on FLUX.2, and a 5.5% increase in tok/s on MiniMax M3. We’ve seen in recent years that agents have become surprisingly capable at kernel development, from ideation to generating kernels from scratch. Existing benchmarks such as KernelBench have made it easier to evaluate how well agents can optimize kernels on isolated general-purpose problems. However, there’s a gap between winning a kernel benchmark and shipping optimizations into production. A few reasons why: The best kernel configuration depends on the production workload. The kernel that wins on a general benchmark may lose on a specific deployment. Optimizations such as tile shapes, warp-specialization strategy, and CTA configurations respond differently to changes in tensor shape, batch size, sequence length, etc. Kernels like MoE and Attention make this especially visible. A faster microbenchmark doesn’t necessarily translate to a faster model. Once your changes are integrated, interactions with downstream dependencies like CUDA graph capture and multi-stream execution can wipe out kernel-level gains or even result in a regression. Optimizing kernels individually can miss higher-level opportunities. End-to-end traces often show that only a small subset of kernels have headroom for improvement. The lower-effort wins may come from restructuring the computation around them: fusing operations, eliminating redundant work, or removing pipeline bubbles. Integrating a new kernel into a production serving engine is nontrivial. Unlike modifying a standalone torch model, serving engines have interconnected execution paths and dependencies. New kernels must be wired into the correct path, replace existing computation cleanly, and remain compatible with the surrounding runtime. With this in mind, we’ve created a solution that bridges the gap between benchmarks and production. Given a model and serving engine, our framework can profile the full workload, reason about the best optimizations, and then generate and ship those kernels straight to production. The stack: model-level + kernel-level optimizations The optimization stack divides into two layers: ✕ Two-track optimization pipeline: model-level work profiles, ranks, and tests candidate changes for end-to-end latency improvements, and kernel-level work benchmarks kernels across variants. Both feed into engine integration and production. The dotted testing loop (microbenchmark, correctness check, ablation tests) determines whether a candidate gets archived or recorded as a dead end. Model-level optimization: Understands the full model workload, profiles where time is spent, and proposes changes such as fusion and redundant work elimination. Per-kernel optimization: Takes generated and other performance-critical kernels identified in the trace, explores several implementations in parallel, and iterates on the strongest candidate. The first layer helps expand the search space beyond one-for-one kernel improvements. Rather than only optimizing kernels in isolation, the framework can restructure the execution graph by removing redundant work, reducing intermediate materialization, or combining operations before generating and improving the underlying kernels. Learning across optimization runs Our framework also has a self-improving mechanism: kernels that pass correctness and end-to-end performance checks are retained as reusable candidates, while lessons from both successful and failed attempts are added to an evolving knowledge base alongside workload constraints and integration findings. This creates a self-improvement loop where each optimization iteration starts from accumulated experience, enabling the agent to generate stronger candidates and converge faster over time. ✕ Persistent knowledge architecture: successful optimizations get stored in the kernel database with their patch, test cases, benchmarks, and workload data, then feed into the next optimization loop. Both successful and unsuccessful optimizations get summarized into the knowledge base, with successes captured as reusable patterns and failures captured as caveats and root causes. Results + case studies Our initial experiment targeted diffusion models, namely Qwen-Image and FLUX.2 served with SGLang on B300 GPUs. The optimizations highlighted below were identified, proposed, and implemented entirely by our agentic framework. ✕ Median per-step denoising time across four model configurations (FLUX.2 FP8, FLUX.2 NVF4, Qwen-Image FP8, Qwen-Image NVFP4). Every model gets faster moving left to right, with Qwen FP8 showing the largest drop, from 245.6 ms to 141.8 ms. Optimizations on both models Optimization #1: Pre-packed FP8 scales The FP8 paths in Qwen-Image and FLUX.2 were wasting launches converting scale metadata into DeepGEMM’s required format before matrix multiplications. Constant weight scales were repeatedly repacked through sequences of small kernel launches. We eliminate this overhead by changing the main FP8 activation producers to emit packed scales directly while also moving weight-scale packing to model load time. The numerical computation is unchanged, so outputs remain bit-identical. For example, at FLUX.2 attention projections: ✕ Baseline ✕ Optimized Another example at the Qwen-Image feed-forward layer: ✕ Baseline ✕ Optimized The optimization reduced end-to-end latency by 7.3% on Qwen-Image and 6.1% on FLUX.2, with these gains persisting throughout the subsequent FP8 optimizations. Optimization #2: Fused QKV projection and epilogue Both models’ original attention paths compute the image query, key, and value projections independently, despite them all using the same input. This resulted in repeated activation quantization and GEMM setup throughout every attention block. The optimization merges the three FP8 projections into one GEMM, then fuses bias addition, QK normalization, RoPE, and writes to the joint image-text attention buffers in a single Triton epilogue. NVFP4 still uses separate Q, K, and V GEMMs, as each projection uses a different scale. ✕ Baseline ✕ Optimized Optimization #3: Normalization + quantization kernel fusion In both models, normalization previously produced a large BF16 tensor that the following quantization kernel immediately reads back. Thus, the fix was to simply fuse these together, eliminating the intermediate BF16 write-and-read round trip. ✕ Baseline ✕ Optimized On Qwen-Image, the fused kernel emits both the original BF16 result and pre-quantized FP8 activations for the QKV and feed-forward GEMMs. This reduces latency by 4.3% and creates the producer path used by the packed-scale optimization. On FLUX.2’s residual path, the fused kernel emits the normalized output, updated residual, packed E2M1 values, and swizzled E4M3 scales in one pass. This improves end-to-end latency by 0.7%. Qwen-Image Optimization #1: Bias absorption After the previous optimization, there are two standalone bias additions remaining after the attention and feed-forward output projections, which account for roughly 11% of Qwen’s FP8 step time. To address this, we fold each bias into the next fused operation (residual normalization scale and residual update), reducing latency by 5.2%. Optimization #2: CFG modulation cache Classifier-free guidance runs two denoiser passes at the same timestep. Each pass uses different conditioning (one receives the prompt ccc, while the other receives an empty or negative prompt ∅\varnothing∅). The previous implementation recomputed the same timestep-only image and text modulation branches in both passes: ϵcond=F(xt,t,c),ϵuncond=F(xt,t,∅)\epsilon_{\text{cond}} = F(x_t, t, c), \qquad \epsilon_{\text{uncond}} = F(x_t, t, \varnothing)ϵcond​=F(xt​,t,c),ϵuncond​=F(xt​,t,∅) ϵCFG=ϵuncond+w(ϵcond−ϵuncond)\epsilon_{\text{CFG}} = \epsilon_{\text{uncond}} + w\left(\epsilon_{\text{cond}} - \epsilon_{\text{uncond}}\right)ϵCFG​=ϵuncond​+w(ϵcond​−ϵuncond​) The noisy latent xtx_txt​ and timestep ttt are shared by both passes. The image and text modulation branches are functions only of the timestep embedding and fixed model parameters, not the prompt: et=Embed(t)e_t = \text{Embed}(t)et​=Embed(t) and so: mimage=Wimageet+bimage,mtext=Wtextet+btextm_{\text{image}} = W_{\text{image}} e_t + b_{\text{image}}, \qquad m_{\text{text}} = W_{\text{text}} e_t + b_{\text{text}}mimage​=Wimage​et​+bimage​,mtext​=Wtext​et​+btext​ Because these modulation branches depend only on ete_tet​ and fixed weights, their outputs are identical across the conditional and unconditional passes at the same timestep, making it cacheable. Prompt-dependent outputs like hidden states and attention are computed separately. Create cache key at DiT entry (the same timestep object is passed to both CFG branches): 1def _cfg_cache_optimization_enabled(active) -> bool: 2 return active 3 4# QwenImageTransformer2DModel.forward 5if _cfg_cache_optimization_enabled() and isinstance(timestep, torch.Tensor): 6 # Both CFG branches receive the same timestep tensor. 7 # Keep a reference to it so tensor identity can be used safely. 8 cache_key = { 9 "timestep": timestep, 10 "version": version_if_available(timestep), 11 } Cache the image and text modulation outputs in each block: 1# QwenImageTransformerBlock.forward 2cached = getattr(self, "modulation_cache", None) 3 4cache_hit = ( 5 cache_key is not None 6 and cached is not None 7 and cached["timestep"] is cache_key["timestep"] 8 and cached["version"] == cache_key["version"] 9) 10 11if cache_hit: 12 # Second CFG pass: reuse the cached outputs. 13 image_modulation = cached["image_modulation"] 14 text_modulation = cached["text_modulation"] 15 16else: 17 # First CFG pass: compute the modulation outputs. 18 image_modulation = image_modulation_GEMM(timestep_embedding) 19 text_modulation = text_modulation_GEMM(timestep_embedding) 20 21 # Cache them for the second CFG pass. 22 if cache_key is not None: 23 self.modulation_cache = { 24 "timestep": cache_key["timestep"], 25 "version": cache_key["version"], 26 "image_modulation": image_modulation, 27 "text_modulation": text_modulation, 28 } This contributes to a reduction in latency of 2.1% for FP8 and 3.1% for NVFP4. Optimization #3: Per-kernel optimization We then run an optimization pass on performance-critical and previously fused kernels, producing the following improvements: These per-kernel optimizations improve 7.6% for FP8 and 13.4% for NVFP4. ✕ FLUX.2 Optimization #1: Single-block QK normalization + RoPE FLUX.2’s single-stream transformer block didn’t use the production fused QK-normalization and RoPE kernel because of a Python contiguity guard that rejected the merged-GEMM views. The fallback ran QK RMSNorm and interleaved RoPE as separate passes, which repeatedly concatenated the cosine and sine caches. The new replacement is a per-token-CTA kernel that loads each contiguous 12 KB Q/K head tile, performs RMSNorm in FP32, rounds the result to BF16, and applies interleaved RoPE in the same pass. It reads the cosine and sine tensors directly, eliminating 48 of 60 cache concatenations per step. ✕ Baseline ✕ Optimized The fused kernel offers a 2× speedup, resulting in end-to-end latency improvements of 2.3% for FP8 and 4.0% for NVFP4. Optimization #2: Fused SwiGLU + FP8/NVFP4 quantization Each invocation of SwiGLU previously produced a large BF16 intermediate that a separate FP8 or NVFP4 quantization kernel r [truncated for AI cost control]

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Model performance Agentic kernels in production Baseten's agentic kernel optimization framework cuts latency by 42.3% on Qwen-Image and by 15.2% on FLUX.2. Authors Brian Li Faraz…

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。