AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
All posts Back Engineering September 23, 2026 •20 minute read How to serve trillions of tokens for trillion-parameter coding agents Janelle Cai Member of Technical Staff @JanelleCai Charles Frye Member of Technical Staff @charles_irl James Liu Member of Technical Staff @JamesLiuID Timothy Feng Member of Technical Staff Richard Gong Member of Technical Staff @_gongy Software is eating the world, and agents are eating software engineering. It is imperative that software engineers develop an understanding not just of the agent software that is now their most important tool but also of how the intelligent core of agent software works, through inference by large generative models of language — if not out of the engineer’s need to understand and control their tools, then at least because inference is poised to consume more computing power and produce more benefit than all other uses of computers. The central fact about inference services for coding agents is that they must operate at extremely high relative and absolute performance. By relative performance, we mean large fractions of the peak rate or “speed of light” of the hardware that it uses. By absolute performance, we mean that the scale of that peak rate and the amount of work done per request is large. Contemporary matrix math accelerators like Tensor Cores operate at the petaFLOP per second scale. Large generative sequence models with sufficient intelligence to automate software development have trillions of floating point parameters, and each of them must be accessed many times per second, even when serving just a single request. Due to these requirements, economically viable coding agent inference services are currently only feasible by operating at a scale sufficient to amortize hardware and engineering costs — roughly, at the scale of trillions of input and output tokens. We’ve done this, and we’d like to share how. At Modal, we operate a number of such inference services for coding agents at this scale and work with a number of customers who do the same. You can use our services indirectly via inference routing platforms like OpenRouter or Vercel AI Gateway or directly through our Shared Endpoints. In this blog post, we will walk through how we optimized inference performance when serving inferences from Moonshot AI’s Kimi K2.6 model to power coding agents. Though this model is “old” by this field’s standards (literally hundreds of days old!), the fundamentals of sequence modeling, hardware, and scaling change slowly enough that the core story and many of the details match what we have done for more recent models that have superseded K2.6 in intelligence and cost-performance, like Kimi K3. Our optimizations allowed us to scale per-replica performance of inference replicas by 2.8x per user and 5.6x across users on the replica: This chart relates the individual user’s experience (decode tokens per second per user, aka interactivity) on the x-axis with the cost-performance of the overall system on the y-axis (total tokens per minute per GPU, aka token throughput), with the number of concurrent users indicated at each point. More intuitively, that’s the difference between a ruinously expensive service with the UX on the right below and a price-competitive service with the UX on the left: We then scaled those single-container replicas into deployments and services. One particular service processed hundreds of billions of tokens a day and trillions in aggregate: Below, we aim to make this performance engineering legible to a general software engineering audience. By sharing how we, and our customers, are able to operate these services, we hope it enables you to do the same — perhaps by deploying a Dedicated Endpoint on Modal. First, understand the workload. We break this down into two sections: understanding the sequence model that infers the response to each request and understanding workload structure across requests. State-of-the-art coding agents are supported by trillion-parameter neural sequence models that process input in parallel and infer output sequentially. Contemporary coding agents are powered by probabilistic generative models of unicode sequences pre-trained mainly via unsupervised masked sequence prediction and post-trained mainly by reinforcement of output software correctness. Like the parser of a compiler, they operate not on raw strings but on tokenized sequences, so we call their inputs and outputs tokens. Because we are, in the end, guessing what output tokens should be, this is called inference. If you prefer deduction, stick to databases and operating systems. The underlying sequence models these days are hybrid-attention, mixture-of-experts Transformer neural networks. These networks apply computations both per token in the sequence and across tokens in the sequence. Attention has evolved into a generic term for cross-token computation. Mixture-of-experts refers to the dynamically routed block-sparse matrix multiplication that applies the majority of the per-token computation. These computations iteratively update the network’s internal, or latent, representation. For insight into what’s going on inside these sequence models, see Anthropic’s “A Mathematical Framework for Transformer Circuits” (2021, but still undefeated). A single forward pass through such a neural network produces both substantial internal state and a probability distribution over the next token(s) in the sequence for each sequence position. Because we predict (”regress”) based on our own outputs (”auto”), this is autoregressive sequence modeling. To respond to a client request, we generally chain multiple forward passes together like this: Forward passes are expensive, so we want to amortize this work as much as possible. Much of the work in per-token computation amortizes by batching several sequences together. Much of the work in cross-token computation amortizes by caching the internal state. For historical reasons, this is called the key-value cache (KV cache or just KV), even though contemporary models like Kimi don’t have distinct keys and values. You can read more about the “napkin math” here in Kipply’s excellent “Transformer Inference Arithmetic” blogpost (2022, but still undefeated). When a forward pass processes a request’s input tokens, we call it a prefill, because it is “prefilling” the KV cache. When a forward pass produces a response’s output tokens, we call it a decode, because we are “decoding” the model’s “encoding” of past state into predicted future. What about forward passes that do both? Yeah, we don’t like the terminology either. Prefill performance is mostly tracked by the latency to complete all prefills for a request, aka time-to-first-token (TTFT). Decode performance is mostly measured by the rate at which output tokens are produced after that, aka output tokens per second (TPS). Both can be measured client-side or server-side, causing no end of confusion. The particular sequence model covered in this post is Kimi K2.6 by Moonshot AI. This model parametrizes its matrix multiplications with approximately one trillion numbers (weights in its matrices), the majority of which are stored as four bit integers (INT4). We serve the model, however, with four bit floating point numbers (FP4). Four bits only gives you sixteen distinct values, so you further need a micro-scaling format to scale individual blocks within tensors independently. We chose the NVFP4 micro-scaling format, which has native hardware support at the petaFLOP/s scale in the Tensor Cores of Blackwell Streaming Multiprocessor Architecture GPUs like the B200 and B300. Because we operate a dynamic GPU fleet in a time of constrained compute supply, we prepare our deployment to run on both B200 and B300 GPUs. Results below are all for B200 GPUs; B300s are substantively similar but operate at higher request concurrency because they have more high-bandwidth memory (HBM) available for caching. We chose the SGLang inference engine as our base. We found several opportunities to improve performance by patching the engine. As contributors to the SGLang project, we upstreamed these patches, described and linked in the post below. To optimize UX and cost-performance, you must understand the structure of these sequences across requests. When you serve such models on coding agent traffic naïvely, you get bad results. This chart indicates that throughput and interactivity rapidly collapse above 6 concurrent users. Furthermore, even before that peak, the interactivity is below user expectations and the system is below acceptable efficiency. So from here, you need to increase interactivity and throughput to deliver better outcomes to users while decreasing your own costs. To do that, you need to understand the sequences in this workload deeper than just “tokens in and tokens out”. Individual requests for output tokens are created in “sessions”: the user, the generative model, and the tool calls chain together iteratively to construct a tower of input sequences, accumulating context — and value — over time. The iterative process of meaning construction, information discovery, and sense-making strikes us as fundamental to the nature of sequence modeling and sequential action, so we expect this pattern to far outlast “coding agents”. Concretely, a single session looks something like this: That is, the input sequence (green) for each turn T is the entire session history up to T (darker green), plus something new (lighter green). This has two key consequences. First, it means requests inherently have long input sequences relative to their output sequences (pink, above) — there are T-1 past output sequences in the input to turn T, and T is in the dozens. For the core workload we used in optimization and served in production, this ratio was 200:1; requests contain roughly 100k input tokens and produce roughly 500 output tokens. That means the majority of processed tokens will be input tokens (just check the token usage numbers in your coding agent software). Second, it means the input sequences have high overlap with previously processed input sequences — the ones from turns 1 to T-1. That means that on the way to serving turn T, the tokens in turn 1 are processed T times. This makes caching absolutely critical — we can avoid linearly-scaling recomputation to save effort, but we introduce linearly-scaling state that must be managed and has its own performance characteristics. Navigating this tradeoff is the core engineering problem we’ll tackle in this post. With this picture of the workload in mind, we turn to optimization. Then, optimize a single replica. To optimize performance, build a working system, identify the bottleneck, then lift it. Repeat as needed until you’ve won. Though our ultimate goal was to optimize an entire service, we decomposed that problem into two simpler problems: optimize a single replica first, then scale from one to many replicas. We further split the problem of single replica performance into two sub-problems: first maximize interactivity, then maximize throughput without losing interactivity. Interactivity primarily impacts request latency. Request latency and throughput interact through concurrency, the number of in-flight requests, by a rearrangement of Little’s Law: Our key bottlenecks for latency, concurrency, and throughput started in the GPU HBM. Our key bottleneck on latency was HBM bandwidth during decode. We lifted it by parallelizing matrix multiplication across GPUs (tensor parallelism, TP) and by applying custom DFlash speculative decoding — doing more computation per memory load, even when that computation may not be needed. That created a bottleneck on concurrency through HBM capacity: how much work can we keep in a cache that loads faster than we could just recompute results. We lifted it by clearing up intermediates in HBM [truncated for AI cost control]