跳到主要内容
AI News HubLIVE
站内改写6 分钟阅读

待翻译:Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

待翻译:Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

All posts Back Research September 24, 2026 •15 minute read Hitting a billion tokens per minute on one GPU by combining a query planner and an inference engine Charles Frye Member of Technical Staff @charles_irl Shreya Shankar Asst Professor, CMU FSD Lab @sh_reya I see it as a point on the LLM pareto optimal curve in a regime that had a large revealed latent demand (no thinking, single token, low latency acceptable intelligence) that was under-invested into because of a race to higher intelligence. - Karpathy-san, on Jev While everyone and their cousin is loudly building coding agents and chatbots, there’s a quieter inference revolution going on in the backend. Simple LLM transformations of data can be incredibly powerful, provided the cost-performance is good enough — just scroll social media and catch a few of the eye-popping, hack-inspiring demos of TypeSafe AI’s Jev model. Jev implements these transformations at what you might call the “JSON layer”, Web-style interfaces between clients and services. AI-SQL implements it at the analytic SQL layer, at the interface between business intelligence and the database: -- get hot leads with AI™ SELECT customers.id, products.id FROM customers JOIN products ON -- for each row in both tables AI.IF( -- run the prompt below and filter by truthiness PROMPT("{customers.profile} might buy this: {products.description}") ) Different inference applications produce different inference workloads, and AI-SQL is no exception. A query like the one above might produce millions of sequences of thousands of tokens — RIP your token budget. These queries often require much less than frontier intelligence, so small open-weights models can crush. But naïvely delivering these sequences directly to an inference engine optimized for agentic inference through interfaces for arbitrary user-controlled requests is inherently and massively inefficient. So we built an inference engine to fix this: the QUery-Aware Inference Layer (Quail). On one multi-join query where planning is particularly important, Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware. On Modal, that comes out to under 6¢ per billion tokens. On our newly-released benchmark for AI-SQL queries, Quail runs 1.84x faster than vLLM, geometrically averaged over tasks -- including two queries we designed to demonstrate areas for future improvement in AI-SQL inference. You can take it for a spin on Modal right now: # uvx modal run try_quail.py import modal app = modal.App("try-quail") image = ( modal.Image.from_registry("nvidia/cuda:13.0.1-devel-ubuntu24.04", add_python="3.12") .entrypoint([]) .apt_install("git") .uv_pip_install("quail-engine==0.1.0") ) @app.function(gpu="H100!", image=image, timeout=600) def run(sql=None, documents=None): from datasets import load_dataset import pyarrow as pa import quail if sql is None: sql = """ // no spoilers! SELECT r.id FROM reviews r WHERE AI_FILTER(PROMPT('Does this review discuss the ending?\n\n{0}', r.review)) """ if documents is None: imdb = load_dataset("stanfordnlp/imdb")["train"] documents = pa.table( { "id": pa.array(f"review-{i}" for i in range(len(imdb))), "review": imdb.data.table.column("text"), } ) config = quail.EngineConfig( gpus=1, model="qwen3-4b-fp8", backend="quail", device="h100-sxm", ) with quail.Session(config) as session: session.register( "reviews", quail.DocumentProvider.from_table(documents, id_col="id"), ) result = session.sql(sql).run() print(result.collect()) print(result.report) In this blog, we’ll give a quick overview of the problem we’re solving and how Quail works today. Spoilers: the big win is that with a structured query in hand, you can order requests to better cache (and evict) KV. This requires a slight revision of Hydragen-style cascade attention. Large numbers of small requests for small models can also incur lots of host overhead, aka have low GPU utilization, which can be avoided when you know the structure of the requests ahead of time. This was a collaboration between inference researchers at Modal and database researchers Carnegie Mellon University’s Full Stack Data Lab — call it a “mixture of experts”. We’re sharing what we did because we’d like to make this work more “expert-parallel”, as it were. We believe this is only the beginning for open source performance engineering at the intersection of inference and databases — two of the most important applications of computing. In this post, we’ll focus more on considerations for inference engineers. You can read more, from a database engineer’s perspective, at the Full Stack Data Lab blog. You can also check out the code for Quail here or the docs here. And if you run AI-SQL queries at scale and are interested in improving performance and cutting costs, get in touch with us. What are AI Functions and AI-SQL? First, a bit more background on the workload. This is emphatically not prompting AI systems to produce SQL based on natural language inputs — that’s NL2SQL. That looks a lot like a traditional chatbot or coding agent workload, so existing inference engines work well. It’s actually the other way around! In AI-SQL, we use an extension of SQL to programmatically produce (and consume) prompts for AI systems. Prompts are constructed from database entries and produce tables. Like this: -- get hot leads with AI™ SELECT customers.id, products.id FROM customers JOIN products ON -- for each row in both tables AI.IF( -- run the prompt below and filter by truthiness PROMPT("{customers.profile} might buy this: {products.description}") ) AI-SQL is primarily used inside of business intelligence (BI) platforms to help data scientists and stakeholders ask more “fuzzy” questions of their semi-structured data, like documents and free-text fields. There’s not a standard (yet), but major managed analytical database platforms have their own flavor: Snowflake Cortex AI-SQL, Databricks AI Functions, BigQuery AI functions. Unlike the rest of SQL, this problem actually needs GPUs. Consider the following plan for a query over the BioDEX dataset, which selects reports of serious adverse events in response to drugs that include both a neurological and a cardiovascular component: If these were normal filters and joins, say based on string matching and logical equality, there’d be no good reason to use a high-throughput numerical accelerator like a GPU, even though this is an analytical query, which seems “throughput-y”. This may be obvious to some, but let’s step through the logic anyway. Each byte loaded from durable storage to memory (or from memory to registers) would need at most a handful of arithmetic/logic operations to implement comparisons. GPUs are designed for workloads with high arithmetic intensity — many operations per byte loaded. And the latest GPUs have most of their arithmetic bandwidth in specialized hardware for large matrix multiplications, aka Tensor Cores. Normal filtering/joining requires no large matrix multiplications. But this query plan uses AI_FILTER and AI_JOIN, which instead pass the inputs through a large language model. An LLM is a sequence of numerical operations, the bottleneck for which is large matrix multiplications. Each byte loaded from durable storage will be subject to on the order of billions of operations before a byte is written to storage. Why is this interesting to inference engineers? Most inference engineering these days is focused on one workload shape in particular: iterative construction of long input sequences by users and tool calls external to the inference service. This is the shape of workloads from chatbots and agents — and of “rollout” inference during the reinforcement learning runs that fine-tune models to be chatbots or agents. Don’t get us wrong, this is very important work! We’ve written about our approach to it here. But for the hardcore inference engineer, it’s honestly starting to feel a little… played out. There’s also some work on ultra low-latency inference where speed matters as much as intelligence. We’ve written about our techniques for this here. In general, these workloads use structured outputs/tool-calling. They end up as something like the “OLTP” of inference, slotting into other computer applications more easily than open-ended agents. The recent popularity of Jev demonstrates the importance of these workloads — and that we are still so early! AI-SQL workloads haven’t gotten so much attention — yet — but we think they are interesting for inference engineers for a number of fundamental reasons, quite outside their importance to applications. Most intriguingly, they are an incredible fit for transformers (because they enable “perfect” KV cache use) and for transformers-on-GPUs (because they don’t require decode). Manage a KV cache without all the regrets. In typical inference, requests are client-controlled and arbitrary. This causes no end of pain. But in AI-SQL inference, clients only control SQL queries, which create many requests, and the combined query planner/inference engine has substantial control over the processing of those requests. This makes it particularly easy to operate a cache that amortizes more work. For instance, we know exactly when any cache entry is no longer needed, so we can fearlessly evict it. We also know quite a bit about what the cache demand will look like, since we get an entire query plan’s worth of requests up front. And we badly need caching for Transformers, because their forward passes are naïvely quadratic in the sequence length. We can exchange that for linear time and linear storage with KV caching. KV caches can be tricky to operate for agent workloads, because the time between accesses is completely unknown. But for an AI-SQL query, we control the inference engine requests and so can anticipate future accesses and apply optimizations like prefetching. And furthermore, because we are oriented to token throughput, we care less about latency to retrieve KV entries. This makes, for instance, operating a multi-tier KV cache much more feasible. Look, mom, no decode! Sequence model inference is split into two phases: “prefill”, when most of the KV cache is generated, and “decode” phase when most of the output tokens are generated. Decode is kind of a pain. GPUs aren’t particularly good at it. Decode has low arithmetic intensity so even though GPUs provide lots of memory bandwidth, it’s tricky to keep the arithmetic bandwidth saturated. Agentic applications skew heavy on the decode — even though there are more input tokens than output tokens, the decode is so much slower that it takes most of the time. This problem is so bad that inference service deployments are often forced to adopt complex solutions like cross-node prefill-decode disaggregation just to get acceptable perf. But not all tokens are generated during decode. The final “prefill” forward pass during input sequence processing emits a prediction for a single token. And for Boolean classification of a sequence, aka AI.IF, a single token is all you need — literally. This matters because AI.IF isn’t a sideshow. It’s how joins are implemented in AI-SQL (JOIN ON AI.IF). With a bit of cleverness in prompt construction, AI.CLASSIFY can be mapped onto a single token as well, for a number of classes up to the size of the vocabulary (we’ve left that one for future work!). Presently, we don’t take much advantage of this, except in what we don’t implement: Separate prefill and decode phases (let alone disaggregation), because there is no decode Sampling, because there are no generated tokens, only probabilities CUDA Graph capture, because prefills have long enough durations that launch overhead is negligible, even for small models on big GPUs Speculative decoding, because that accelerates decodes of more than one token But we anticipate deeper oppor [truncated for AI cost control]

展开要点与分析

文章情报

工程师进阶

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。