AI News HubLIVE
In-site rewrite6 min read

The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.

We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3: And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF: Three years ago, inference engineering barely existed as a category. Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem. baseten.com/inference-engi… ","username":"philipkiely","name":"Philip Kiely","profile_image_url":"https://pbs.substack.com/profile_images/1644827140641153024/ExLuda2F_normal.jpg","date":"2026-02-23T18:03:01.000Z","photos":[{"img_url":"https://substackcdn.com/image/fetch/$s_!1BR1!,w_1028,c_limit,f_auto,q_auto:best,fl_progressive:steep/l_play_button_usfui2,w_88,e_colorize:0/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fss-rehosttw-video-preview-13_2025989166333616128.jpg","link_url":"https://t.co/QTNdMrypqR"}],"quoted_tweet":{},"reply_count":190,"retweet_count":230,"like_count":2367,"impression_count":1396109,"expanded_url":null,"video_url":"https://video.twimg.com/amplify_video/2025989166333616128/vid/avc1/1280x720/fBFnlcAf_0wCVPNv.mp4","video_preview_media_key":"13_2025989166333616128","belowTheFold":false}" data-component-name="Twitter2ToDOM"> In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out. Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles. In this episode, Baseten’s Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API. We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model. The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them. We discuss: What happens when a 200,000-token request enters an inference system Cache-aware routing and reusing previously computed KV cache Why prefill and decode are increasingly handled by different GPUs When dedicated deployments become cheaper and more reliable than shared APIs How speculative decoding uses a smaller model to accelerate a larger one Tool calling, structured outputs, and what LLMs actually do What it takes to support a new open model on day zero Grafting Kimi’s vision encoder onto GLM-5.2 Retrofitting inefficient model layers with components from other architectures Why models sometimes collapse into repeating the same token How hardware, kernels, and race conditions create nondeterministic failures Preserving model fidelity while making inference faster How quantization errors can cancel each other out Why inference optimizations still deliver gains of 20%, 100%, and 200% How optimized serving can make a model up to 10× faster NVIDIA Dynamo, KV-aware routing, and distributed model serving Speculative decoding the speculative decoder Why local AI is about making models less dumb while data-center AI is about making them less slow Tensor, expert, and pipeline parallelism across GPUs Hardware-aware model design, auto-tuning, and the case against mega kernels Rubin and why inference is becoming a systems problem Whether modern GPUs are evolving into programmable AI ASICs Why enormous models like Kimi K3 require GB300-class hardware Why open-source video generation still trails Veo, Kling, and other closed models The quadratic attention bottleneck behind long-form AI video Autoregressive video, real-time generation, and compounding quality drift Why future video systems may combine autoregressive and diffusion architectures Training for inference and inference for training Continuous post-training, deployment, evaluation, and improvement loops How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself Why faster networking could unlock dramatically faster decoding Continual learning, KV-cache compaction, and persistent model memory Show Notes How to build a day-0 API for Kimi K3 22580: From GPT2 to Kimi3, Explained Philip Kiely LinkedIn: https://www.linkedin.com/in/philipkiely X: https://x.com/philipkiely Inference Engineering: https://www.baseten.co/inference-engineering/ Ali Taha LinkedIn: https://www.linkedin.com/in/aliestaha/ X: https://x.com/waterloointern Timestamps 00:00:00 Introduction and the 200K-Token Prompt 00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling 00:11:26 Launching Production-Ready Open Models 00:19:06 Model Retrofits, Failure Modes, and Nondeterminism 00:28:22 Quantization and Canceling Errors 00:32:15 The Race to 10× Faster Inference 00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI 00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels 01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips 01:10:03 Giant Models and the Limits of GPU Memory 01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation 01:21:47 Audio, Images, and Diffusion Models 01:27:32 Training, Self-Optimizing Models, and Continual Learning 01:40:06 Closing Thoughts Transcript Introduction: Baseten, Waterloo Intern, and Inference Engineering Swyx [00:00:00]: Okay, we’re here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you’ve done, you and I have done before, as well as Ali. Welcome. Ali [00:00:15]: Pleasure to meet you. Swyx [00:00:15]: Waterloo intern. Ali [00:00:16]: Waterloo intern, always. Swyx [00:00:17]: When did you get “Waterloo intern” as a handle? Ali [00:00:19]: As a handle? Oh. Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.” Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer. Philip [00:00:30]: So we have to figure out who’s gonna get the handle. Ali [00:00:33]: Well, I’ll pass the torch over to the next intern. Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad. Ali [00:00:37]: To another Waterloo intern. No, bruh. Philip [00:00:39]: Yeah. Ali [00:00:39]: Intern. Swyx [00:00:40]: Intern, yeah. Ali [00:00:40]: And no. Philip [00:00:41]: You gotta get an intern from Waterloo. Ali [00:00:42]: Yeah, I’ve gotta get an intern from Waterloo. Swyx [00:00:44]: Right. Ali [00:00:44]: But they have to follow the path. Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it’s like whoever Baseten gets from Waterloo. Ali [00:00:48]: Right. Swyx [00:00:49]: Has the title of Waterloo. Ali [00:00:50]: It stays in the ecosystem. Philip [00:00:51]: Exactly. Ali [00:00:52]: Halfway through the internship, you either get it or you’re out. Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle. Ali [00:00:59]: Just say it. Philip [00:00:59]: For everybody. Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you’re an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten’s inference? What’s the process of query through GPU model routing, balancing, all that? What is all the stuff that we don’t think about? Long Context Requests, KV Cache, and Cache-Aware Routing Philip [00:01:26]: With a long query specifically, the first thing that I’m gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it’s gonna be a lot easier for me and a lot cheaper for you. So the first thing that we’re gonna look at is some cache-aware routing, where we’re going to see, we probably have a number of instances, a number of replicas up serving whatever model you’re hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you’re doing two hundred thousand tokens, it’s probably coding or a multi-turn agent or something where you would expect to have that cached. If you don’t, we’re gonna have to send it to a prefill worker. We’ve at least on certain models disaggregated prefill and decode, so you’re going to have one set of GPUs that’s solely going to process the input, create the KV cache, and get you your first token, and then that’s going to be passed over to a separate set of GPUs, which is going to run decode. We’re going to iteratively make those tokens. We’re probably going to have some speculator model in front of that. I’m going to assume that you’re doing coding, and because of that, our speculator model, which assumes you’re doing coding, is gonna have a high draft token acceptance rate. If I’m wrong and you’re asking me to summarize every Harry Potter book, it’s gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?” Swyx [00:03:04]: Except Baseten doesn’t charge by pennies. Philip [00:03:07]: Well, yeah, we charge. I’m assuming that we’re talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it’s not pennies. Public APIs vs. Dedicated Deployments Swyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it’s up to you to figure out how to saturate the box. Ali [00:03:31]: And more often than not, it’s, like, way cheaper if you’re pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token. Philip [00:03:37]: Yeah, they do. I think that we’ve increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that’s really sticky, then they move over to dedicated. Swyx [00:03:51]: Is there a best practice on when it’s time to swap over? Philip [00:03:54]: Couple reasons. Yeah, reliability, that’s a big one, right? Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like th [truncated for AI cost control]