Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks! While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning. This year we are proud to feature the work of Alex Zhang of MIT. From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. RLMs took over the timeline early this year: and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra: and is even today, influencing new research that has more extreme implications than RLMs: We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie. We discuss: Why AI-generated GPU kernels still leave substantial room for human expertise How one expert insight can potentially replace enormous amounts of brute-force token search Why PhD students should take research bets that initially look trivial, weird, or pointless What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste GEV and why a language model does not have to mean an autoregressive text-to-text decoder Why Claude Code, Codex, and Pi are structurally more similar than they look How harness design can improve compositional generalization across tasks and domains RLMs: context offloading, code execution, recursive subagents, and shared memory Prime Agent, continual harnesses, and persistent agent-to-agent communication Why the model you query in the future may secretly be an entire swarm or scaffold OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving Why much of an agent swarm may be wasted search — and why convergence is still hard Kimi versus OpenAI and different approaches to multi-agent systems Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work Why current frontier models may already have a large capability overhang Speculative programmatic tool calling and overlapping tool execution with generation Whether English, code, or an entirely new “Neuralese” constrains how models reason AI for science, fast-moving benchmarks, and how Alex chooses what research problems to bet on Alex Zhang Website: alexzhang13.github.io X: @a1zhang Timestamps 00:00:00 Introduction 00:00:49 GPU Mode, KernelBench, and AI-Written Kernels 00:07:38 Human Expertise vs. Brute-Force AI Search 00:13:20 Research Taste and Taking Big Bets 00:19:28 GEV and Rethinking the Language Model 00:29:03 Video Game Agents and the Harness Problem 00:31:01 Why Claude Code, Codex, and Pi Are So Similar 00:36:42 Harnesses as Compositional Generalizers 00:44:24 RLMs Explained 00:52:01 Prime Agent and Persistent Subagents 00:57:41 RLMs in the Wild 01:00:30 OpenAI Swarms and the Future of Language Models 01:07:26 Open-Endedness and Sakana AI 01:15:52 Kimi vs. OpenAI Agent Swarms 01:20:06 Capability Overhang and Speculative Tool Calling 01:28:19 Neuralese, Future Research, and AI for Science Transcript Introduction: Alex Zhang, RLMs, and GPU Mode Swyx [00:00:00]: All right, we’re here in the studio with Alex Zhang, I guess most famously of, RLMs, but you have a few other affiliations. Welcome to the show. Alex Zhang [00:00:12]: Yeah. Thank you for having me. Swyx [00:00:13]: Yeah. I guess GPU Mode as well? Alex Zhang [00:00:15]: Yes, GPU Mode as well. Swyx [00:00:16]: You were shepherded in by Mark Saroufim. Not everyone gets that kind of welcome. Alex Zhang [00:00:19]: Yep. Yeah. Yeah. I’m very close to all the people in GPU Mode, so yeah Swyx [00:00:23]: Yeah Alex Zhang [00:00:23]: We often end up working together in various capacities, like even beyond just GPU Mode itself, so. Swyx [00:00:29]: Yeah. Can we explain, so people who are not that close Don’t know about this. It’s just a-- it’s a Discord. It used to be focused on, I guess, CUDA Mode, and then generalized a little bit. it was started by Mark. Alex Zhang [00:00:41]: Yep. Swyx [00:00:42]: It was basically like. To me, it’s like the hiring pipeline of the PyTorch team. Swyx [00:00:45]: And then you left PyTorch. Alex Zhang [00:00:47]: Yep. Alex Zhang [00:00:49]: Yeah. So it used to be, I think, it started actually around when I was in college, in like 2023. I think it was started by Mark, Andreas, and Jeremy Howard. The original premise was just like, it was a GPU-- or it was a Discord dedicated to learning how to write GPU kernels, and they had, like, lectures. That was basically the extent of it. and I got interested in it because I was writing GPU kernels. It was actually out of, like, pure chance. I was interning at Snapchat at the time, and From CUDA Mode to GPU Mode Swyx [00:01:22]: Rexis. Alex Zhang [00:01:23]: Yeah, I was very bored with Rexis. So, they had a project where, like, they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper. Swyx [00:01:35]: Yes, we’ve covered it on Paper Club. Alex Zhang [00:01:36]: Yes, yeah. So I was interested in whether or not you could write specialized kernels for it at Snapchat. it didn’t. Nothing really came of it, but I joined GPU Mode. At the time, it was called CUDA Mode, I think for, like, legal reasons or something, they changed the name. But I met Mark, I met Matei, I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn, which was now what you see as the leaderboard today. But the general idea was like. I think all of us had this, like, intuition that GPU programming is, like, very similar to if you guys have done, like, competitive programming. It’s a, it’s a not. I don’t mean to say, like, they’re transferable skills. Popcorn, KernelBench, and Automating GPU Kernels Swyx [00:02:17]: You have constraints. You code golf a little bit. Alex Zhang [00:02:19]: Yep. Swyx [00:02:19]: Yeah. Alex Zhang [00:02:19]: Yeah. And there’s like. There’s actually a surprisingly small space of optimizations that people do. and there’s actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that, like, if you had enough data, like, in the same way that Codeforces, there’s millions of problems. If you could do this with GPU code, like, you could scale and automate kind of GPU kernel development, which for researchers is a huge deal. ‘Cause I think one of the bigger bottlenecks. Like, if you look at, like, Mamba, for example, like, they release the paper with kernels because otherwise, like, you can’t really use it in any meaningful way? And not everyone has, like, a Tri Dao on their team. So we’re very interested in this. KernelBench kind of spawned from that too, of like, can we get LLMs to automate, GPU kernel code? And I think that was like a. It was a very fun time. It was like between college and my PhD, and yeah, I had a really pleasant time doing stuff with GPU Mode. Now I kind of just help with the lectures sometimes. I’m not as involved, and I think in general, like, we don’t have as many competitions as we used to. But, yeah, I still keep in touch a lot with everyone there. Swyx [00:03:27]: Is there a friendly rivalry? Because I think the previous community that used to do this was like MLSys, MLPerf Alex Zhang [00:03:32]: Yeah Swyx [00:03:32]: Kind of thing. Is there a friendly rivalry? Is this like just new generation MLPerf, or what’s going on? Alex Zhang [00:03:38]: The nice thing about GPU Mode is that it is also a community in the sense that, like, a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that. And, like, the competitions are, like, somewhat, not secondary, but, like, you can participate in them to learn. I think with a lot of, like, MLSys, MLPerf kind of benchmarks, like, for the most part, like, only serious labs and companies participate. Like, seriously in them, at least. That was my understanding of it. I could be wrong. But I think also beyond GPU Mode now, one thing that has been really exciting is there’s a lot more websites and, like, people that work on hosting competitions. Like, I think there’s this. Alex Zhang [00:04:21]: I think there’s this website called, like, LeetGPU or something, and it’s like leet code for GPU problems. Swyx [00:04:26]: Wow. Alex Zhang [00:04:26]: There’s, like, other ones that I’ve. Like, we’ve, we’ve seen. Like, there’s many that have kind of spawned and, like, talked on GPU Mode, and like, it’s very exciting in the sense that I think GPU programming used to be super niche, like when I was interested in it. And the only reason I got interested in it was Tri Dao gave a talk at Princeton because he was, applying for faculty. he is faculty there now, but I listened to his talk on FlashAttention in like 2023, and I was like, “Wow, this is like the coolest thing ever.” Alex Zhang [00:04:56]: And I was like, “This is like. This is what everyone should be working on.” I guess, like, vLLM and stuff had come out too, and it was like, “Oh, we should be writing kernels.” But now it’s like, everyone writes kernels. Like, everyone. It’s, it’s. I think it’s actually almost saturated in some sense, as a field. AI-Written Kernels and the Verification Gap Vibhu [00:05:11]: Any interesting takes for people that wanna get into it? So I think one of the biggest news is GPT-5.6 Wrote more efficient kernels, so Terra and, Luna could be 80% cheaper. Alex Zhang [00:05:24]: Yeah. Vibhu [00:05:25]: And then we’ve seen other competitions where people are, like, setting records, and they’re like, “We’re doing some auto research loop,” and these are people that don’t have a background Alex Zhang [00:05:34]: Yep Vibhu [00:05:34]: In any kernel writing, right? Alex Zhang [00:05:36]: Yeah. So even on the GPU Mode leaderboard, if you look at, like, a lot of the recent problems, almost all the solutions are AI generated. However, you’ll notice on the leaderboard. So there’s this guy named Gauners who’s, like, a very, like, regular member of GPU Mode. We’ve always known for a long time that he’s, like, a super cracked, like, GPU kernel writer. One thing we discovered on this leaderboard is, like, almost ev-- Like, he also used AI to help him with these solutions, but for the most part, like, he helped prompt and move it in certain directions. we found that, like, his kernel was, like, basically the only one in, like, the top 10 that was actually stable in, like, actual, like. end-to-end systems. And it does bring into question, like, it’s not. Like, GPU kernels have a verification problem. Like, we’ve kind of known this. It’s been a problem since KernelBench was released. Like, there’s a lot of reward hacking that goes on. But you al-- Yeah, you also notice, like, the lines of code is a lot smaller, bu [truncated for AI cost control]
Academia is for Ambition — Alex Zhang, MIT
Summary
We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.
SourceLatent Space
Academia is for Ambition — Alex Zhang, MIT
Report an error
The correction channel is not available yet. You can copy the article reference below for later.
Correction instructions