This episode is with Jean-Stanislas “JS” Denain of Epoch AI, who leads their Insights Team and is one of the people I find myself debating the state and trajectory of AI with more and more. We’ve had follow-on discussions of many of my favorite recent posts online and/or in private, so I wanted to dig into the nuance in a public episode.
A big takeaway of this podcast is how JS and I both have so much uncertainty with exactly where we are heading, and this was our best effort at stating our observations today.
Chapters / topics include:
00:00 Predictions for RSI
18:15 The role of robotics in an AI acceleration
24:20 How far behind are Chinese models?
27:39 Does distillation explain the gap?
40:58 What Chinese job postings reveal about their labs
48:13 Are open or closed models safer?
58:10 How Epoch AI ticks
1:00:55 What a frontier post-training recipe looks like
Enjoy!
More from JS: Epoch AI profile and writing, X, LinkedIn
Share
Listen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here.
Transcript
00:00:00 Nathan Lambert: I’m here with JS Denain, who is a senior researcher at Epoch AI. He leads the insights team. He is one of the people who I feel like I get the best feedback on my writing from, whether it’s from US-China AI capabilities, now RSI. And I just wanted to open this discussion and honestly go deeper with him, trying to understand how he thinks about these various things. And I think you have a very useful, moderate point of view, which I feel like you’re probably a step further into what would be called faster scenarios for AI progress. But let’s get into this, and it’s like, what measurements do you think OpenAI and Anthropic are seeing when we get all these proclamations on RSI happening very imminently?
00:00:47 JS Denain: Yeah. I think, so there’s the measurements they’ve published, right? So, OpenAI and Anthropic both had blog posts, I mean, Anthropic two at least, on the effect AI has on accelerating AI progress. I think at least the things they publish, I don’t think are super strong evidence of imminent self-sustaining acceleration AI capabilities, or full automation of the job of AI researcher. But I think the kinds of things we see are, I think probably the most striking thing I saw in the OpenAI blog post was increasing usage of AI systems in model deployment, like the increase in spending on Codex that we saw. And it’s kind of unclear how exactly to interpret this, because maybe it’s a measurement artifact where they’re only looking at Codex, but in fact, there was a bunch of ChatGPT usage before from the researchers. But overall, that plot, for example, just shows a 2X a month increase in Codex spending by researchers, and that does seem to me to be some evidence of they’re getting a lot of value out of this probably. I don’t think this is strong evidence that in six months we have a software intelligence explosion.
00:01:57 Nathan Lambert: Do you think this is the same? So, what is the information they have internally relative to what we have? And this is obviously hypothetical. We don’t have this internal information. Because I get the sense that a lot of people are more scared in their updates from the labs than the information we have. And I try to take this very seriously of, what will they be seeing that is making the acceleration of risk comments go faster, and how much of this is material evidence versus how cultures evolve over time? And I’m much more interested in evidence.
00:02:29 JS Denain: Yeah. So two things. I think, first of all, I guess I don’t think, I don’t know, right? I don’t have full information here. I don’t currently think that either there’s some specific thing that people at OpenAI or Anthropic are seeing right now that we don’t have access to that warrants being way more freaked out about this. I also don’t think that... I think the public evidence we have right now, more general on AI progress and just a priori case for this being an important dynamic, I think is enough. I think to care about this particular dynamic of AI accelerating AI progress, that being a big deal and worth tracking. And then it’s kind of unclear what the urgency is of when the feedback loop really kicks in. So, basically on the what is there on the inside that people have access to, I could give examples of kinds of metrics, right, they could be looking at. It’s plausible that we have access to the capabilities of AI systems, but internal teams have their KPIs, and maybe they’re seeing compute multipliers in the pre-training team or other kinds of metrics that people are tracking going crazy. And then the combination of this plus some intuitions of how the different outputs of different teams combine yields a prediction on the trend in actual performance of the end AI systems. So that could be an early warning sign. It’s unclear to me that the recent discourse we’ve seen is evidence of things going crazy on those metrics.
00:04:09 Nathan Lambert: And how do you think of the link between RSI and existential risk? So I would posit that you agree. I think that there are very real risks of AI, and I’m curious on how you think these, what I would describe as very, very early measurements change anything on the scope of risk. Because I don’t think if you had asked people six months ago, it would be as immediate to x-risk among people are very reasonable. I think there’s more people that are reasonable talking about x-risk again, which was a little surprising to me.
00:04:43 JS Denain: Yeah. So, okay, my sense is something like... So personally, I feel very uncertain about this, but I do feel, yeah, basically bought into there’s, I know Evan Hubinger was like, at least 10% of x-risk within I don’t know what timeframe. I think I’m like, yeah, I know, and this seems pretty reasonable over a decade-long timeframe. I just feel extremely uncertain about it, but I’m definitely very worried about this. Now, why am I worried about this, and where do I think the disagreements come from? And then how do I relate this to the early sense of RSI? My sense, I’m kind of a capabilities theory of everything person. I think, and I think some people disagree here, but I really think that principal component of disagreement between everyone is how huge do the capabilities get, how soon, of AI systems? And I sort of agree that there’s other factors that come in play for how big your economic growth gets, also depend on the diffusion you get. And you could possibly you could think that capabilities are going to get crazy, but the AI system’s going to be just aligned and benign and stuff, and so there’s no huge risk. But my sense is concretely, when I look at the main kinds of disagreements between people, most of the people who I see who are very skeptical of those most extreme scenarios... I think just expect capabilities to not be as huge or as I think folks who are—
00:06:18 Nathan Lambert: What does being a capabilities maximalist look like in a few years? Because I think I’m probably on the skeptical side, so please continue.
00:06:28 JS Denain: I think it looks, for example, something like the AI 2027 scenario, right? I think it looks like the mechanism for this is AI is automating the AI research process, I think, and that’s a reason to pay attention to it. But in terms of effect on the real world, I think it’s like massive progress on robotics. I think a huge industrial explosion, AI systems are just managing factories. You have this kind of self-sustaining economy that just is able to make a large scientific progress much faster than you would have expected. And so I think concretely, the kinds of disagreements I would expect are on, yeah, if you have AI systems that are both very intelligent in the book smart sense, but also have been trained to have more affordances and use them astutely, have been trained to kind of manage projects in efficient ways and stuff like that. How big are the real-world bottlenecks to making very fast R&D progress or getting hard power over humans?
00:07:26 Nathan Lambert: Yeah. Can we go into some of these in detail? Did you listen to the Dwarkesh podcast with Charlie, Beren, and John?
00:07:32 JS Denain: Yes. Yeah.
00:07:33 Nathan Lambert: Yeah, because they had at the end, they had this section on various capability levels and timelines for getting them. And I feel like I agreed with most... I was very in agreement on the distribution they had up to this, and then was surprised by the timelines. And one of them was the 10X productivity for the AI researchers. And I think Beren and John were faster than I think. And my kind of statement is that I think the cycle from of having an idea and doing the experimentation to test it, I agree will be 10X faster very soon. But I don’t necessarily agree that I would say that AI researchers will be 10X more productive in net, which I would describe as the pace of the field’s complete understanding. And understanding is a different axis from just continuing to scale models. I think that’s one of my core confusions on the AI research side. So I’m just kind of curious how you think about this type of thing and how you might specify a 10X improvement in AI research into subcategories.
00:08:37 JS Denain: Yeah. So maybe there’s a scale you could have here, which is the end thing that you might care about is how much faster is AI research overall? Or how much faster is Anthropic’s overall output? And then Anthropic as a company is producing some things, and it’s doing in one year what it would have taken it 10 years to do. And that’s pretty different from individual researcher productivities, where I think you could... So if most of what AI researchers right now are doing is this loop that you were describing, then it’s possible that you get a 10X productivity improvement for the median researcher based on the tasks they’re doing right now. But first of all, that doesn’t mean you get a 10X productivity improvement for all researchers. And even if you did, right, there’s other bottlenecks that hit such that that needn’t convert into a 10X productivity improvement for Anthropic as a whole, right? You could have all the researchers be 10X more productive, but because of compute or other things, the company itself still doesn’t move as fast.
00:09:38 Nathan Lambert: I think an analogy I have is, I think that junior PhD students will be 10X as productive, but from the advisor’s perspective, their research agenda will not proceed 10X as fast. And it’s like to the extent that that contributes to Anthropic’s progress is another hard thing to jump on AI capabilities, where I think that listening to the Noam podcast. This is me, I’m just thinking, was thinking about this when writing about this, is there’s such an amount of inference compute coming online, and I think Dwarkesh highlights this very well, that it’s very hard for me to disambiguate massive speed-up in AI research from the fact that we have way more compute and can now much more effectively spend it on related problems. And I think we’re going to get all of this at once.
00:10:27 JS Denain: Yeah. So definitely I think this question of... So when you look at the OpenAI blog post they had on their acceleration, right, they do point out huge surge in Codex spending from researchers. Interestingly, actually, Codex spending in other parts of the company kind of had a huge surge in the spring and then kind of plateaued in the summer. But for researchers or the data team or engineers, it keeps growing, even accelerates sometimes. And so they point this out, and then they try to look at where there was an increase, which kinds of tasks had an increase in usage. And a lot of them are engineering tasks. There’s an increase in troubleshooting tasks. But they definitely point out that for a lot of the high-level strategic decision-making, they both anecdotally and also when they look at sessions, don’t seem to find a huge uplift in making better compute allocation decisions or deciding on research directions. And so to me, that’s pretty similar to the PI case. And so I think there’s a first question, which is, how much of an improvement... Imagine you just didn’t get that much AI uplift on that component of AI research, but the rest of AI research really went crazy. Then how much faster do things go? And the separate question is, I don’t know, how hard is this strategic decision-making? Can’t you just have a bit longer horizon RL? Or maybe you can bet on decent transfer from other fields where those kind of decisions are important, and then you do actually get that kind of uplift at the end.
00:11:59 Nathan Lambert: I think part of my intuition is that science will look so fundamentally different that it’s almost hard to put a number on it. And it’s the pre and post-AI era, and we’re just in the rapid transition to what is a new method, new way of doing science, because I think all the conferences are ready to burn down and struggle through the next few years. I hope that they collectively figure out a way to like add AI oversight into reviewing and things that are scalable, because they have so many slop papers that they need this type of gate. So I don’t really know. And I think on the capability side, I’m like, what an AI progress is like clearly translated into new capabilities. There are two things. One is like pre-training scaling laws. Our loss is proportional to like an exponential increase in compute. And on the research side, what we are doing is we’re shifting the line so that it has a better offset and potentially a better slope. And then on the other side is the RL environments, and I think the RL environments we’re building now are very comparable to valuable work. So I expect the AI models to get much, much better at like knowledge work that can be scoped. But I don’t know if we have a good process for like churning out an order of magnitude harder environments, which would be closer to like cure cancer, solve these open math problems. I think math is a case that we could talk about. But that’s kind of like, I think there are unknowns on scaling the like raw intelligence more than efficiency. So I’m very optimistic in scaling efficiency.
00:13:29 JS Denain: Yeah. I agree that like in some sense, right, like inference efficiency is like a very like hill-climbing task. It’s like pretty well-scoped. And yeah, so I mean, it’s already something that has very, very fast trends, but I could imagine those trends. Yeah, I imagine those trends will go even faster. It seems really like the kind of... I mean, indeed, we have evidence from OpenAI, right? Like, saving on like serving costs, et cetera, through like building better kernels and like if you’re new to this. They don’t give that many details, but that’s already happening. One thing I’m curious about actually in your case is like, so here’s one way of defining like Anthropic, for example, like accelerates overall, which is you could look at the like ECI trend in like, the ECI of the best Anthropic quality point in time. And you can like look at the current trend line and you can ask the question, like over the next year, will we see a 5X acceleration? Like, will the slope be like 5X larger than it was, say, in like 2025? And I think it’s like pretty likely we see like a huge increase in this because it is like kind of a legible like KPI that like, I mean, it’s not literally KPI, but it is a KPI that like the company is aiming for, modulo like safety considerations, et cetera. And I’m curious about whether you think it’s very unlikely we get this or whether it’s more like we might get this, but like if we do get this, it’s mostly that like ECI has been Goodharted as a metric and like the implications for like real-world capabilities aren’t that huge.
00:15:03 Nathan Lambert: I wouldn’t be surprised if we got this, but I think that it’s going to be like we’re on a slope and then we could get an uptick in slope of hill climbing, but then we like saturate what we know how to hill climb and then it goes to be lower. So it’s like all the things that we could measure I think are going to be getting pulled up very quickly by being measurable. And then we’re in the domain of like, how do we measure it? Because I think, like I talk to people that are trying... Like evals are so expensive to build now, and I do think that evaluations are going to be like how good and efficient it is at coding, how good and efficient it is at ML research, how good and efficient it is at knowledge work. But I don’t know how to like... Building those evals all seems tractable but hard. But then like how to make a breakthrough in fundamental chemistry seems really, really, really hard to measure. I was going to draw on like maybe frontier math as an example, but I think math is such an exception as like one of the most jagged pieces of AI. I think especially like open problems in mathematics are like the perfect target for rapidly improving AI because it’s like a falsifiable thing. And it’s like if we were to, say, see that in something that’s much more open-ended, I think I would update a lot. Or if the labs were like to come out and say, “Using Claude, we have a very, very big change in what our architecture of AI is,” to like there’s the famous like Jonathan Frankle–Sasha Rush bet, and it’s like, and the transformer is no longer like the lineage we are on. I think any of those things being very AI-driven would make me update a lot. But seeing more math, like I think I was surprised by the pace of math, but like not astonished.
00:16:49 JS Denain: That’s interesting to me. I definitely agree with this general sense. So like I think METR folks looking at like your nanoGPT results from autoresearch-style things compared to like what the humans were doing, it does seem like there’s this, I think Tom Cunningham calls this like the apple-picking model where AI is like much more efficient at the start, but then doesn’t actually like uncover as many new ideas. And you see this in this kind of optimizer research. Yeah, I mean, one thing I will say on this like verifiability point is like, I think a pretty common trend is like you’ll have some task that’s like not verifiable and you’re like, maybe you struggle to build an environment for it. But actually it’s like it’s a subset of a larger task that is itself like verifiable. It’s just like longer range. An example of this is like there are many like hard to verify tasks out of like companies. But in some sense, like revenue or like other like metrics, like valuations are like pretty legible. So that’s like one thing. I mean, the other thing is like expect things to be pretty jagged. But I think a big question is like, yeah, can you get, for a crazy world, can you get like a large, like self-sustaining industrial kind of explosion?
00:18:05 Nathan Lambert: Yeah. Well, can we talk about robotics and industry? Because I have a background in physical robots and like I think the robotics trends will look much closer to self-driving cars than LLMs. And I think that a lot of the singularity arguments are based on robotics being able to look much closer to LLMs than the self-driving cars roll out. So, why would you disagree? Or, what is the argument that mass industrialization and robotic expansion is doable? Because my prior is so suspicious that I maybe even haven’t given it enough justice, but I’m very suspicious of this being a viability, and mostly in terms of being a relative timeline. I think it could happen over decades, but I don’t think it’s a two to five-year concern.
00:19:01 JS Denain: Two to five years seems rough, to be clear. I think I just don’t know as much about robotics here. I think is your main concern just reliability is really rough to get right in the same way that it was just a long tail of scenarios where things are, or was it more like a real-world thing where there’s much more regulation that comes up?
00:19:21 Nathan Lambert: I think it’s building things is hard. I think that, let’s see. I’ll talk us through some of this. For example, I know places like Amazon, they build new factories to be robotic first, and those are more effective for them. And what this would take then is building a robotics factory. In the case of the US, it’s like you have to build a robotics factory that builds robots very efficiently in the US and then transition that or make a new one that is built by said robots. And I think the re-industrialization of the US is something that I think is like, there’s a lot of reasons why it is not happening. I think potentially in China it is more likely, but I also just haven’t been convinced by AI results on visual and action models that they’re progressing fast enough. I think I’ve had discussions with people in the multimodal field have described the techniques as being much more rudimentary and less developed than the text language models, and in need of much more fundamental innovation, where something like code plus RL is a very natural match that the hill climbing is very predictable. So—
00:20:35 JS Denain: So, it seems like there’s two things. There’s the trends in robot capabilities is not as fast as you would expect for LLMs, and also even if robot capabilities were huge, it takes a while to build factories. I think I’m sort of skeptical of the second one. I’m just like, if robot capabilities are sufficient, the total addressable market for this is massive. And if you look at data centers in the US, there has been extremely fast build-out. If you had robots that were just literally able to substitute for blue-collar human workers, I feel like the financial incentives would be huge. And I think a lot of the reason why in some cases, the US doesn’t have huge build-out is just a demand thing. I think that’s the case for power, for example. So I think in that case, I’m just like, yeah, I feel like we just, what is the Tyler Cowen thing? Don’t underestimate the elasticity of supply is the main thing I would point to. I think on the capabilities front, I’m more uncertain. In particular, I’m sort of still confused and haven’t really looked into the, how much do you get directly actually from LLMs and foundation models for robotic capabilities? In particular, the other uncertainty I have is, it’s not clear to me that extremely fine-grained, extremely dexterous capabilities are the main thing you need for massive industrial explosions. And so, this longer tail of the hardest part of robotics, I’m not sure if that’s the biggest blocker for massive industrial explosion. Overall, robotics is something I have less expertise in. I’m interested in how many of the scenarios for doom ultimately kind of route through hard power acquired through robotics. I think part of my uncertainty also comes from, is it plausible to me that the minimum abilities that you need to acquire a lot of hard power and pose pretty catastrophic possibly extinction risks is more like, have access to nuclear codes or something like that? I don’t feel like I have great thoughts on this. I’m interested in more threat modeling, but I think that’s part of the thing is, what are the capabilities trends is something that people have disagreements about, and so what’s the minimum capability that’s necessary to cause these extinction-level harms, or harms that are sufficiently catastrophic, they just permanently alter the human trajectory?
00:22:46 Nathan Lambert: Yeah. The last point I would make on—
00:22:47 JS Denain: I think that’s the kind of questions I want to see a bit more thinking on, but yeah.
00:22:50 Nathan Lambert: The last point I would make on robotics is that I think the robots will be very good in constrained and repetitive environments, like manufacturing robots, and I think much longer until there are robots walking around the street cohabitating with humans. And if I were to go deep on this, I would want studies and discussions with people that are building multiple different data centers to how much variety is in their job versus how much of it is you take box off of truck and you put box in location. And I don’t have any good signal on how clearly repetitive that is now versus more human dexterous and ingenuity in problem-solving.
00:23:34 JS Denain: Yeah. Although also currently, the environments have been designed around humans who are pretty flexible in those ways and have other constraints. You could imagine designing factories to be robot first is a thing that I think you’re mentioning Amazon, right? Sort of does that more now. And so you could also... Yeah, I think that’s the kind of stuff that jaggedness would get you, which is you might just have massive accelerations, including in the physical world of some industries, even if capabilities for fully being as dexterous or flexible as a human might not be there. Mostly, yeah. I think mostly, yeah, robot capabilities seem very important to track to me. We had some piece about this at Epoch, but we don’t claim robotics expertise.
00:24:21 Nathan Lambert: Yeah. I agree. I think we could shift to another capabilities topic, which is how far behind do you think the top Chinese labs are of OpenAI and Anthropic? And you can define how you want to measure it, whether it’s like public models or like internal models. I think doing both is actually pretty interesting to think about. Like, how would you describe the gap? Like, choose your... This is one of the few things I’ll push you for a number on.
00:24:47 JS Denain: Yeah. I think I would go with like, I don’t know, six to eight months or something, roughly.
00:24:53 Nathan Lambert: From public to public?
00:24:56 JS Denain: Yeah. Like, release date to release dates. I don’t have a great catch numbers on like how long the internal to public deployment. Yeah, I’m sorry if it increases a bit if you’re counting when the model was built, because I would expect the delay to be larger for OpenAI and Anthropic than it is for Chinese labs. But yeah, I don’t know, just looking at ECI and there’s a few different methods you can do there, and there’s a good amount of noise between. Those methods are kind of reasonable, and my sense is they give you something on the order of like six to eight months-ish.
00:25:28 Nathan Lambert: I would say that like Kimi K3 and GLM 5.2 are closer to like two to three or four. So do you think those are anomalies down to measurement or overfitting? So I have this discussion a lot with Florian. And Florian, who helps me with Interconnects, is constantly badgering me down. And I think I intuitively have landed something closer to you. And I think those two models, and in particular Artificial Analysis, were very close. I think Kimi might have been even under two in the Artificial Analysis Index at the time. And I’m just putting this out there, and I go back and forth all the time on it.
00:26:12 JS Denain: So I don’t remember the especially most up-to-date details on the Artificial Analysis Index. My sense is that it might understate the gap for curation reasons or something, but I don’t want to be too confident there because also they update methodology and stuff like that. I think even in ECI, though, there are some cases where it was more like four months. Yeah, I don’t know. Four to eight months seems reasonable to me. I think like two months, I would be like, “Nah, that seems like a bit more overfitting-y.” We did internally look into a bit how much overfitting explains the gap. Like, just doing the kind of like, what’s the ECI lag analysis for, if you just take an ECI based on private benchmarks versus not, or private benchmarks plus results were from models that were from benchmarks that were published after a model was released or something like that, where there’s no risk of contamination. I think overall, that effect was basically not statistically significant, which was interesting. I think if you did rely just results in model cards as opposed to the results being put in ECI, we probably get some contamination or overfitting effects. I still think there’s some amount of, yeah, something like ECI might underestimate the gap just because I do think OpenAI and Anthropic are probably like... There’s probably other capabilities just would appear less in benchmarks and where there’s been more optimization or serving a broader set of users. But yeah, I don’t know. That’s my overall number.
00:27:39 Nathan Lambert: If we were to ban distillation effectively, not even just ban, but if distillation were effectively be stopped, where do you think the number would be in six to 12 months? I think without RSI being super crazy. I think RSI going super crazy, the labs pull ahead by a lot more.
00:28:01 JS Denain: Yeah. On current trends, I think I’ve actually updated towards distillation is actually a really big factor. We can talk about the different factors. And so, okay, so currently if I’m saying six months gap, then in six months, it’s probably not literally 12 months, right? So, my guess is, yeah, you are at like eight or nine, something like that. You go from like six to nine maybe.
00:28:31 Nathan Lambert: Why have you updated on distillation being effective?
00:28:35 JS Denain: Yeah. So I think it’s like, yeah, there’s two reasons. I think, a lame reason is, I don’t know, more people seem to be saying it’s this.
00:28:45 Nathan Lambert: They have been. I’m surprised by it.
00:28:47 JS Denain: Yeah. The Stolen Thoughts paper did suggest it was actually relatively easy to get the stuff. Yeah, random gossip. I think also a big one is, we can talk about this, thinking about the other explanations for the lag and them not seeming as compelling as when I first thought of them. I do think the Claude routers thing also just seems like a pretty big deal. And initially when I thought of distillation, I hadn’t considered that. But I do think Claude routers giving the kind of right prompt distribution for realistic usages, use cases and stuff like that just seems quite useful.
00:29:26 Nathan Lambert: But it’s also like these companies have their own usage at this point. And it’s also like we have the benchmarks. So the benchmark distribution we already have. So I don’t necessarily think distillation is helping with... I think the routers could be very helpful, but I was confused by that prompt distribution thing because once you have the benchmark distribution, you just make similar examples or find similar examples.
00:29:50 JS Denain: Yeah, so I agree. So I think the Claude routers just give you a bunch of useful data to train on, and that will generally improve capabilities. I think that’s one explanation. I agree that I think giving you the right prompt distribution is, I think, going to be useful for real-world capabilities. I agree that purely if I’m discussing this ECI lag or something, then yeah, you can just use the prompt distribution for benchmarks. And so I think it’s not going to be that good of an explanation there. Yeah. I think just... Yeah. Also, so as an example, right, people saying that mid-training is a really important thing or something. I think actually in that conversation, right, Beren was like, “Mid-training is actually getting you 80% of the way there.” And more claims of this form and generally of RL being very useful but not doing that much of the exploration work, I think is also evidence that doing SFT on language trajectories is really useful.
00:30:48 Nathan Lambert: Yeah. I want to go through—
00:30:50 JS Denain: I think he mentioned on a post on this is, it’s a good initialization, but I guess I’m updated that this initialization is really, really important. And yeah.
00:30:59 Nathan Lambert: Yeah, because I’m still on the side that I think doing RL well and doing RL faster so you can go bigger is the way people are getting capabilities. But I have been hearing more about this repeated... The framework would be it’s a very repeated cycle between mid-training SFT and RL, and then your peak RL helps you feed into mid-training and less of a very big RL run, which I never thought... The RL runs are a week probably. I don’t think they’re insane, especially like Zhipu or Kimi. And that’s... So I could see this, and I was like, “I don’t really know.” I generally thought that because the models are getting so useful, I don’t know how... It just seems like the scale of distillation would need to be so big to really do that for mid-training. And the counterargument is Kimi and GLM 5.3 are also strong models, where it’s like, if the gap is that small, it just doesn’t make sense to me that the help would be so big. That’s why I’m still on the not convinced, but looking for more evidence. And then some of the other re... I have this blog post that I wrote in front of me. If you want to talk about distillation more, you can get that comment in before I switch to other topics.
00:32:21 JS Denain: I guess I’m pretty interested in talking about the alternative explanations or something for... I think there’s this general phenomenon of Chinese labs have way less capital and particularly way less compute than US labs. And that delay is huge. And then this, whether it’s four months or eight months, it’s a much shorter delay in capabilities than in capital. And so, this requires some explanation.
00:32:48 Nathan Lambert: Yeah. Do we think the Chinese labs are still releasing meaningfully faster from time that RL is done to public gets the API and evaluation scores? In the past, I thought that the time to release was much faster, and I think the labs are releasing intermediate versions faster now, so I don’t know if that explains as much of it. But I would say a year ago, I would put a lot more to Chinese models get model out within days, American labs could take months, and that would artificially squeeze the gap very substantially.
00:33:26 JS Denain: Why would it squeeze the gap that much?
00:33:28 Nathan Lambert: Because you have the trajectory of capabilities, and higher one is OpenAI and Anthropic, and they stop training here in time, and then the Chinese labs are here. But then it’s just like the model is stagnant before it gets released.
00:33:44 JS Denain: Yeah, I think this depends on the method that... This both depends on the method that you use for computing the lag or something, right? So my sense, I think when I was saying four to eight months, I think a lot of that uncertainty interval also comes from, are you measuring the method by looking forward or backward or something? Or how are you resolving the uncertainty? Are you taking this kind of staircase or this kind of staircase, or interpolating or something like that? So I think that’s kind of still accounted for in the uncertainty. So my guess is that’s not a huge... Or that’s still, even with this, I think there’s still a large effect to explain. And I think, yeah, these are other potential explanations, right? So one is maybe there’s a big delay in compute, but in labor or potentially data, there’s less of a delay, and those are important factors. And so even if you’re two years behind in compute, if you’re six months behind in data and in fact ahead in number of researchers or something, or not that far behind in number of researchers, maybe that helps catch up, and that’s not even a spillover effect, right? That’s just a reason. Yeah, there’s distillation. Other explanations people have are just ideas might leak, and I don’t suspect that’s a huge effect. People can play with a model, and through playing with the model, they kind of infer maybe what it was trained on or what’s useful to push on is another explanation. Those things are not quite like distillation, but using Claude as your reward model or using Claude to clean data or something, maybe have some spillovers.
00:35:23 Nathan Lambert: Let’s see. I’m trying to quantify some. I would say that I think the Chinese labs care about benchmarks a bit more for financial... Some of them are public companies, and showing close benchmarks helps them a lot. So I do think they care about it more. I don’t know if that’s going to give you a month or two. I don’t know what are you going to give back, like a month or two. I think there’s a good chance that the Chinese labs are better at organizing talent on just doing really mundane hill-climbing data work. I don’t think that’s a huge effect, but it’s just like they have a ton of talent, and culturally through the way that... I think that there’s a good chance that it’s just more grind. They might actually grind more effectively than the American labs, which is... I don’t know how. I wouldn’t put a lot to this because I think it’s fairly close, but I think that’s a chance. But that’s a very marginal... That’s not a gigantic lead cause. I would put more of it to compute difference than to distillation. But maybe distillation is very related as a compute difference because it’s a way to turn... It is a definition of way to turn money into very high-quality data, which is something you would otherwise need compute for.
00:36:35 JS Denain: Yeah. Another theory that I’ve vetted about, and I’m not sure how big of a deal it is, is there’s this data market in the US, and maybe one thing that happens is the best RL environments or evals, it kind of takes a lot of compute actually to figure out which ones they are. And so maybe there’s kind of a collaboration between data providers and AI labs to actually test and figure out what RL environment really worked. And then maybe that data provider, once they’ve done this iteration, which implied a lot of R&D compute spending from a frontier lab, then they’ll build a bunch of similar RL environments because they’ve learned the lesson, and then they’ll sell them with some delay, but still sell them to other people, including in China.
00:37:21 Nathan Lambert: I’ve heard this from people.
00:37:22 JS Denain: And so there’s kind of this implicit R&D compute that was spent and that it gets saved. I’m not sure how big of a deal that is.
00:37:29 Nathan Lambert: I’ve heard this multiple times, including from people in China that have... This is like multi-hop type of rumor that I’ve heard a few times is like, we just have to wait a certain amount of time and then we buy the RL environments that Anthropic bought for a 10th of the price. And I think this could contribute a lot to capital efficiency, but also as somebody who buys into human factors being very important in model progress, which is just like competition, I think the analog to the competition that OpenAI and Anthropic feel right now is proof of concept as a way to make something way easier. And I do think the OpenAIs and Anthropics of the world are much more likely to innovate in open-ended things like Navier-Stokes, multi-agent mega-scaling than the Chinese labs. And having talked to many of the labs, I don’t think they would contest this, but they probably won’t put it on the record. But they’re just like, trying to keep up is so much easier.
00:38:27 JS Denain: Yeah. So that’s more like the... Yeah, so the way I describe this was like four-minute mile style effects of like, you show that something is possible at all and then people have the conviction to go down that route, don’t waste a ton of resources exploring a bunch of other things, is the thing you’re pointing at here.
00:38:43 Nathan Lambert: Yeah. And I think this—
00:38:44 JS Denain: I guess I’m looking for examples where you think this... What are examples? Because for example, reasoning models is not this, right? Like, DeepSeek Math or whatever came way before o1. Yeah, I’m kind of curious if... Yeah, I agree this seems kind of plausible, and I’m not sure I can come up with examples of big innovations that a frontier lab has—
00:39:04 Nathan