AI News HubLIVE
In-site rewrite6 min read

Inside the Model Factory — Eiso Kant, Poolside AI

Poolside's co-CEO on how his small team of top researchers built a model factory capable of training Laguna S - a 118B MOE beating Thinky's ~1T open weights model... and this is just the beginning.

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines’ recent release nearly 10 times their size.

Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club:

@latentspacepod breakdown of our Laguna M.1/XS.2 Technical Report! The Latent Space paper club just did a deep dive, and their takeaways perfectly capture what we set out to build with our Model Factory. A few quotes from the video 🧵👇 (1/6)\nyoutu.be/QLfZamyMls0","username":"eisokant","name":"Eiso Kant","profile_image_url":"https://pbs.substack.com/profile_images/1842230143965675520/j6mVG2Py_normal.jpg","date":"2026-05-28T20:34:07.000Z","photos":[],"quoted_tweet":{},"reply_count":1,"retweet_count":8,"like_count":47,"impression_count":11267,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}" data-component-name="Twitter2ToDOM">

From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.

We go deep on Poolside’s Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.

We also discuss model-harness co-design, Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside’s $500 million raise, open-source AI, regulation, NVIDIA and TSMC’s influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.

We discuss:

How Andrej Karpathy’s RNN work inspired Eiso to start building language models for code in 2015

Why Eiso spent four years and $12 million pursuing an idea before the market cared

Why ChatGPT felt like vindication and brought Poolside back to open source

Why Eiso would prefer 100 foundation model companies over an oligopoly of five

The difference between releasing open weights and publishing genuinely open research

Why Poolside deliberately built a global research organization outside the Bay Area talent war

Why model building is ultimately 90% engineering

The Model Factory: Poolside’s end-to-end system for rapidly training and improving models

How fewer than 70 researchers run roughly 10,000–20,000 experiments each month

How Poolside moved from six-month model cycles to five- and eight-week launches

Why streaming data directly into training unlocked faster experimentation

How immutable data, versioned code, and reproducibility enable rigorous model research

Why Eiso wants capable researchers to leave their labs and become Poolside’s competitors

Why 95% of model building can be reduced to better data or compute efficiency

Laguna S and why persistence, verification, and backtracking can outperform raw intelligence

Why smaller models may handle far more knowledge work than previously expected

Why reinforcement learning will move earlier into pre-training

Why next-token prediction is still failing to extract enough knowledge from the web

Why distillation and environments have become the AI industry’s favorite “drugs”

Why mid-training is really an early form of curriculum design

Low-precision training, networking bottlenecks, and the next gains in compute efficiency

Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch

Why model builders can often evaluate a new checkpoint within its first 30 minutes

Model versus harness: where agent capabilities actually come from

Why Poolside sees coding and long-horizon software tasks as a path to AGI

Why Eiso thinks MCP and traditional tool calls are “stupid”

Why future agents will write scripts instead of choosing from dozens of predefined tools

The case for minimal harnesses, containers, and model freedom

Why Poolside is prioritizing vision but does not expect to work on audio soon

Why language may be the most compute-efficient modality for encoding knowledge and reasoning

The real cost of model development and why the final training run is anticlimactic

The story behind the Poolside name and why it represents refusing to lower ambitions

How Poolside raised $500 million while investors still questioned whether AGI was real

Why intelligence could become the world’s most demanded and commoditized resource

When open models may become too capable to release without restrictions

Why unilateral AI safety does not work in a globally competitive environment

How regulation could accidentally lock in an oligopoly of two or three AI companies

NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress

Why reinforcement-learning wall-clock time is one of Poolside’s biggest bottlenecks

Why Poolside trains models from scratch instead of simply distilling larger models

How AI changes the way companies should measure engineering productivity

Why agency may become the most important quality for employees in the AI era

How leaders align high-agency people through shared goals and clear constraints

Hiring across research, post-training, pre-training, architecture, evals, and engineering at Poolside

Eiso Kant

LinkedIn: https://www.linkedin.com/in/eisokant

X: https://x.com/eisokant

Poolside: https://poolside.ai

Timestamps

00:00:00 Introduction

00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers

00:02:26 The $12M Failure and ChatGPT Vindication

00:03:39 Open Source and the Case for 100 Foundation Model Companies

00:09:22 Open Weights, Open Research, and Poolside’s Global Team

00:16:04 The Model Factory: Why Model Building Is 90% Engineering

00:20:19 Agents, Automated Experiments, and Early Signs of RSI

00:24:04 Streaming Data, Reproducibility, and Scientific Rigor

00:30:35 Creating More Foundation Model Companies

00:36:07 Laguna S: Persistence vs. Raw Intelligence

00:43:01 Reinventing Pre-Training, RL, and Curriculum Design

00:52:33 Low-Precision Training and Squeezing More From Smaller Models

00:58:37 Model Harnesses, Coding Agents, and the Path to AGI

01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”

01:13:04 Vision, Multimodality, and Why Language Still Matters

01:18:15 Scaling Models and the Real Economics of Training

01:20:40 Why Poolside Is Called Poolside and Raising $500M

01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly

01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck

01:41:52 Smaller Models, Distillation, Engineering Productivity, and Hiring

Transcript

Introduction: Eiso Kant, Poolside, and Open Models

Swyx [00:00:00]: All right, we’re here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.

Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.

Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I’m on my way to SF.” I was like, “You’re on a plane right now, right?” Like, hey.

Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let’s do it.

Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don’t have to prepare that much because if you’re truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don’t live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we’re gonna talk about. But I guess, like, what got you into democratization of AI? Like, it’s not obvious from your LinkedIn or something.

From Karpathy’s RNN Post to Sourced

Eiso Kant [00:00:57]: No, it’s not at all. I don’t think it’s obvious how I got in this space. I owe getting into this space to Andrej Karpathy.

Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”

Swyx [00:01:10]: Neural Nets, yep.

Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There’s a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn’t. and there’s a little-- There’s an example of code a little bit further down. Yeah, so Shakespeare.

Swyx [00:01:47]: Shakespeare.

Swyx [00:01:49]: Cool

Eiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.

Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn’t obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn’t obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors’ money, which was a lot back then.

Swyx [00:03:18]: Yep.

Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn’t really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was

[truncated for AI cost control]