翻訳待ち:The Sequence Knowledge #920: The Physics of Teaching: Distillation Scaling Laws
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:The Sequence — Distillation Series
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
Every field becomes a science at the moment it stops collecting anecdotes and starts fitting curves. For most of its history, distillation was an anecdote field. It worked, often spectacularly, and nobody could tell you in advance by how much. Should the teacher be as strong as possible? Folk wisdom said yes; practitioners kept tripping over cases where a stronger teacher produced a worse student. How much data does distillation need? Depends who you asked. Was distilling ever actually cheaper than just training the small model longer? Shrug. The field ran on vibes and ablations, which is a fine way to write papers and a terrifying way to spend ten million dollars on a training run. Meanwhile, right next door, pretraining had undergone exactly the transformation distillation lacked. The Kaplan scaling laws, then Chinchilla, turned “how big a model should I train, on how much data?” from a matter of taste into a matter of arithmetic. Loss became a predictable function of parameters and tokens. Budgets became optimization problems. The single most consequential number in the industry — twenty-ish tokens per parameter — fell out of a fitted curve. The obvious question hung there for three years: where is the Chinchilla of distillation? If a student’s loss is a function of its size and its data, it must also be a function of its teacher. What does that function look like? In early 2025, a team at Apple led by Dan Busbridge answered it, with the most compute-intensive controlled study of distillation ever run — students from 143 million to 12.6 billion parameters, teachers spanning a similar range, up to 512 billion training tokens. The resulting paper, Distillation Scaling Laws, is the closest thing the field now has to physics. This essay is about what the curve says, and what it quietly settles. The Shape of the Law Read more