本文にスキップ
AI News HubLIVE
サイト内リライト1 分で読了

翻訳待ち:The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Inside the RSI post-training pipeline

ソースTheSequence著者: Jesus Rodriguez
翻訳待ち:The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Forget the agent that rewrites its own code. The most economically important self-improvement loop in AI is the post-training pipeline, and every frontier lab has been running it at industrial scale for two years. Last time I argued that the factory got automated before the design office, and that the line between them tracks whether the work comes with an answer key. This time I want to look inside the part of the factory that matters most, because there is a self-improvement loop running in there that gets almost no attention in the RSI conversation, even though it is the one actually producing the models. Here is the loop. A model writes many candidate answers to a problem. Something grades them. The good ones become training data. The model trains on them and gets slightly better at producing good ones. Repeat. That is it. It is called STaR in the 2022 paper that first stated it cleanly, RLVR in the 2025 vocabulary, and post-training in the org chart. It is the same loop, and it is the loop that took frontier models from chatbots to agents. Sourdough, not a screwdriver Read more

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • Inside the RSI post-training pipeline

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。