跳到主要內容
AI News HubLIVE
站內改寫1 分鐘閱讀

待翻譯:The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Inside the RSI post-training pipeline

來源TheSequence作者: Jesus Rodriguez
待翻譯:The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Forget the agent that rewrites its own code. The most economically important self-improvement loop in AI is the post-training pipeline, and every frontier lab has been running it at industrial scale for two years. Last time I argued that the factory got automated before the design office, and that the line between them tracks whether the work comes with an answer key. This time I want to look inside the part of the factory that matters most, because there is a self-improvement loop running in there that gets almost no attention in the RSI conversation, even though it is the one actually producing the models. Here is the loop. A model writes many candidate answers to a problem. Something grades them. The good ones become training data. The model trains on them and gets slightly better at producing good ones. Repeat. That is it. It is called STaR in the 2022 paper that first stated it cleanly, RLVR in the 2025 vocabulary, and post-training in the org chart. It is the same loop, and it is the loop that took frontier models from chatbots to agents. Sourdough, not a screwdriver Read more

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Inside the RSI post-training pipeline

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。