跳到主要内容
AI News HubLIVE
站内改写1 分钟阅读

待翻译:The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Inside the RSI post-training pipeline

来源TheSequence作者: Jesus Rodriguez
待翻译:The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Forget the agent that rewrites its own code. The most economically important self-improvement loop in AI is the post-training pipeline, and every frontier lab has been running it at industrial scale for two years. Last time I argued that the factory got automated before the design office, and that the line between them tracks whether the work comes with an answer key. This time I want to look inside the part of the factory that matters most, because there is a self-improvement loop running in there that gets almost no attention in the RSI conversation, even though it is the one actually producing the models. Here is the loop. A model writes many candidate answers to a problem. Something grades them. The good ones become training data. The model trains on them and gets slightly better at producing good ones. Repeat. That is it. It is called STaR in the 2022 paper that first stated it cleanly, RLVR in the 2025 vocabulary, and post-training in the org chart. It is the same loop, and it is the loop that took frontier models from chatbots to agents. Sourdough, not a screwdriver Read more

展开要点与分析

文章情报

工程师进阶

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Inside the RSI post-training pipeline

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。