跳到主要內容
AI News HubLIVE
來源內容 · 翻譯待補全2 分鐘閱讀

待翻譯:RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned…

待翻譯:RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

content type paperpublished October 2026 RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback AuthorsMichael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi, Abbas Kazerouni, Omar Attia, Sanjoy Chowdhury, Alexander Toshev View publication The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task → insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128 = 0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14–31% with insights in context during training and, crucially, 12–13% when no insight is in context at eval time. We identify that the key is the task → insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form “on this sort of task, keep this sort of thing in mind”, which we hope to inspire future research on. DR-MPC: Deep Residual Model Predictive Control for Real-World Social Navigation March 6, 2025research area Computer Vision, research area Methods and Algorithmsconference IEEE Robotics and Automation Letters How can a robot safely navigate around people with complex motion patterns? Deep Reinforcement Learning (DRL) in simulation holds some promise, but much prior work relies on simulators that fail to capture the nuances of real human motion. Thus, we propose Deep Residual Model Predictive Control (DR-MPC) to enable robots to quickly and safely perform DRL from real-world crowd navigation data. By blending MPC with model-free DRL, DR-MPC overcomes… Read more Apple Workshop on Machine Learning for Health 2023 November 8, 2023research area Health Earlier this year, Apple hosted the Workshop on Machine Learning for Health. This two-day hybrid event brought together Apple and the academic research community and clinicians to discuss state-of-the-art machine learning (ML) research in health. Read more

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This beco…

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。