跳到主要内容
AI News HubLIVE
站内改写4 分钟阅读

待翻译:Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google Cloud AI Research has open-sourced RRSI, a framework that lets LLM agents rewrite their own prompts, tools and memory while model weights stay frozen. It adds a leakage critic, a noise floor, a cost rule and pruning so gains carry over to new tasks. With Claude Opus 4.8, Terminal-Bench 2.1 rose from 74.2% to 80.2%, and all 6 held-out splits improved. The post Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting appeared first on MarkTechPost.

来源MarkTechPost作者: Asif Razzaq
待翻译:Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Google Cloud AI Research, with UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement). It lets an LLM agent rewrite its own harness: prompts, tools, memory, control flow and sub-agents. Model weights never change. RRSI constrains the improvement loop itself, so gains hold on benchmarks the agent never optimized against. Deployable? Yes, as a research framework. The code is Apache 2.0, needs Python 3.10+, and accepts any LiteLLM model string. Defaults assume Claude Opus 4.8 on Vertex AI. Why Self-Improving Harnesses Overfit Harness evolution loops propose edits, score them on a fixed evolve set and keep the winner. The same tasks are reused every round, so the loop can memorize them. The RRSI research names 3 failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer. How RRSI Works RRSI keeps every harness component editable. It regularizes how the search moves instead. Proposal side Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change. Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried. Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched. Selection side Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring. Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness. Cost rule: extra inference tokens must be paid for by measured gain. Pruning: components that stop producing gains become deletion targets. The research team frame these as analogies to classic regularizers. The edit budget maps to L0, pruning to Lasso (L1) and the cost rule to Ridge (L2). Results Across 8 Benchmarks Terminal-Bench 2.1 (evolve split): 74.2% to 80.2%. SWE-bench Verified (never used for selection): 82.0% to 83.8%. Out of distribution: JobBench +4.7, GDPval +3.5 and APEX-Agents +3.7 points. EngDesign (evolve) +4.9; Frontier-Eng +4.3 Medal points. Harvey LAB: +1.1 on the evolve split, +2.3 on its held-out split. All 6 held-out splits improved. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. SWE-bench Verified rose from 76.8 to 79.0. The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial. Unregularized evolution uses 3.80M. The abstract reports this as 30% fewer; the project page says 36%. RRSI vs Closest Competitors Scores come from Table 1 of the RRSI research paper. All methods share the same starting harness, policy, evolve split and candidate budget. FeatureRRSIMeta-HarnessAHETTHEHarnessX Core ideaRegularized proposal and selectionAgentic proposer over code, scores and traces of all prior candidatesObservability-driven loop; edits paired with verified predictionsEvolves harness during test time, no gold labelsModular typed primitives, trace-driven adaptation Model weightsFrozenFrozenFrozenFrozenFrozen Cost rule and pruningYesNo*No*No*No* Harvey LAB evolve score90.593.090.791.191.8 OOD average (H0 = 39.7)43.640.639.238.039.7 *Per the RRSI research team. OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1. Meta-Harness leads on the Harvey LAB evolve split. RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0. Interactive Explainer How RRSI Regularizes Agent Self-Improvement The model stays frozen. The harness (prompts, tools, memory, control flow) evolves, but every edit must pass the regularizers. Pick a candidate edit, then press Run. Watch where RRSI stops it. ProposerReads full edit ledger, shrinking edit budget Leakage criticScreens diff before scoring EvaluateRun on evolve set GateNoise floor + cost rule Harness Ht+1Accepted, then pruning check Candidates are illustrative. The rules they hit are the ones described in the RRSI paper. // edit ledger: component | hypothesis | Δscore | Δcost | verdict bt = ⌈ bmin + (bmax − bmin) · ½(1 + cos(πt / T)) ⌉ bmax (edits per candidate, early) 5 bmin (edits per candidate, late) 1 T (rounds) 16 Early rounds may bundle coordinated edits to find a mechanism. Late rounds get single, attributable changes. round 0round T−1 Cosine schedule from the paper (Eq. 4). Slider values are for exploration; the paper lists its own settings in Appendix D. ΔS, score gain vs incumbent (pts) 2.0 ΔC, token cost change (%) 20 δ, noise band from repeated base runs (pts) 1.5 β0 + β1ΔS cost allowance: β1 (% per pt) 8 ?Noise-adjusted floor: score must not drop more than δ below the best so far ?Clear gain: ΔS > δ, otherwise the paper’s within-band rule applies ?Cost rule: ΔC ≤ β0 + β1ΔS … Illustrative thresholds (β0 fixed at 5% here). In RRSI, δ is estimated on the unchanged base harness and β values are tuned on the evolve set, then frozen. Data: RRSI paper · GitHub · project pageBuilt by Marktechpost Getting Started Copy CodeCopiedUse a different Browser git clone https://github.com/google-research/rrsi.git && cd rrsi pip install -e ".[dev]" python3 rrsi.py --domain coding baseline python3 rrsi.py --domain coding run Each round drafts 2 candidates in separate git worktrees, screens them, evaluates both and fast-forwards the branch to the winner. The coding instance also needs Docker and harbor. New domains plug in through a single adapter module. Key Takeaways RRSI evolves prompts, tools, memory and workflows while model weights stay frozen. A leakage critic, noise floor, cost rule and pruning decide which edits stick. Terminal-Bench 2.1 rose from 74.2% to 80.2% with Claude Opus 4.8. SWE-bench Verified, never used for selection, rose from 82.0% to 83.8%. Apache 2.0 code on GitHub; research-grade, not an official Google product. Check out the Paper, Codes and Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting appeared first on MarkTechPost.

展开要点与分析

文章情报

工程师进阶

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Google Cloud AI Research has open-sourced RRSI, a framework that lets LLM agents rewrite their own prompts, tools and memory while model weights stay frozen. It adds a leakage cri…

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。