跳到主要內容
AI News HubLIVE
來源內容 · 翻譯待補全6 分鐘閱讀

待翻譯:Google Research RRSI Guide: Mastering Self-Improving AI Agents

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Explore a comprehensive coding guide to Google Research's RRSI (Regularized Recursive Self-Improvement), detailing how noise bands, cost rules, and leakage screens enable safe, efficient, and self-improving AI agents. The post Google Research RRSI Guide: Mastering Self-Improving AI Agents appeared first on MarkTechPost.

來源MarkTechPost作者: Sana Hassan
待翻譯:Google Research RRSI Guide: Mastering Self-Improving AI Agents
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores them inside Docker benchmarks, which is not something a free notebook can run. The part of RRSI that actually carries the paper’s idea, the rules that decide which proposed edits to keep, is plain Python, and that is what we drive directly. We install the package from the official repository, walk through its estimator, its calibrated noise band, both branches of its selection algorithm, its annealed edit budget, its deterministic leakage screen, and its edit history, and then plug a simulated agent into RRSI’s own Domain interface. Because we built the simulated environment ourselves, we know the true effect of every edit, which lets us audit RRSI’s decisions against ground truth and compare them with an unregularized search that simply keeps whatever scores highest. Copy CodeCopiedUse a different Browser import os import sys import json import math import copy import random import tempfile import textwrap import traceback import subprocess import statistics as st from pathlib import Path RESULTS = {} def banner(title): print("\n" + "=" * 78) print(title) print("=" * 78) def section(name): def wrap(fn): def run(*a, kw): banner(name) try: out = fn(*a, kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).name}: {e}" print(f"\n[!] {name} did not complete: {type(e).name}: {e}") traceback.print_exc(limit=3) return None return run return wrap banner("1. Install RRSI and map the paper onto the code") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "git+https://github.com/google-research/rrsi.git@be50316e1db05914068a973f322770ef08ed7ba1"], check=True) from importlib.metadata import version from rrsi.config import RRSIConfig from rrsi.evaluate import TaskResult, EvalResult, aggregate, evaluate from rrsi.calibrate import calibrate from rrsi.selection import Candidate, cost_rule, judge, select_round from rrsi.schedule import edit_budget, budget_table from rrsi.history import History, stall_flag, exploration from rrsi.components import K, K_STR, normalize, novelty from rrsi.critic import precheck, review from rrsi.domain import Domain print(f" rrsi {version('rrsi')} | anthropic {version('anthropic')} | Python {sys.version.split()[0]}") print("\n RRSI evolves an agent's HARNESS (prompts, tools, memory, control flow, sub-agents) around a") print(" frozen model. The full loop drafts edits with Claude Opus on Vertex AI and scores them in Docker") print(" benchmarks. The part that decides which edits to KEEP is plain Python, and that is what we drive:") for symbol, where in [ ("S_hat, C_hat Eq. (estimate)", "rrsi.evaluate.aggregate"), ("delta noise band", "rrsi.calibrate.calibrate"), ("Algorithm 2 floor + cost rule", "rrsi.selection.judge / select_round"), ("b_t Eq. (anneal)", "rrsi.schedule.edit_budget"), ("Critic leakage screen", "rrsi.critic.precheck / review"), ("L_t, g_t, B_t history, yield, prune", "rrsi.history.History"), ("sigma_t, U_t stall + exploration", "rrsi.history.stall_flag / exploration"), ("nu structural novelty", "rrsi.components.novelty"), ]: print(f" {symbol:40s} -> {where}") CFG = RRSIConfig() print(f"\n paper defaults: T={CFG.T} rounds, k={CFG.k} trials/task, m={CFG.m} candidates/round," f" b in [{CFG.b_min},{CFG.b_max}]") print(f" beta0={CFG.beta0} beta1={CFG.beta1} w_s={CFG.w_s} w_c={CFG.w_c} w_n={CFG.w_n}" f" delta_z={CFG.delta_z}") print("\n Nothing below needs an API key, a GPU or a dataset download.") We install RRSI from the google-research repository, pinned to the commit this notebook was written against, since the package is not on PyPI. Its only dependency is the Anthropic client, which the search roles use to call Claude and which we never exercise. We then print the mapping the repository itself documents between the paper’s symbols and the functions that implement them: the empirical score and cost estimate in evaluate, the noise band in calibrate, Algorithm 2 in selection, the annealed edit budget in schedule, the leakage screen in critic, and the edit history with its yield, prune, stall and exploration summaries in history. RRSIConfig holds the paper’s hyperparameters, and every function below receives it exactly as the real loop does. Copy CodeCopiedUse a different Browser @section("2. Evaluate(H): a score and a cost, and why a crash counts as zero") def estimator(): base = { "task_000": TaskResult(rewards=[1, 1], tokens=[11_800, 12_400]), "task_001": TaskResult(rewards=[1, 0], tokens=[15_100, 14_600]), "task_002": TaskResult(rewards=[0, 0], tokens=[21_000, 19_500]), } ev = aggregate("H0", 2, base) print(f" three tasks x k=2 trials -> S_hat = {ev.S:.3f} C_hat = {ev.C:,.0f} tokens/trial" f" ({ev.n_expected} trials expected, {ev.missing} missing)") crashy = dict(base) crashy["task_002"] = TaskResult(rewards=[0.0, 0.0], tokens=[None, None], missing=2) ev_crash = aggregate("crashy", 2, crashy) dropped = {t: r for t, r in crashy.items() if not r.missing} naive = sum(sum(r.rewards) for r in dropped.values()) / sum(len(r.rewards) for r in dropped.values()) print("\n A candidate crashes on the hardest task instead of failing it:") print(f" an estimator that drops missing trials reports {naive:.3f} 12s} {'trials':>7s} {'delta':>8s} method") deltas = {} for n, k in [(40, 2), (100, 4), (400, 8)]: cal, _ = calibrated_delta(make_world(0, n), k, seed=0) deltas[f"{n}x{k}"] = cal["delta"] print(f" {f'{n} x k={k}':>12s} {n * k:>7d} {cal['delta']:>8.4f} {cal['method']}") print("\n The paper's calibrated bands are 0.017 (coding), 0.004 (workspace) and 0.020 (engineering).") print(" A gain smaller than delta is indistinguishable from re-running the same harness, and") print(" Algorithm 2 treats it that way. Remember the first row: it becomes the lesson of step 11.") return "delta " + ", ".join(f"{k}={v:.3f}" for k, v in deltas.items()) noise_band() Before any rule can separate a real gain from luck, it needs to know how far one harness’s score moves on its own. We build a small simulated agent, whose success on each task is a logistic function of harness skill minus task difficulty, and evaluate the unchanged starting harness six times on forty tasks with two trials each: the scores spread by 0.113 although nothing changed. calibrate turns repeated evaluations of the same harness into delta, twice the standard deviation of the difference between two runs. With 80 trials delta is about 0.108; with 3,200 trials it falls to about 0.013, in the range the paper reports for its instances (0.004 to 0.020). For selection purposes, any gain smaller than delta is indistinguishable from re-running the same harness. Copy CodeCopiedUse a different Browser def ev_at(S, C, job, n=200): """An EvalResult with exactly score S (in steps of 1/n) and C tokens per trial.""" hits = round(S * n) return aggregate(job, 1, {f"task_{i:03d}": TaskResult(rewards=[1.0 if i 6,} {'ADMIT ' if dec.admissible else 'reject'}") print(textwrap.indent(textwrap.fill(dec.reason, 88), " " * 6)) print("\n Three rules, in the order RRSI applies them:") print(" 1. never fall below the best score ever seen, minus the noise band (A)") print(" 2. a gain bigger than delta must pay for any extra tokens: dC 2d}" for t in range(CFG.T))) print(" b_t : " + " ".join(f"{b:>2d}" for b in table)) print(f" round {CFG.T} (after the run) -> {edit_budget(CFG.T, CFG.T, CFG.b_min, CFG.b_max)}") print("\n Early candidates may bundle up to 4 coordinated edits. Note the ceil(): the cosine term is") print(f" only exactly zero at t = T, so inside a {CFG.T}-round run the budget bottoms out at" f" {min(table)}, not {CFG.b_min}.") print(" Every edit in a bundle inherits the bundle's single measurement, so the shrinking budget is") print(" what makes late history attributable to fewer components. It caps how MANY edits ride") print(" together, never WHICH mechanisms the harness may eventually contain.") return "budget " + "".join(str(b) for b in table) edit_budget_schedule() The proposal side regularizes how edits are drafted, not which ones are kept. edit_budget implements the annealed L0 budget from the paper: a cosine schedule from b_max to b_min, and by default it allows up to four coordinated edits per candidate for the first eight rounds, three for the next five, and two for the last seven. The ceiling in the formula means the budget only reaches its minimum of one at t = T, one step after the run ends, which is easy to miss when reading the equation. Because every edit in a bundle inherits the bundle’s one measurement, the shrinking budget is what makes late-run history attributable to fewer components. The budget limits how many edits travel together and never restricts which mechanisms the harness may eventually contain. Copy CodeCopiedUse a different Browser CRITIC_PATTERNS = [ (r"\btask_\d{3}\b", "hard-codes an evolve-set task id"), (r"expected_output|grader|rubric\[", "reads the grader or the expected answer"), ] COMPONENT_SIGNALS = [ ("control_flow", [r"\bretry\(", r"max_attempts"]), ("context_mgmt", [r"compress_context", r"keep_last"]), ("config", [r"CONFIG\["]), ] class ScreenOnly(Domain): name = "screen" critic_patterns = CRITIC_PATTERNS briefs = {"critic": "A coding agent harness."} @section("7. The critic's deterministic layer, and how edits are tagged") def critic_and_tags(): dom = ScreenOnly() diffs = { "memorise answers": "+ memory = Memory('answers')\n+ memory.remember('task_007', cached_patch)", "peek at the grader": "+ if os.path.exists('/grader/expected_output.txt'): return read_it()", "leaked credential": "+ api_key = 'sk-live-0123456789abcdefghijkl'", "empty diff": " ", "general retry rule": "+ for attempt in range(max_attempts): result = retry(step)", } for label, diff in diffs.items(): try: verdict = review(dom, diff, summary=label, targets_mode="evolve") print(f" {label:20s} -> {verdict['verdict']:6s} {verdict['reasons']}") except ZeroDivisionError as e: print(f" {label:20s} -> passed the deterministic layer; the LLM layer raised" f" ZeroDivisionError: {e}") print("\n Gotcha: with RRSI_VERTEX_PROJECTS unset, rrsi.llm.generate computes x % len(projects)") print(" outside its try block, so the helpful 'set RRSI_VERTEX_PROJECTS' error is never reached.") print(" A clean diff is supposed to go to Claude for an intent review; here there is no Claude.") print("\n Every edit is tagged with the component it touches, and a tag needs evidence in the diff:") tag_cases = [ ("skill", "+ \"Remember to run the tests before finishing.\""), ("control_flow", "+ for attempt in range(max_attempts): result = retry(step)"), (None, "+ review = subcall('reviewer', transcript)"), ("memory", "+ context = compress_context(context, keep_last=8)"), ] for declared, diff in tag_cases: tag = normalize(declared, diff, COMPONENT_SIGNALS) print(f" declared {str(declared):13s} -> tagged {tag:13s} {diff[2:60]!r}") print(" A proposer cannot label a prompt tweak as a new 'skill' to look novel: without evidence") print(" the tag falls back to what the diff actually is.") counts = {"prompt": 3, "subagent": 1} print(f"\n novelty(nu) counts STRUCTURAL components {K_STR} the incumbent has never accepted.") print(f" against an incumbent with accepted edits {counts}:") for comps in (["memory"], ["subagent"], ["prompt", "client_tool", "memory"]): print(f" {str(comps):38s} nu = {novelty(comps, counts)}") return "precheck rejected 4/5 diffs without an LLM call" critic_and_tags() The critic screens every candidate diff before any evaluation is spent, in two layers. The first is a deterministic precheck against a generic credential pattern plus the d [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Explore a comprehensive coding guide to Google Research's RRSI (Regularized Recursive Self-Improvement), detailing how noise bands, cost rules, and leakage screens enable safe, ef…

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。