AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA for Qwen3-1.7B and Qwen3-8B students. The takeaway: recovery is learnable, and standard OPD rarely teaches it. TL;DR Size: A training method, not a model. Tested on Qwen3-1.7B and Qwen3-8B students, plus a Nemotron-3.5-SFT student on SWE-Bench Verified. Runs on: Trained on NVIDIA H100 nodes. Adds 0 inference cost, so the trained agent runs wherever its base model runs. Performance: First on all 8 per-benchmark averages against 13 baselines, across 3 seeds. Best: Recovers from 72.7% of replayed pivotal mistakes, vs 20.3% for standard OPD. Worst: 55.9% on ALFWorld “Look” tasks with the 1.7B student, vs 83.9% for SOD. Bottom line: Best: teaches recovery that outcome-only RL cannot reach. Worst: depends on replayable environments and a teacher whose pivots match the oracle in 77.8% of failed rollouts. What is a pivotal mistake in a multi-turn agent? A pivotal mistake is an action that lengthens the shortest remaining path to finishing a task, or makes it unsolvable. ALFWorld’s symbolic oracle measures this at every turn. Across Qwen3-8B, Qwen3-30B-A3B and Qwen3-235B-A22B, 59% of failed rollouts (155 of 262) contained one. The first pivotal turn arrived early, at a median of turn 8 to 12 out of 30. The agents then wasted 18 to 21 more turns without recovering. In replays of Qwen3-8B failures, correcting the pivotal turn raised success from 8% to 59%. Leaving the mistake in place and forcing the right action for the next 2 turns still reached 58%. Why does standard on-policy distillation miss it? Standard OPD lowered the held-out failure rate from 79% to 56%. Failures after a pivotal turn only fell from 51% to 49%. The correct action stayed below 1% probability at every pivotal turn, so 8 rollouts rarely sample it. Outcome-based RL shares the blind spot: if every rollout fails, the group-relative advantage is 0. How does PivotOPD work? PivotOPD adds 3 components to group-based RL, combined in a single PPO update. Pivot detection: A larger teacher model reads each rollout and its outcome in hindsight. It picks candidate turns and names a gold action at each. A turn counts as pivotal when the student’s action differs from the gold action. On ALFWorld, detected pivots land within 1 turn of the oracle’s pivot in 77.8% of failed rollouts on average. Preventive distillation (reverse KL): A frozen copy of the student, hinted with the gold action, re-scores the student’s own response. This pushes the student away from the committed mistake. Recovery distillation (forward KL): After each pivot, the teacher names a recovery action for up to K turns. The hinted self-teacher writes recovery responses, and the unhinted student trains on them. Forward KL is mass-covering, so it lifts actions the student almost never samples. The teacher only names actions. Token-level targets come from the student’s own hinted distribution. How does PivotOPD perform on agent benchmarks? With the 1.7B student, PivotOPD averages 73.7% on ALFWorld, 5.5 points above SDAR. It averages 44.5% on Search-based QA, 5.9 points above RLSD. On WebShop it beats RLSD by 1.2 in score but by 14.1 in success rate (76.6%). With the 8B student, it reaches 93.0% on ALFWorld, 47.4% on Search-based QA and 81.9% WebShop success. Margins are smaller, at least 1.8 points. With Qwen3-8B as its own teacher, PivotOPD still wins all 3 benchmarks by at least 1.5 points, 3.9 on average. On SWE-Bench Verified, a Nemotron-3.5-SFT student taught by Nemotron-3-Super went from 62.8% to 66.0%. Standard OPD reached 63.0%, and the teacher scores 73.0%. Recovery is the standout. Across 72 replayed pivotal mistakes, PivotOPD recovered 72.7% of the time, vs 8.3% for the base model, 20.3% for standard OPD and 45.8% for preventive-only. It averaged 9.7 turns to recover, against an optimal 6.2. How does PivotOPD compare with other agent distillation methods? MetricsPivotOPDOPSD (standard OPD)OPIDSDARRLSD Core ideaPrevent and recover at teacher-detected pivotal turnsPrivileged self-teacher on every responseHierarchical skills from on-policy hindsightOPSD as sigmoid-gated auxiliary lossSelf-distillation sets update size, RLVR sets direction Trains on teacher-written recovery responsesYes (forward KL)NoNoNoNo ALFWorld avg, Qwen3-1.7B73.712.461.368.260.7 Search QA avg, Qwen3-1.7B44.536.837.537.138.6 WebShop success, Qwen3-1.7B76.610.168.057.062.5 ALFWorld avg, Qwen3-8B93.046.783.783.890.8 CodeComing soonGitHubGitHubNot disclosedNot disclosed Score source: PivotOPD Table 1. What does it cost to train, and can you run it? PivotOPD changes only training, so inference costs nothing extra. On ALFWorld with the 1.7B student, on 4 H100 GPUs, the overhead over GRPO is 12.4% at K = 1. It jumps to 94.2% at K = 2, because later recovery turns need environment replay. The selected budget is K = 2 on ALFWorld and K = 1 on WebShop and Search-based QA. Key Takeaways Over half of failed agent rollouts hinge on 1 early, recoverable mistake. Standard OPD barely touches these failures: 51% to 49%. PivotOPD recovers from 72.7% of replayed mistakes, vs 20.3% for OPD. Best average against 13 baselines on ALFWorld, WebShop and Search QA. 0 inference overhead, but code is not yet released. Check out the paper and the project page. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. [Sponsored] The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn’t. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free. The post NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes appeared first on MarkTechPost.