待翻譯:Pushing GPT-5.6 Luna from 0% to 56% on ARC-AGI-3 Public
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s SwarmOS tested this hypothesis on the AR…
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
Research August 27, 2026 For long-horizon agents, model capability alone does not determine system capability. Infrastructure orchestration multiplies what models can do. INT21’s SwarmOS tested this hypothesis on the ARC-AGI-3 benchmark and found that infrastructure can dramatically unlock a model’s capabilities. Key Takeaways SwarmOS achieved a 100% RHAE (Relative Human Action Efficiency) score on ARC-AGI-3 Public using only GPT-5.6-Sol. This is the first time GPT-5.6 has matched Anthropic models on an ARC-AGI-3-style benchmark with action-efficiency parity. The result shows that infrastructure architecture is the differentiator. We also tested GPT-5.6-Luna, the least expensive frontier model in this comparison, with SwarmOS and improved its RHAE score from 0% to 56%. Model intelligence and agent intelligence are not the same thing. SwarmOS is a large-scale, cloud-native evolution infrastructure for generating system software, including faster video, audio, and music inference compared with SGLang and vLLM. Without any ARC-specific specialization, it achieved a full score on ARC-AGI-3 Public. We believe the unit of intelligence is shifting from the model to the system, and that this is an early sign of AGI. For enterprises budgeting for expensive frontier models, this result suggests that cheaper models paired with superior orchestration can deliver strong long-horizon performance. Results ModelOfficialINT21 SwarmOSAPI Price (Input/Output) GPT-5.6-Sol13.3%100%$4.0/$20.0 GPT-5.6-Luna0.00%56%$0.2/$1.2 Claude Opus 530.16%No Support$5.0/$25.0 SwarmOS reached a full RHAE score on ARC-AGI-3’s public set with Sol, a 7.5x improvement. With Luna, RHAE improved from 0% to 56% despite Luna being a substantially smaller and less expensive model. This suggests that long-horizon intelligence is not solely a property of the foundation model; system architecture can contribute a meaningful share of the capability. SwarmOS with GPT-5.6-Sol solved 25 environments and 183 levels, taking 6,731 environment actions to achieve a 100% RHAE score. For reference, NVIDIA’s AVO research project, co-authored by INT21 founder Bing Xu earlier this year, reported 6,624 environment actions. VISTA reported 7,542 environment actions to reach the same score on the same public set. Both NVIDIA AVO and VISTA are powered by Anthropic Claude Opus 5. These results cover the 25-environment ARC-AGI-3 public set; they are not results from the semi-private or fully private competition sets. We believe SwarmOS achieving a full score on ARC-AGI-3 is an early sign of AGI (artificial general intelligence). The self-improving agent swarms override learned patterns when they reach dead ends, preventing them from drifting toward globally suboptimal systems. This means orchestration systems can work across very different domains, from pixel games to complex infrastructure generation. The ability to challenge assumptions and change direction when evidence demands it mirrors an aspect of human general intelligence. The NVIDIA AVO Context NVIDIA AVO demonstrated that an orchestration harness can unlock frontier performance. NVIDIA connected Claude Opus 5 to the AVO Harness and reached 100% on ARC-AGI-3, a 3.3x improvement. SwarmOS, our cloud-native platform for running self-improving agents, builds on this finding and tests a more challenging setting: using only GPT-5.6-Sol from a 13.3% baseline, we achieved 100%, a 7.5x improvement, with action-efficiency parity against Claude Opus 5-based solutions. What We Tested ARC-AGI-3 is a benchmark that measures how well AI agents learn and reason through unfamiliar, game-like environments. Agents must infer how the environments work without explicit instructions. They succeed by preserving useful knowledge, learning from failures, and progressing across dozens of challenges without human intervention. For this evaluation, we connected the same SwarmOS that powers our system-programming work to the ARC-AGI-3 task interface, with web search turned off and no ARC-specific specialization. We tested two models: GPT-5.6-Sol and GPT-5.6-Luna. SwarmOS coordinates self-improving agent swarms to explore solutions in parallel, share what they learn, and continuously improve results. The architecture was the same across both models; only the underlying LLM changed. Why This Matters NVIDIA’s AVO research showed that an orchestration harness can shift a frontier model, Opus 5, from 30% to 100% on ARC-AGI-3. SwarmOS independently validates this finding and demonstrates a 0% to 56% leap with GPT-5.6-Luna. This changes the competitive conversation because model performance is no longer the only constraint. AI users can stop asking only, “Which frontier model should we use?” and start asking, “Which orchestration layer maximizes our ROI?” The discussion now focuses on infrastructure as the new moat. What’s Next The industry narrative to date has centered on model capability and which company builds the best frontier model. SwarmOS and AVO suggest a different question for the next phase: “Which organization orchestrates its models most effectively?” While competition between frontier models has been fierce, we believe an equally exciting battle in infrastructure is emerging. Enterprises that treat orchestration as a core engineering competency, rather than a bolt-on integration layer, will have a clear competitive advantage in self-improving infrastructure. SwarmOS is purpose-built for this shift: a platform where self-improving infrastructure becomes your competitive moat.