AI News HubLIVE
In-site rewrite3 min read

DeepSeek V4 Flash 0731: 82.7% on Terminal-Bench 2.1 with a public harness

Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-Bench 2.1 One harness to unlock the pote…

SourceHacker News AIAuthor: ubermon

Ante Terminal Bench 2.1 Results | Coding Agent Benchmark Skip to main content We just open sourced a tiny GPT-style cognitive core built in pure Rust.See our repository→ Terminal-Bench 2.1 One harness to unlock the potential of all models Compare Ante runs across models on the same Terminal-Bench 2.1 task set, using consistent parameters and verified benchmark results. Best Accuracy82.7% Leading Model DeepSeek V4 Flash 0731max Task Set89 tasks Trials368 passed / 445 trials We benchmark what we ship.Every eval uses a pinned public . No eval-only branches or benchmark-specific prompts. The runs are auditable.Every result links its raw Harbor run, so anyone can inspect the trials behind the number. The constraints are official.All trials follow the official Terminal-Bench parameters: 89 tasks, 5 trials per task, strict timeouts, and hardware limits. Model org All model orgsTB 2.1 · Same parameters: 89 tasks · 5 trials/task · Updated Aug 9, 2026 #ModelSame-modelAgentSource 1 DeepSeek V4 Flash 0731max 82.7% ±1.79 SE $68.4138.9 min#1 same-model Ante0.preview.710.preview.71 Harbor linkAug 9, 2026› 2 Grok 4.5mediumⓘReward hacking was identified in 20 trajectories. The displayed score excludes them; see PR #129 in Source for details. 80.9% ±1.27 SE $242.578.4 min#1 same-model Ante0.preview.560.preview.56 PR #129Jul 10, 2026› 3 GLM 5.2 74.6% ±2.06 SE $260.1111.3 min#1 same-model Ante0.preview.430.preview.43 Harbor linkJun 20, 2026› 4 DeepSeek V4 Pro 69.1% ±2.24 SE $26.3448.4 min#1 same-model Ante0.preview.540.preview.54 Harbor linkJul 7, 2026› 5 DeepSeek V4 Flash 66.4% ±2.27 SE $49.9841.4 minNo public rows Ante0.preview.530.preview.53 Harbor linkJul 5, 2026› 6 MiMo V2.5 65.8% ±2.30 SE $73.7612.5 min#1 same-model Ante20260625-0824-e383a9220260625-0824-e383a92 Harbor linkJun 25, 2026› 7 MiniMax M3 62.1% ±2.33 SE $121.0013.9 min#1 same-model Ante20260623-0825-ff174ee20260623-0825-ff174ee Harbor linkJun 23, 2026› 8 Qwen3.6 27B 56.2% ±2.36 SE Local61.6 minNo public rows Ante20260701-0837-70d2aac20260701-0837-70d2aac Harbor linkJul 3, 2026› Same parameters: 89 tasks · 5 trials/task · Updated Jul 17, 2026 · Source: Terminal-Bench official verified rows For how different models perform on TB 2.1, see Vals AI's Terminal-Bench 2.1 benchmark. 17 official rows #AgentModelAccuracyRun date 1 Claude CodeAnthropic Fable 5xhigh 83.8% ±1.16 SE Jun 7, 2026› 2 DeepSeek V4 Flash 0731Grok 4.5Ante + DeepSeek V4 Flash 0731 model · 82.7%Ante + Grok 4.5 model · 80.9%would slot between public #2 and #3 CodexOpenAI GPT-5.5xhigh 83.2% ±1.13 SE May 1, 2026› 3 Terminus 2Terminal-Bench Fable 5high 80.5% ±1.16 SE Jun 5, 2026› 4 Cursor CLICursor Grok 4.5high 79.3% ±1.46 SE Jul 9, 2026› 5 Claude CodeAnthropic Opus 4.8high 78.9% ±1.31 SE Jul 9, 2026› 6 CodexOpenAI GPT-5.6 Terramax 78.4% ±1.25 SE Jul 11, 2026› 7 Terminus 2Terminal-Bench GPT-5.5xhigh 78.0% ±1.22 SE May 1, 2026› 8 mini-SWE-agentPrinceton Muse Spark 1.1xhigh 76.2% ±1.23 SE Jul 9, 2026› 9 CodexOpenAI GPT-5.6 Lunamax 75.7% ±1.32 SE Jul 11, 2026› 10 GLM 5.2Ante + GLM 5.2 model · 74.6%would slot between public #10 and #11 Claude CodeAnthropic Sonnet 5high 74.6% ±1.64 SE Jul 9, 2026› 11 DeepSeek V4 ProAnte + DeepSeek V4 Pro model · 69.1%would slot between public #11 and #12 Terminus 2Terminal-Bench Gemini 3 Prohigh 73.9% ±1.29 SE May 1, 2026› 12 DeepSeek V4 FlashAnte + DeepSeek V4 Flash model · 66.4%would slot between public #12 and #13 Claude CodeAnthropic Opus 4.7max 68.9% ±1.41 SE May 1, 2026› 13 Terminus 2Terminal-Bench Opus 4.7max 66.1% ±1.37 SE May 1, 2026› 14 Gemini CLIGoogle Gemini 3 Prohigh 65.8% ±1.38 SE May 1, 2026› 15 MiMo V2.5Ante + MiMo V2.5 model · 65.8%would slot between public #15 and #16 Gemini CLIGoogle Gemini 3.1 Prohigh 65.8% ±1.67 SE May 5, 2026› 16 MiniMax M3Ante + MiniMax M3 model · 62.1%would slot between public #16 and #17 Terminus 2Terminal-Bench Gemini 3.1 Prohigh 65.6% ±1.65 SE May 5, 2026› 17 Qwen3.6 27BAnte + Qwen3.6 27B model · 56.2%would slot below public #17 Claude CodeAnthropic GLM-5.1max 58.6% ±1.24 SE May 1, 2026›