AI News HubLIVE
サイト内リライト8 分で読了

翻訳待ち:Who is at the frontier of terminal tasks?

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Newly released models repeatedly appear near GPT-5.6 Sol on Artificial Analysis’s Terminal-Bench 2.1 leaderboard. Are these models actually on par with Sol on terminal tasks? We picked the top 20 models from that leader…

ソースHacker News AI著者: amrrs

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Newly released models repeatedly appear near GPT-5.6 Sol on Artificial Analysis’s Terminal-Bench 2.1 leaderboard. Are these models actually on par with Sol on terminal tasks? We picked the top 20 models from that leaderboard and benchmarked them on TB-fn to answer that question. TB-fn is Fidian’s variant of Terminal-Bench, built from the same 89 tasks as TB-2.1. TL;DR TB-fn separates models that look comparable on TB-2.1. The score range widens from ~15 to ~22 points, and the top tier narrows from five labs to OpenAI and Anthropic. Six models lose ~6–11 points. Grok 4.5 loses ~11 points, Kimi K3 ~10, and GLM-5.2 ~9. The other three are Gemini 3.6 Flash (~8), GLM-5.3 (~7), and DeepSeek V4 Pro (~6). These models span a wide spectrum on TB-2.1, from top-tier performers like GLM-5.3 down to lower-ranked options such as Gemini 3.6 Flash. Per-attempt cost rises on TB-fn in every task for every model. Individual increases range from about 12% to 97%. Median turns rise from 16 to 19, with average increases of 45% for input tokens and 32% for output tokens. Claude Opus 5 max effort costs 3x the medium effort with no observed TB-fn gain. All three Opus 5 effort settings score ~82% on TB-fn. The medium effort costs $1 per task, xhigh costs 2x, and max costs 3x. Confirmed security refusals affect Anthropic’s Fable 5 and Opus 5, Google’s Gemini Flash 3.6 and 3.7, and OpenAI’s GPT-5.6 Sol. If Fable 5 were to fall back to Opus 4.8, its performance would improve by approximately 7 points on TB-2.1 and 10 points on TB-fn, reaching roughly 82% on both benchmarks. Open-weight models show reduced performance on TB-fn compared to TB-2.1. While GLM-5.3 maintains its lead within this category, its performance gap versus Sol max expands significantly, and drops out of Tier 1 status. Scores, rank intervals and tiers across both benchmarks Order by Models Metric How to read this chart Circle = pass@1 · square = pass@3 · filled indigo = TB-fn · violet outline = TB-2.1. The bold mark is the ranking series; the pale mark is the comparison; the hairline connects them. When comparing the same metric across benchmarks, the connector is black if the drop is statistically significant (95% CI of the difference excludes zero) and red if it exceeds 4σ. Soft band: 95% CI of the ranking series. On pass@1, TB-2.1 is shown as two broad presentation tiers and TB-fn as four ordered groups labeled to roughly match the TB-2.1 tiers’ score ranges: Tier 1 aligns with TB-2.1 Tier 1, Tiers 2a and 2b subdivide the TB-2.1 Tier 2 range, and Tier 3 falls below it. Cross-benchmark tiers are shown for pass@1 only; pass@3 ranks as a continuous list. Boundaries are maximum-likelihood contiguous partitions of the sorted scores, with models weighted by bootstrap SE. The displayed tier counts are fixed at two and four for cross-benchmark interpretation rather than selected by BIC. The labels express ordering and rough score alignment, not exact nesting: models reorder between benchmarks. Tiers are descriptive groupings, not significance statements. The score axis is truncated at 60. Refusals score 0, as served. CIs capture within-task trial noise only. pass@3 = bootstrap: 3 trials drawn with replacement per task per replicate, scored if any passes (B=10,000). Rank intervals and tooltips carry the statistics. Table view Figure 1: Combined leaderboard showing both benchmarks. Each row shows one evaluated model at one effort setting. Pass@1 is the average of the 89 task-level pass rates. A filled dot marks the hardened TB-fn pass@1 score, joined by a line to a hollow dot for original TB-2.1. A band indicates the 95% score interval, with the rank interval at left. The All/Open-weight families control filters the rows while retaining ranks and tiers calculated over the full set of models. GPT-5.6 Sol (max) +0.6 86.2 85.6 10 10 0.46M·91% 0.40M·92% 35.7k 29.6k 2.2k 1.8k 6.0 4.0 $1.52 $1.27 $1.76 $1.49 Claude Opus 5 (xhigh) +1.2 82.5 81.3 15 13 0.89M·94% 0.68M·94% 64.7k 50.7k 2.8k 2.3k 8.1 5.0 $2.37 $1.84 $2.87 $2.26 GPT-5.6 Sol (xhigh) −2.3 82.3 84.6 8 8 0.41M·92% 0.25M·90% 22.4k 15.3k 1.4k 1.2k 3.9 2.9 $1.07 $0.73 $1.30 $0.86 Claude Opus 5 (max) −0.4 81.8 82.2 17 16 1.06M·94% 0.90M·95% 92.9k 74.6k 3.7k 3.1k 13.2 8.2 $3.23 $2.59 $3.95 $3.16 Claude Opus 5 (medium) +0.0 81.8 81.8 11 10 0.64M·96% 0.36M·95% 22.6k 14.1k 1.2k 0.9k 2.2 1.5 $1.04 $0.65 $1.27 $0.79 Grok 4.6 (high) −4.4 78.7 83.1 15 13 1.07M·94% 0.70M·92% 43.4k 27.8k 1.6k 1.2k 4.0 2.9 $0.90 $0.60 $1.15 $0.73 GLM-5.3 −7.0 77.6 84.6 20 14 1.25M·96% 0.91M·96% 67.0k 50.2k 2.3k 2.1k 12.7 6.9 $0.67 $0.50 $0.86 $0.60 GPT-5.6 Terra (max) −1.8 75.7 77.5 11 11 0.91M·95% 0.81M·95% 80.9k 71.9k 3.5k 3.4k 7.5 5.9 $1.25 $1.12 $1.66 $1.45 Claude Opus 4.8 (max) −0.2 75.5 75.7 18 17 0.93M·92% 0.77M·93% 116.2k 86.7k 4.4k 3.5k 19.5 12.2 $3.80 $2.85 $5.04 $3.77 Qwen 3.8 Max −2.4 74.5 76.9 15 14 0.72M·87% 0.58M·86% 80.3k 60.9k 3.4k 2.8k 17.1 10.4 $0.83 $0.65 $1.11 $0.85 Kimi K3 (max) −10.2 72.4 82.6 11 9 0.94M·90% 0.58M·87% 26.4k 18.2k 1.1k 1.0k 3.0 2.6 $0.83 $0.57 $1.14 $0.69 Gemini 3.7 Flash (high) −4.6 72.4 77.0 33 30 3.54M·86% 2.63M·86% 43.4k 34.4k 0.8k 0.8k 3.5 3.2 $0.76 $0.58 $1.05 $0.75 Claude Fable 5 −3.0 72.1 75.1 10 10 0.36M·91% 0.22M·89% 27.4k 18.3k 2.0k 1.4k 2.7 1.9 $2.10 $1.41 $2.92 $1.87 GPT-5.6 Luna (max) −3.3 71.5 74.8 13 12 1.70M·97% 1.27M·96% 70.4k 56.7k 2.5k 2.2k 5.1 3.5 $0.13 $0.11 $0.19 $0.14 Claude Sonnet 5 (max) −2.2 69.9 72.1 26 19 2.49M·96% 1.56M·96% 174.3k 120.6k 4.4k 3.9k 27.4 17.1 $2.46 $1.67 $3.53 $2.31 Gemini 3.7 Flash (medium) −4.5 69.7 74.2 32 27 3.80M·88% 2.02M·88% 34.0k 22.9k 0.6k 0.5k 2.6 1.9 $0.71 $0.40 $1.02 $0.54 DeepSeek V4 Pro (0813) −6.1 68.6 74.7 16 14 1.92M·84% 1.18M·75% 97.4k 71.6k 2.8k 2.6k 9.6 6.6 $0.48 $0.37 $0.70 $0.50 GLM-5.2 (max) −8.5 67.4 75.9 12 12 2.08M·82% 1.23M·84% 65.8k 43.7k 2.0k 1.8k 5.8 4.4 $0.53 $0.28 $0.79 $0.36 Grok 4.5 (high) −10.7 66.6 77.3 15 12 1.54M·95% 0.79M·94% 27.7k 16.1k 0.9k 0.7k 3.2 2.0 $0.76 $0.42 $1.13 $0.54 GLM-5.3 Flash −5.3 66.6 71.9 39 19 1.13M·35% 0.98M·83% 53.3k 50.8k 1.4k 1.9k 54.7 13.4 $0.15 $0.07 $0.22 $0.10 Muse Spark 1.2 (xhigh) −5.4 66.3 71.7 20 18 3.01M·61% 2.85M·70% 129.2k 116.3k 2.8k 2.7k 6.3 4.8 $2.29 $1.87 $3.45 $2.61 DeepSeek V4 Flash −4.8 65.5 70.3 23 20 2.06M·90% 1.23M·87% 130.9k 88.9k 3.4k 3.0k 7.2 4.7 $0.03 $0.02 $0.05 $0.03 Gemini 3.6 Flash (medium) −7.9 64.5 72.4 46 40 6.25M·87% 4.09M·86% 51.6k 36.7k 0.7k 0.6k 4.0 2.9 $1.23 $0.84 $1.91 $1.16 Each cell shows TB-fn above TB-2.1. Input tokens carry their cache-hit rate. Drop is TB-fn minus TB-2.1 in pass@1 points, black once its 95% interval excludes zero and red beyond four standard errors. Ordered by TB-fn pass@1; click any column to sort, and a third click restores this order. Sorting uses the TB-fn value. GPT-5.6 Sol (max) scores ~86%, while Sol xhigh and all three Opus 5 effort settings score ~82%. OpenAI and Anthropic are the only labs in the tier. Grok 4.6 (~79%), GLM-5.3 (~78%), and Kimi K3 (~72%) fall below that tier; xAI, Z.ai, and Moonshot leave it. Which models hold and which fall on TB-fn Dot = TB-fn − TB-2.1 (percentage points) · soft band = 95% CI of the difference · dashed line = no change. Black dot = significant drop (95% CI excludes zero); red dot = drop beyond 4σ. Figure 2: Change in pass@1 from TB-2.1 to TB-fn for the evaluated models and effort settings. Dots show differences with shaded 95% confidence intervals. Violet dots indicate intervals including zero, black exclude zero, and red mark falls exceeding four standard errors. Claude Opus 5 (xhigh) gains ~1 point, and Grok 4.5 has the largest fall at ~11. The cross-model score range widens from ~15 to ~22 points, which naturally leads to more tiers. Each model is evaluated at the same effort setting, through the same serving route, and under the same scoring rule on both benchmarks; only the task set changes. Six drops exceed four standard errors: Grok 4.5 falls ~11 points, Kimi K3 ~10, GLM-5.2 ~9, Gemini 3.6 Flash ~8, GLM-5.3 ~7, and DeepSeek V4 Pro ~6. These models span ~72% to ~85% on TB-2.1, and the gap between Grok 4.6 and Grok 4.5 roughly doubles from ~6 to ~12 points. Which rewrite causes each fall remains unknown. TB-fn makes every model work harder Benchmark Models Dot + logo per config · vertical bar = 95% CI · dashed line = Pareto frontier among the models shown · shaded bands retain the full-field pass@1 lineage tiers · hover for score, CI, and $/attempt. GLM-5.3 is included because the GLM family is open-weight; its own weights were not public at the time of writing. Figure 3: TB-fn pass@1 is plotted against cost per task attempt. The cost axis is inverted so cheaper is to the right. The efficiency frontier is drawn through its members in bold, with trails linking each result to its TB-2.1 position. The All/Open-weight families control filters the points and recomputes the efficiency frontier among the models shown. Per-attempt cost rises on TB-fn in every task for every model. This increase occurs because median turns rise from 16 to 19, mean input tokens per attempt rise 45%, and mean output tokens rise 32%. A model whose pass rate holds still moves toward higher cost. The performance-cost frontier consists of the models for which no cheaper alternative has a higher pass@1. GLM-5.2, Grok 4.5, and GLM-5.3 Flash leave the frontier on TB-fn, while Grok 4.6, Opus 5 medium, and Sol xhigh enter it. Open-weight families occupy part of the low- and mid-cost frontier, while OpenAI and Anthropic occupy its highest-scoring end. Open-weight families fall further behind on TB-fn We group GLM, Qwen, Kimi, and DeepSeek as open-weight families. GLM-5.3 is included because the GLM family is open-weight, but its own weights were not public at the time of writing. GLM-5.3 leads the open-weight families on both benchmarks. On TB-2.1 it scores 84.6%, 1.0 point behind Sol max at 85.6%. On TB-fn it scores 77.6%, 8.6 points behind Sol max at 86.2%, ranks seventh, and is in Tier 2a. No open-weight family reaches Tier 1 on TB-fn. Four of the six models whose scores fall by more than four standard errors are from open-weight families: Kimi K3 (-10.2 points), GLM-5.2 (-8.5), GLM-5.3 (-7.0), and DeepSeek V4 Pro (-6.1). GLM-5.3 costs $0.67 per attempt on TB-fn versus $1.52 for Sol max, and DeepSeek V4 Flash costs $0.03 per attempt. On the open-weight-family efficiency frontier, TB-2.1 contains DeepSeek V4 Flash, GLM-5.3 Flash, GLM-5.2, and GLM-5.3, while TB-fn contains DeepSeek V4 Flash, GLM-5.3 Flash, DeepSeek V4 Pro, and GLM-5.3. Claude Opus 5 max costs 3x medium for the same TB-fn score Opus 5 at medium gives up about two points on a SWE-bench Pro subset for roughly half the cost of default high, according to Anthropic’s current cost guidance. Opus 5 xhigh outperforms max on Anthropic’s official Frontier-Bench v0.1 effort curve. All three Opus 5 effort settings score ~82% on both benchmarks. The medium setting costs $1.04 per task, xhigh costs $2.37, and max costs $3.23. Medium achieves the same pass rate as max at one-third the cost. Security refusals by model and effort setting Security refusals affect Anthropic’s Fable 5 and Opus 5, Google’s Gemini Flash 3.6 and 3.7, and OpenAI’s GPT-5.6 Sol. vulnerable-secret is the one task all five refuse on both benchmarks. The table shows the number of tasks refused at least once and the pass-rate headroom lost to refusal averaged across all 89 tasks. Claude Fable 5 13 12 12.4 12.6 Gemini 3.7 Flash 4 8 3.5 7.0 Claude Opus 5 xhigh 3 3 3.4 3.4 Claude Opus 5 max 3 3 3.4 3.4 Claude Opus 5 medium 3 3 3.4 3.4 Gemini 3.6 Flash 2 2 1.4 2.2 GPT-5.6 Sol max 2 1 1.1 0.7 GPT-5.6 Sol xhigh 2 2 0.6 1.1 TB-fn is shown above TB-2.1 in each cell. Select a heading to sort. To compute the degree of performance degradation, we calculate a hypothetical fallback by routing every refused attempt from Fable and all three Opus 5 effort setting [truncated for AI cost control]