本文にスキップ
AI News HubLIVE
サイト内リライト2 分で読了

翻訳待ち:OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship

記事の要約

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship appeared first on The New Stack.

ソースThe New Stack AI著者: Frederic Lardinois
翻訳待ち:OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship
誤りを報告

訂正窓口はまだ利用できません。記事情報をコピーして保存できます。

訂正案内
本文へ

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The company describes the new model as “an upgrade to GPT-6 Sol that nearly matches GPT-6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices.” Same price, near-Astra performance Pricing for GPT-6.1 Sol remains the same as before, at $2 per million input tokens and $10 per million output tokens, with cached input significantly discounted to $0.10 per million tokens. Click image to enlarge. (Credit: OpenAI.) GPT-6.1 Sol is now available in the API, as well as to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. For now, though, the model isn’t available in Chat. One new feature is that GPT-6.1 Sol will also come in an Ultrafast version in Codex, with token generation that is up to 8x faster than the standard speed. Although its predecessor is only a week old, the updated model shows significant improvements. In virtually every benchmark OpenAI provided ahead of the announcement, the new model ranks similarly to OpenAI’s costly GPT-6 Astra flagship model, but at a significantly lower cost. Click image to enlarge. (Credit: OpenAI.) For example, on coding benchmarks, GPT-6.1 Sol scores 6.4 percentage points higher than GPT-6 Sol on DeepSWE 1.1, with results that essentially match GPT-6 Astra—but at one-fifth the cost. In some benchmarks, the new Sol model also beats Anthropic’s Opus 5.5 (with fallbacks), which launched on the same day as GPT-6 Sol. On the GDP.pdf benchmark, for example, which tests how the models answer questions about complex PDF documents, GPT-6.1 Sol tops out around 32%, while Opus 5.5 hits about 29%. Here, too, the results are similar to GPT-6 Astra at about one-fifth the cost per task. One area where GPT-6.1 Sol performs especially well is computer use. Here, the new model outperforms its predecessor by seven percentage points at maximum reasoning, at half the cost — and once again with performance in line with Astra. Indeed, given these results, it’ll be hard to justify using Astra for most use cases. Mixed results against Sonnet 5.5 Sadly, only a few comparison benchmarks exist for Sonnet 5.5, which was released on Monday and costs the same $2/$10 per million input/output tokens. Where benchmarks exist for both models, the results are mixed. Sonnet 5.5 scores 71% on DeepSWE, while GPT-6.1 Sol scores about 75%. On AutomationBench, GPT-6.1 Sol scores around 36% compared to 44.7% for Sonnet 5.5, but Sol’s price per task is significantly lower ($0.30 vs. $1.14). Fewer errors, better alignment OpenAI also says GPT-6.1 Sol makes fewer factual errors. At low reasoning effort, the share of responses containing at least one factual error fell from 11.4 percent with GPT-6 Sol to 7.7 percent, a reduction of about 32 percent. These results come from deliberately difficult conversations where users flagged mistakes by earlier models, and they do not represent error rates in typical use. Given that we’re not quite pacing the frontier (yet), it’s good to see that GPT-6.1 Sol brings Sol’s alignment in line with Astra. In general, it is better at respecting user intent and safety constraints, OpenAI says, and only fails to disclose broken search tools 2.8% of the time (instead of guessing). In the company’s tests, the GPT-6.1 Sol-based agents also never tried to work around an automated safety reviewer’s decision to block their agents. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship appeared first on The New Stack.

要点と分析を開く

記事インテリジェンス

エンジニア上級

要点

  • AI 生成が一時的に利用できないため、ソース内容とフォールバックメタデータを保存しました。
  • OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra…

要点と分析は自動生成され、誤りを含む場合があります。原典をご確認ください。