跳到主要內容
AI News HubLIVE
站內改寫2 分鐘閱讀

待翻譯:OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship appeared first on The New Stack.

來源The New Stack AI作者: Frederic Lardinois
待翻譯:OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The company describes the new model as “an upgrade to GPT-6 Sol that nearly matches GPT-6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices.” Same price, near-Astra performance Pricing for GPT-6.1 Sol remains the same as before, at $2 per million input tokens and $10 per million output tokens, with cached input significantly discounted to $0.10 per million tokens. Click image to enlarge. (Credit: OpenAI.) GPT-6.1 Sol is now available in the API, as well as to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex. For now, though, the model isn’t available in Chat. One new feature is that GPT-6.1 Sol will also come in an Ultrafast version in Codex, with token generation that is up to 8x faster than the standard speed. Although its predecessor is only a week old, the updated model shows significant improvements. In virtually every benchmark OpenAI provided ahead of the announcement, the new model ranks similarly to OpenAI’s costly GPT-6 Astra flagship model, but at a significantly lower cost. Click image to enlarge. (Credit: OpenAI.) For example, on coding benchmarks, GPT-6.1 Sol scores 6.4 percentage points higher than GPT-6 Sol on DeepSWE 1.1, with results that essentially match GPT-6 Astra—but at one-fifth the cost. In some benchmarks, the new Sol model also beats Anthropic’s Opus 5.5 (with fallbacks), which launched on the same day as GPT-6 Sol. On the GDP.pdf benchmark, for example, which tests how the models answer questions about complex PDF documents, GPT-6.1 Sol tops out around 32%, while Opus 5.5 hits about 29%. Here, too, the results are similar to GPT-6 Astra at about one-fifth the cost per task. One area where GPT-6.1 Sol performs especially well is computer use. Here, the new model outperforms its predecessor by seven percentage points at maximum reasoning, at half the cost — and once again with performance in line with Astra. Indeed, given these results, it’ll be hard to justify using Astra for most use cases. Mixed results against Sonnet 5.5 Sadly, only a few comparison benchmarks exist for Sonnet 5.5, which was released on Monday and costs the same $2/$10 per million input/output tokens. Where benchmarks exist for both models, the results are mixed. Sonnet 5.5 scores 71% on DeepSWE, while GPT-6.1 Sol scores about 75%. On AutomationBench, GPT-6.1 Sol scores around 36% compared to 44.7% for Sonnet 5.5, but Sol’s price per task is significantly lower ($0.30 vs. $1.14). Fewer errors, better alignment OpenAI also says GPT-6.1 Sol makes fewer factual errors. At low reasoning effort, the share of responses containing at least one factual error fell from 11.4 percent with GPT-6 Sol to 7.7 percent, a reduction of about 32 percent. These results come from deliberately difficult conversations where users flagged mistakes by earlier models, and they do not represent error rates in typical use. Given that we’re not quite pacing the frontier (yet), it’s good to see that GPT-6.1 Sol brings Sol’s alignment in line with Astra. In general, it is better at respecting user intent and safety constraints, OpenAI says, and only fails to disclose broken search tools 2.8% of the time (instead of guessing). In the company’s tests, the GPT-6.1 Sol-based agents also never tried to work around an automated safety reviewer’s decision to block their agents. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra flagship appeared first on The New Stack.

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • OpenAI on Tuesday launched GPT-6.1 Sol, the newest version of its workhorse model, only a week after launching GPT-6 Sol. The post OpenAI’s new GPT-6.1 Sol undercuts its own Astra…

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。