跳到主要內容
AI News HubLIVE
站內改寫3 分鐘閱讀

待翻譯:Anthropic launches Claude Sonnet 5.5 with near-Opus performance at half the price

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Anthropic on Monday launched Claude Sonnet 5.5, the latest version of its workhorse model and the second model in the The post Anthropic launches Claude Sonnet 5.5 with near-Opus performance at half the price appeared first on The New Stack.

來源The New Stack AI作者: Frederic Lardinois
待翻譯:Anthropic launches Claude Sonnet 5.5 with near-Opus performance at half the price
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Anthropic on Monday launched Claude Sonnet 5.5, the latest version of its workhorse model and the second model in the Claude 5.5 family (after Opus 5.5, which launched last week). A new version of Claude Haiku, the smallest and most affordable model in the family, is on the roadmap and will launch “in the coming weeks.” 30% faster, up to 30% cheaper The new Sonnet model, Anthropic says, generates output more than 30% faster than its predecessor, Sonnet 5, while lowering the cost per task by up to 30%. That makes Sonnet 5.5 the company’s fastest Sonnet model yet and, as Anthropic puts it, “a faster, lower-cost complement to Claude Opus 5.5.” Anthropic argues that the model is “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.” Early testers also noted that it “adds polish to user interfaces and can follow slide templates to create decks that require minimal editing.” Like Opus 5.5, Sonnet 5.5 also writes more clearly than Anthropic’s previous generation of models. With Opus 5.5, the company promised a clearer, more natural-sounding voice, and Sonnet 5.5 now shares that trait. Close to Opus on most benchmarks Anthropic says the biggest gains are in coding. On Terminal-Bench 4.0, which tests agents on command-line tasks, Sonnet 5.5 scores 70.6%, up from 10.3% for Sonnet 5 and ahead of Opus 5.5’s 66.4%. (The benchmark’s maintainers noted that Sonnet 5 sometimes ran into timeouts and token limits, which helps explain its low score.) On CursorBench, which uses tasks from real Cursor coding sessions, Sonnet 5.5 comes within about two points of Opus 5.5. On Cognition’s FrontierCode, which checks whether a code change could be merged without human edits, it scores 52.1% at its second-highest effort setting, compared to 54.4% for Opus 5.5 and 49.3% for OpenAI’s GPT-6 Sol. For knowledge work, Sonnet 5.5 scores 1,844 on Artificial Analysis’ GDPval-AA, which ranks models by Elo score on real-world tasks across 44 occupations. That’s only two points behind Opus 5.5 and about 400 points ahead of Sonnet 5. Sonnet 5.5 also bests GPT-6 Sol’s score of 1,487. Credit: Anthropic. The new model also comes close to Opus 5.5 in computer use and on Humanity’s Last Exam. And in a less formal test of long-horizon work and image understanding, Anthropic says it’s the first Sonnet model to beat Pokémon Red working only from screenshots. On several benchmarks, Anthropic says, Sonnet 5.5 at low or medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. Still, the company says Opus 5.5 remains “clearly stronger at complex, open-ended work requiring sustained judgment.” Pricing and availability While Anthropic cut prices for Opus 5.5 (to $4 per million input tokens and $20 per million output tokens), and OpenAI halved prices for its GPT-6 Sol and Luna models on the same day, the pricing for Sonnet 5.5 remains unchanged at $2 per million input tokens and $10 per million output tokens. Cache reads come in at $0.20 per million tokens. That’s the same list price as GPT-6 Sol, OpenAI’s second-best model after Astra. Since the new model uses far fewer tokens, though, at least according to the company’s own measurements, running it should be cheaper than running Sonnet 5. Sonnet 5.5 is now available on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure. Developers who run Sonnet with thinking turned off will need to switch to a new between_tools setting before moving to Sonnet 5.5. Opus 5.5 already rejects requests that turn thinking off entirely. Safeguards Because Anthropic believes Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5.5’s, this is first Sonnet model to launch with the same kind of cyber safeguards Anthropic also uses for its most capable models. Routine bug fixing isn’t affected, the company says, but higher-risk cybersecurity requests will fall back to Sonnet 5. It’s also the first Sonnet model with classifiers designed to stop attackers from extracting its reasoning to train their own models. The post Anthropic launches Claude Sonnet 5.5 with near-Opus performance at half the price appeared first on The New Stack.

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Anthropic on Monday launched Claude Sonnet 5.5, the latest version of its workhorse model and the second model in the The post Anthropic launches Claude Sonnet 5.5 with near-Opus…

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。