Skip to content
AI News HubLIVE
In-site rewrite6 min read

[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%

Summary

overshadowing more efficient GPT6 models from OpenAI

[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

OpenAI made a valiant effort with GPT-6 Sol and Luna launching 50% lower than GPT-5.6, but with 17M views on the launch and counting, today was always going to belong to Claude Opus 5.5, “the first model in our new Claude 5.5 family” performing like “Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.” Opus 5.5 beats Fable or challenges Astra at most benchmarks, and both labs credited efficiency work for the API price cuts, but there are HUGE double digit gains everywhere from prefill to decode to overall compute… … with offsetting inefficiency in token usage on some frontier tasks. HOWEVER something that is a rare emphasis in the Claude launch was the writing improvements: “It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.” We can confirm - here is today’s AINews section run on Opus 5.5 and Sol 6. The difference is night and day - we are migrating to Opus 5.5 immediately for AINews going forward until we reach the next model/version of AINews. They have also published initial work on large multiagent swarms (and efficiency): AI News for 9/21/2026-9/22/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies! AI Twitter Recap Top Story: Claude Opus 5.5 launch, numbers, and reactions What happened Anthropic shipped Claude Opus 5.5, the first model in a new Claude 5.5 family. Its pitch is Fable 5.1‑level capability at Opus pricing, with more speed and better writing. OpenAI released GPT‑6 Sol and Luna about an hour later. Launch claims. Opus 5.5 “performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5” (@claudeai; @AnthropicAI). Where it leads. Anthropic says it leads on agentic coding, computer use, and knowledge work (@claudeai). Speed and cost. It is about 30% faster and about 40% cheaper per task than Opus 5 (@ClaudeDevs, @lydiahallie). Communication fixes. The model puts the most important information up front and follows user writing rules. This targets the most common feedback on Opus 5 (@claudeai). Subscription changes: 5‑hour session limits are up 20%. Lower pricing means limits go 25% further. Pro, Max, and Team users get a banked rate‑limit reset they can use whenever they choose (@claudeai, @ClaudeDevs, @trq212). New defaults. Opus 5.5 is now the default in Claude Code and the Claude app, including Cowork. Default effort is medium, described as “comparable to Fable 5.1 on intelligence but faster” (@_catwu). Availability. It is live in Claude Code and the Claude Platform API (@ClaudeDevs), and in Claude Tag for Slack (@_catwu). Roadmap. Sonnet 5.5 and Haiku 5.5 follow “in the coming weeks” (@mikeyk, @AiBattle_). This contradicts rumors that Haiku was discontinued (@kimmonismus). Safeguards. Opus 5.5 is the first Opus with Fable 5.1‑class safeguards on cyber, bio, and frontier LLM development. Flagged requests fall back to another model, and Anthropic says it is “working to reduce incorrect flags” (@ClaudeDevs). Pre-release signals. The model was spotted in Claude Code shortly before the announcement (@kimmonismus). System card. It was published at launch (@scaling01). Pricing and token economics (facts) List price. Token pricing was cut 20%, from $5/$25 to $4/$20 per 1M input/output tokens (@ValsAI). Offset by higher token use. Vals notes Opus 5.5 often uses more tokens, especially on coding, where it posts its largest gains. The lower sticker price is partly offset by usage. Artificial Analysis cost breakdown. At max effort, Opus 5.5 costs $5.98 per Intelligence Index task versus $5.86 for Opus 5 (max). Their decomposition (@ArtificialAnlys): Higher token usage alone would raise cost per task about 80%, to $10.51. The 20% base-price cut brings that to $8.41. Cheaper cache reads ($0.20) bring it to $5.98. What that means. At max effort, the per‑task saving over Opus 5 disappears. The “40% cheaper” claim applies to default (medium) settings. Relative to Fable 5.1. Cline reports Opus 5.5 beats Fable 5.1 on the Artificial Analysis Intelligence Index at about 2.5x lower cost (@cline). Prompt caching. Switching effort mid‑session does not break the prompt cache on Claude Code v2.1.280+ (@lydiahallie). Model size (speculation). @theo claimed Opus 5.5 is smaller than Opus 5 and credited post‑training. This was not confirmed in official posts. Benchmarks and independent evals Anthropic’s own table. Opus 5.5 beats Fable 5.1 on every row of Anthropic’s headline comparison and beats GPT‑6 Astra on most (@kimmonismus, @synthwavedd, @scaling01). @ShayneRedford (Anthropic) summarized the claimed gains: Stronger than Astra on CursorBench, KWBench, and OSWorld. Much better style and instruction following. Stronger science and health capabilities. More robust against cyber and bio misuse. Third‑party and partner evals: EvalResultSourceVals Index#1, up 2 spots / 2 pts vs Opus 5; Anthropic holds the top three spots (GPT‑6 Sol pending)@ValsAIVals RSI Index#1; first model to beat the published reference on LM Training under their protocol; beats Fable 5.1@ValsAI, @ValsAIFrontierSWE (Proximal)62.3%, #2 behind GPT‑6 Astra (65.5%); ahead of Fable 5.1 (56.3%) and Opus 5 (52.0%)@ProximalHQFrontierCode 1.1 (Cognition)65.3% on Extended; takes #1 from Fable 5 “at a fraction of the cost”@cognitionCursorBench57.8% (Max), new top model; 40% less per task than Opus 5@cursor_aiPerplexity WANDR0.610 at $4.13/task; slightly above Fable 5.1 at 67.6% lower cost@perplexity_aiParseBench (tables)93.9%, +7 pts over Opus 5; beats Fable, Gemini, Astra@jerryjliu0Roboflow vision/detection”By far the best vision model from Anthropic”; now among the models ahead of Google on the Playground leaderboard@skalskip92, @skalskip92 Eval details and caveats: Vals run settings. RSI was run in native Claude Code at max effort, with 1M context, 128K max output tokens, and temperature 1 (@ValsAI). ParseBench caveats. The model still struggles on charts, formatting, and layout. At 5.8¢/page, LlamaIndex calls it too expensive for production OCR. That verdict comes from a vendor with a competing product. AI R&D vs coding. @eliebakouch reads the system card as “roughly similar on AI R&D but a beast on agentic coding.” Saturation. @scaling01 asked whether CoBench is “cooked.” @synthwavedd joked about a new benchmark that launched already saturated. Arena. Opus 5.5 is in Agent Arena and in Battle Mode for WebDev, Text, Vision, and Document. No scores yet (@arena). Effort‑scaling anomaly. On an agentic coding chart, xhigh effort costs about 2.8x more than medium for a 3.2‑point lower score (@LearnOpenCV). @Yuchenj_UW called it the “most bizarre benchmark result” and advised sticking with medium. @nrehiew_ offered an explanation: Opus 5 showed the same pattern on FrontierCode. FrontierCode penalizes unnecessary changes, and higher effort produces scope creep. As a result, models “consistently perform worse at higher reasoning efforts.” System card details Multi‑agent scaling. The system card reports scaling up to 100 parallel agents in Section 8.12. @scaling01 called it the first lab report of its kind. @maksym_andr highlighted it as evidence on multi-agent scaling laws. ProgramBench caveats. ProgramBench author @OfirPress flagged that Anthropic’s near‑100% solve rate comes from a 166/200 subset. That subset likely excludes the hardest programs, such as FFmpeg and the PHP compiler. He also flagged a metric mismatch (@OfirPress, @OfirPress): Anthropic reports average test pass rate. ProgramBench reports full task completion. Partial solves often pass 60–70% of tests, which inflates the pass-rate metric. Comparison with Mythos 5.1. Opus 5.5 outscores Mythos 5.1 on Anthropic’s ECI and beats it on every tested cyber eval (@scaling01, @scaling01). Odd misalignment finding. @teortaxesTex quoted a passage: malicious output occurred “almost exclusively in cases where, prior to the malicious output, Claude made an improbable, innocuous mistake.” He asked whether Anthropic had “sleeper-agent[ed] themselves.” “Trained from RSI.” He separately quoted a line about “the first model trained from RSI” and called it concerning (@teortaxesTex). Biomedical imaging. @iScienceLuvr welcomed the reported biomedical image analysis capabilities. Requests for more. @scaling01 asked for time horizons without chain-of-thought. Safety posture and safeguard controversy Official position: Sam Bowman: Opus 5.5 is “sufficiently safer than its predecessors that releasing it, more likely than not, reduces risks related to misalignment,” especially for the most extreme alignment risks (@sleepinyourhat, @sleepinyourhat). He also acknowledged worry about keeping pace with escalating risk, while saying current tools remain trustworthy at this capability level (@sleepinyourhat). Mike Krieger cited extensive alignment testing and outside evaluation, including by METR (@mikeyk). Friction: Over-triggering fallback. @iScienceLuvr got downgraded to the fallback model after asking Opus 5.5 to cure cancer. China targeting (single test). @xlr8harder says a quick test suggests the frontier-LLM-development classifiers target Chinese hardware. He calls for more probing. Reactions to the China angle. @teortaxesTex framed this as Anthropic undermining Chinese AI. @jakehalloran1 read it as protecting Trainium know‑how. “Pacing the frontier” framing: @theo argued none of today’s releases were Astra‑ or Fable‑tier and that this is deliberate pacing. @goodside said lab calls to pace the frontier have weakened his “pause and do what?” stance. @dejavucoder mocked the framing, given that Opus 5.5 outperforms Fable 5.1. Writing, prompting, and behavior Writing fixes from staff. “We fixed the writing” (@_sholtodouglas) and “we fixed the accent” (@NotTomBrown). Unusual candor. @nmca (Anthropic) posted: “way, way, way better than Opus 5. Sorry about that model.” @theo called it a wild tweet that signals looser comms. Em dashes. @theo reports they are gone from output. It was the most‑engaged reaction post. Anthropic’s prompting playbook (@ClaudeDevs): Hand over a whole task and define “done” and check‑in points. Drop “think carefully,” since the model always thinks first. After a long run, ask what it needs to go further. Why old tricks break. @dbreunig notes old prompt tricks now clash with the model’s training, an argument for re‑compilable prompt optimization. Long-run steering. @omarsar0 highlights Anthropic’s prompt for long runs, where the model sometimes stops to report instead of continuing. Bug report. The live model sometimes generates user turns (@BlackHC). Writing quality in practice. Hamel Husain livestreamed “Is Slop Dead?” testing its writing (@HamelHusain). @nptacek shared a one‑shot result from a personal writing eval. Vision, 3D, and code-as-art demos Improved perception. Sholto Douglas says the 5.5 series has “a serious step up” in 3D understanding and modeling, and that the model “can see now; it was a bit blind before” (@_sholtodouglas, @_sholtodouglas). Painting in code. @jkeatn had the model generate paintings with pure Python, pixel by pixel: About 7,500 lines of code using standard libraries to emulate brush styles. No image model and no reference images. Sholto contrasts this “manual brush” creativity with diffusion models (@_sholtodouglas). Blender scenes. Alex Albert showed Blender claymations from one prompt on claude.ai (@alexalbert). He also built a source‑grounded 1906 San Francisco Market Street: Built from Sanborn maps, period film, and archival photos. Procedural generators only, with no downloaded meshes or textures (@alexalbert, prompt). @karpathy riffed on the idea: turn historical images or video into custom GTA‑style w [truncated for AI cost control]

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • overshadowing more efficient GPT6 models from OpenAI

Highlights and analysis are generated automatically and may contain errors. Check the original source.