AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
Claude Sonnet 5.5 Review: Faster Agentic Coding & Visual QA India's Most Futuristic AI Conference Is Back – Bigger, Sharper, Bolder d : h : m : s Career GenAI Prompt Engg ChatGPT LLM Langchain RAG AI Agents Machine Learning Deep Learning GenAI Tools LLMOps Python NLP SQL AIML Projects Reading list How to Become a Data Analyst in 2025: A Complete RoadMap A Comprehensive Learning Path to Tableau in 2025 A Comprehensive NLP Learning Path 2025 Learning Path to Become a Data Scientist in 2025 Step-by-Step Roadmap to Become a Data Engineer in 2025 A Comprehensive MLOps Learning Path: 2025 Edition Roadmap to Become an AI Engineer in 2025 A Comprehensive Learning Path to Master Computer Vision in 2025 Best Roadmap to Learn Generative AI in 2025 GenAI Roadmap for Enterprises Large Language Models Demystified: A Beginner’s Roadmap Learning Path to Become a Prompt Engineering Specialist Claude Sonnet 5.5 Review: Faster Agentic Coding & Visual QA Vasu Deo Sankrityayan Last Updated : 29 Sep, 2026 7 min read A practical guide to what changed, the benchmarks that matter, the cost controls developers should not miss, and one hands-on demo worth building. Anthropic has just released Claude Sonnet 5.5. It is the middle child of the Claude family, and the one most people will actually use. It is quick, capable, cheap to run, and free to use for all users without any subscription. In this article, we go over the latest iteration of the Claude’s Sonnet family. We put it to test to see whether its agentic claims had any truth to them or not. And how a regular user of Claude app benefit with this free upgrade. Table of contents What’s new in Sonnet 5.5? Claude Sonnet 5.5 Features How to Access Claude Sonnet 5.5? Pricing, Speed, and Effort Controls Where Sonnet 5.5 Is Strongest Hands-On: Build a Visual Bug-Fixing Copilot Benchmarks That Matter Conclusion Frequently Asked Questions What’s new in Sonnet 5.5? Available to all users Sonnet 5.5 is now the default model for all users of Claude App. If you use Claude without a subscription, this is the model you are talking to. Opus 5.5 stays behind a paid plan, so for most people, Sonnet 5.5 is simply what Claude is. In short, the following improvements have been made: Task Follow Through: completes complex multi-step tasks fully instead of stopping early. Self-Verification: checks and confirms its own work without being prompted to. Agentic Tool Use: plans, uses tools, executes, and reviews its own output. Lower Cost: cheaper per token than Opus, with a discounted launch price. Improved Reliability: declines bad requests better and hallucinates less often. Why this release matters Sonnet 5.5 is not a replacement for Opus 5.5 on the hardest open-ended work. It is the model to look at when the task is well-scoped, repeatable, tool-heavy, or latency-sensitive. Anthropic positions Sonnet 5.5 as a fast low-cost complement to Opus 5.5. In the Claude apps, Medium effort is the default. On the Claude Platform, High is the default. That difference matters because effort changes latency, token use, and how much the model verifies its own work. Claude Sonnet 5.5 Features Anthropic does not publish Sonnet 5.5 parameter count, layer count, mixture-of-experts layout, or other internal model architecture details… which is expected for any proprietary model. Any article that gives those numbers is speculating. Its key features, are out in the open though: Adaptive thinking lets the model spend more or less reasoning effort depending on the request. The effort control exposes five levels: low, medium, high, xhigh, and max. Adaptive thinking makes it so that you can’t disable effort/reasoning. The 1M-token context window is the default, not a special beta path. A single request supports up to 128K output tokens. The model is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. For technical leaders, this is a useful architecture view: input context, reasoning budget, tool loop, verification behavior, and output. Those are the levers that determine reliability and cost in production. How to Access Claude Sonnet 5.5? One of the biggest advantages of Sonnet 5.5 is that you don’t need a paid Claude subscription to try it. You can access the model through Claude’s free tier, although free users have usage limits that reset every five hours. Claude.ai Webapp: Anyone can sign up for Claude and use Sonnet 5.5 without a Pro subscription. The free tier has usage limits, but you don’t need to pay to access the model. Claude Platform: Developers can access Sonnet 5.5 through the Claude Platform using the model ID claude-sonnet-5-5. API usage is billed separately based on token consumption. Cloud Platforms: Sonnet 5.5 is also available through Amazon Web Services, Google Cloud, and Microsoft Azure for developers and organizations that want to integrate the model into their applications. Pricing, Speed, and Effort Controls Cost item Sonnet 5.5 Opus 5.5 Input tokens $2 / MTok $4 / MTok Output tokens $10 / MTok $20 / MTok Cache write $2.50 / MTok (5m); $4 / MTok (1h) $5 / MTok (5m); $8 / MTok (1h) Cache read $0.20 / MTok $0.20 / MTok The key change is cost per task, not cost per token. Sonnet 5.5 keeps Sonnet 5 pricing, but Anthropic says it often finishes with fewer tokens and fewer tool calls, which can lower the total bill by up to 30%. Use Low or Medium for chat, fast iteration, and clearly scoped agent steps. Start at Medium for well-specified coding and multi-step tool use. Move to High for harder or longer coding work. Reserve Xhigh and Max for workloads where your own evaluations show a measurable gain. Do not assume the effort setting you used on Sonnet 5 should carry over. Anthropic recommends re-running your eval sweep. Where Sonnet 5.5 Is Strongest Here are some of the tasks/domains across which Sonnet 5.5 has delivered state-of-the-art performance: Coding and repo-level work: The release is especially strong for bug fixes, multi-file changes, code review, and tool-driven engineering tasks. Anthropic says early testers saw fewer steps because Sonnet 5.5 batches tool calls more efficiently. Every unnecessary tool round-trip costs time and tokens. Documents, slides, and spreadsheets: Anthropic highlights polished documents, slides, and spreadsheets as a sweet spot. This is useful for teams that want a model to turn research or analysis into business-ready artifacts without paying Opus-class pricing for every request. Vision and computer use: Sonnet 5.5 shows a large gain on Chartography and OSWorld. Anthropic also says it is the first Sonnet model to beat Pokémon Red using only screenshots. The lesson is broader than gaming: screenshot-driven workflows, visual QA, desktop automation, and chart interpretation are now much more credible Sonnet use cases. Long-context work: A 1M-token context window is useful, but it does not remove the need for context engineering. For repeated sessions, caching and selective retrieval can still be cheaper and more controllable than dumping the same giant context into every turn. Hands-On: Build a Visual Bug-Fixing Copilot A basic ‘Hello, Claude‘ example does not show why this model is interesting. A better demo is visual QA: give Sonnet 5.5 a screenshot of a broken web page and the page’s CSS, then ask it to diagnose the mismatch and produce the smallest safe fix. For this test we’d be using a CSS file named styles.css containing the style code for this page: Prompt: “You are debugging this customer-success dashboard. Compare the screenshot with the attached CSS and identify the visual issues. For each issue: Explain the likely CSS rule causing it. Propose the smallest safe change. Avoid redesigning the page or changing unrelated styles. Return a corrected style.css. Briefly explain how you would verify that each fix worked. Preserve the existing visual design and make the layout responsive.” Output: Sonnet 5.5 completed the debugging task in around 10 seconds and did more than just rewrite CSS. It correctly mapped visual issues to specific rules, suggested minimal fixes, added verification steps, and clearly called out areas where it was uncertain. What stood out most was its ability to combine screenshot understanding with code-level reasoning. I would still verify the changes in a browser before production use, but for multimodal debugging and frontend QA, the response was fast, practical, and surprisingly precise. Hands-On 2: Making a Video using Claude Prompt: Make a modern slick and punchy video for a modern startup that works on Artificial Intelligence. It succeeds because the prompt leaves room for creative interpretation while giving the model three strong anchors: modern, slick, and punchy, with AI/startup as the subject. If the result has strong pacing, clean motion graphics, confident typography, and avoids the usual generic “AI glowing brain” bullshit, it’s a very strong output. Benchmarks That Matter Official benchmark snapshot recreated from Anthropic’s Sonnet 5.5 release blog There is a significant jump over Sonnet 5 is large in agentic coding and computer use (10% -> 70%). Terminal-Bench 4.0 rises from 10.3% to 70.6%, CursorBench 4.0 moves from 34.1% to 55.5%, and OSWorld 2.1 moves from 57.0% to 80.1%. On GDPval-AA and AA-Briefcase, Sonnet 5.5 lands very close to Opus 5.5, which helps explain why Anthropic is positioning it for everyday knowledge work. Benchmark caveat worth remembering More effort is not always better. Anthropic reports that Sonnet 5.5 scored lower at Max than at Xhigh on FrontierCode because extra review sometimes caused timeouts or out-of-scope edits. In agentic systems, overthinking can be a real failure mode. Benchmark score should not be your only selection criterion. Measure completion rate, tool-call count, latency, token usage, and how often a human has to repair the result. Conclusion Claude Sonnet 5.5 is compelling because the upgrade is practical. It is faster, uses fewer tokens on many tasks, is substantially stronger at agentic coding and visual work, and keeps the same per-token price as Sonnet 5. For teams building coding agents, visual QA systems, document workflows, or tool-using assistants, it is an obvious model to evaluate. The important lesson is to evaluate the system, not just the model. Tune effort, preserve prompt-cache behavior, define verification, and control scope. Sonnet 5.5 can be very efficient when the task is clear. It can also spend extra time and tokens when you ask it to be maximally thorough. The best deployments will treat those controls as part of the application architecture. Note: Some of the images used in this article have been sourced from the official Sonnet 5.5 release blog. Frequently Asked Questions Q1. Is Sonnet 5.5 cheaper than Sonnet 5? A. The per-token price is the same, but Anthropic says completed tasks can cost up to 30% less because Sonnet 5.5 often uses fewer tokens and tool calls. Q2. Does Sonnet 5.5 have a 1M-token context window? A. Yes. Anthropic lists 1M tokens as the default context window and 128K tokens as the maximum output for a normal request. Q3. What effort level should I start with? A. For well-scoped agentic coding, Anthropic recommends starting at Medium and moving to High for harder or longer tasks. For general API usage, High is the platform default. Q4. Is Max effort always the best? A. No. Anthropic reports at least one benchmark where Max scored lower than Xhigh because extra review caused timeouts or out-of-scope edits. Q5. Should Sonnet 5.5 replace Opus 5.5? A. Not for every workload. Anthropic still positions Opus 5.5 as stronger for the hardest open-ended work that needs sustained judgment. Vasu Deo Sankrityayan Studying, evaluating, and explainin [truncated for AI cost control]