Sonnet 5 Stress Test: What It's Good At, What It's Not, and How It Compares
A Sonnet 5 stress test across the work I actually do: coding, copywriting, and tool calling, plus cost against Opus 4.8, Sonnet 4.6, GPT-5.5, GLM 5.2, MiniMax M3, and Gemini.
Sonnet 5 dropped today. I ran it against the work I actually do, and against the models I actually pay for.
Sonnet 5 shipped this morning. I run Opus and Sonnet 4.6 in Claude Code every day, so my first question was simple: where does this one fit?
So I stress-tested it across the four things I care about. Coding, copywriting, tool calling, and cost.
One honest note before any numbers. Sonnet 5 is hours old as I write this, so the independent leaderboards are still filling in, and Anthropic published its official benchmark table as an image in the system card. Where a figure is third-party I say so, and I'd treat every cross-model score here as directional, not gospel.
How Sonnet 5 Slots Into My Stack
My daily setup is Opus in Claude Code for the hard stuff and Sonnet 4.6 when I want speed. GLM 5.2 sits next to both as a cost backup, which I covered in my GLM 5.2 review.
Anthropic is pitching Sonnet 5 as its most agentic Sonnet yet, landing near Opus 4.8 quality for a lot less money. That's a big claim. It's also exactly the lane where a model either earns a daily spot or stays a tab I forget about.
Coding: A Real Jump, And Nearly Opus
SWE-bench Pro, the harder variant, from a third-party llm-stats aggregate. Read it as directional. Sonnet 5 lands just under Opus 4.8 and ahead of GPT-5.5.
On SWE-bench Pro, the harder coding benchmark, Sonnet 5 scores 63.2%. That's a real generational jump over Sonnet 4.6 at 58.1%, and it sits about six points under Opus 4.8 at 69.2%.
It also edges GPT-5.5 at 58.6% and noses ahead of the cheap-tier coders, GLM 5.2 and MiniMax M3. On the easier SWE-bench Verified the order reshuffles and GPT-5.5 and Opus lead, but the takeaway holds: this is near-Opus coding at Sonnet money.
I'll be straight about what I can and can't claim. I can't spin up five models and race them myself, so the scores above are published, cited, and directional. What I can vouch for is the Opus baseline they're measured against, because this very post was built by Opus 4.8 in Claude Code, holding research, file edits, and verification across one long session without losing the plot.
Copywriting: I'm Keeping Sonnet 4.6
There's no clean benchmark for copy, so this part is opinion, and mine hasn't moved. Sonnet 4.6 is still my writer.
Here's why. The newest models follow instructions more literally, which is great for structured work and a little flattening for voice. Sonnet 5 drafts clean copy, but it needs a tighter brief, with real examples, to stop sounding like a model.
Sonnet 4.6 gives me that human wobble more naturally. I split my writing and coding tools on purpose, the same way I split Claude and Gemini in my Gemini vs Claude review. For now Sonnet 5 is a capable drafter I have to steer, not the writer I reach for.
Tool Calling: This Is Where Sonnet 5 Shines
Tool use is the headline for me, because most of my real work is agentic. On Terminal-Bench, the agentic command-line test, Sonnet 5 scores 80.4%.
That's a 13-point jump over Sonnet 4.6 at 67.0%, and it's within striking distance of Opus 4.8 at 82.7% and GPT-5.5 at 83.4%. For a mid-tier model, that's the gap closing fast.
In Claude Code that jump means tighter, more reliable agent loops for everyday coding. In CoWork, where the work is MCP-heavy orchestration across business tools, Sonnet 5's tool gains plus its lower price make it a strong default driver, with Opus held back for the gnarliest multi-step jobs. The 1M token context window and 128K output ceiling give it room to run long without losing state.
Cost: Where It Gets Interesting
Standard API pricing per million tokens. Sonnet 5 matches Sonnet 4.6's price for near-Opus quality. Intro pricing of $2/$10 runs through August 31, 2026.
Sonnet 5 lands at $3 per million input and $15 per million output, the exact price of Sonnet 4.6. You're paying Sonnet 4.6 money for a clear step up in coding and agentic work.
Against Opus 4.8 at $5/$25, that's roughly 40% cheaper for about six points less on SWE-bench Pro. Against GPT-5.5, it's half the output price, $15 against $30, while scoring higher on that same benchmark.
Then there's the cheap tier. GLM 5.2 at $1.40/$4.40 and MiniMax M3 at $0.60/$2.40 are a fraction of the price and now genuinely close on benchmarks, which is exactly why GLM stays my cost backup. Sonnet 5's real pitch isn't "cheapest," it's "best quality per dollar in the Claude lineup," and the intro price of $2/$10 through August 31 briefly undercuts even GLM on output.
Which Model For Which Job
My working split after the stress test. It'll shift as the independent benchmarks fill in.
Here's where I've landed after a day with it. Sonnet 5 is my new default for everyday coding and for agentic tool work, because it's close enough to Opus and costs a lot less.
Opus 4.8 stays for the hardest long-horizon builds, where its follow-through and SWE-bench Pro lead earn the premium. Sonnet 4.6 keeps the copywriting seat. And when a side project can't justify Claude-grade spend, GLM 5.2 or MiniMax M3 still win on raw price.
The Verdict
Sonnet 5 is the most exciting Sonnet release in a while, because it moves the value line. Near-Opus coding and agentic ability at Sonnet 4.6 pricing is a genuinely good deal, and the tool-calling jump is the part I'll feel every day.
It didn't take the copywriting crown, and it's not the cheapest option on the board. But for the bulk of what I do in Claude Code and CoWork, it's about to become the model I reach for first.
A fair last word. It's a day old, so I'll revisit this once the independent leaderboards catch up and I've run it hard for a few weeks. Testing Sonnet 5 against your own stack? I'd like to hear where it lands for you. Find me on LinkedIn.
More notes
AI ToolsAI Vendor Lock-In Is the Real Risk Right Now: 3 Smart Moves to Make
Anthropic is running Microsoft's old playbook on AI subscriptions, and the bill is coming. Here's what to do before AI vendor lock-in catches everyone flat-footed.
Building Openclaw: One Agent, Many Pipelines, Zero Mondays at a Blank Doc
Why I stopped trying to be consistent on my own and built an agent to run my owned-media operation instead. The boundary between me and Openclaw, and what it actually catches.
AI ToolsThe Day My AI's 'Thinking' Ate 97% of the Token Budget
A debugging story about a pipeline that logged success and posted nothing. The culprit: a reasoning model that consumed almost the entire budget on intermediate thinking before the content tokens fired.
AI ToolsThe Apple–OpenAI Breakup: What Tech's Messiest Divorce Teaches Marketers
Apple and OpenAI went from the biggest distribution deal in tech to a trade-secrets lawsuit in two years. Here's what the breakup teaches marketers about borrowed distribution.
