BackBlog / AI Tools
8 min read·

Sonnet 5 Stress Test: What It's Good At, What It's Not, and How It Compares

A Sonnet 5 stress test across the work I actually do: coding, copywriting, and tool calling, plus cost against Opus 4.8, Sonnet 4.6, GPT-5.5, GLM 5.2, MiniMax M3, and Gemini.

Preston Vawdrey

Preston Vawdrey

SEO Marketing Expert

Claude Sonnet 5 stress test cover: coding, copy, tools, and cost versus Opus 4.8, Sonnet 4.6, GPT-5.5, GLM 5.2, MiniMax M3, and Gemini Sonnet 5 dropped today. I ran it against the work I actually do, and against the models I actually pay for.

Sonnet 5 shipped this morning. I run Opus and Sonnet 4.6 in Claude Code every day, so my first question was simple: where does this one fit?

So I stress-tested it across the four things I care about. Coding, copywriting, tool calling, and cost.

One honest note before any numbers. Sonnet 5 is hours old as I write this, so the independent leaderboards are still filling in, and Anthropic published its official benchmark table as an image in the system card. Where a figure is third-party I say so, and I'd treat every cross-model score here as directional, not gospel.

How Sonnet 5 Slots Into My Stack

My daily setup is Opus in Claude Code for the hard stuff and Sonnet 4.6 when I want speed. GLM 5.2 sits next to both as a cost backup, which I covered in my GLM 5.2 review.

Anthropic is pitching Sonnet 5 as its most agentic Sonnet yet, landing near Opus 4.8 quality for a lot less money. That's a big claim. It's also exactly the lane where a model either earns a daily spot or stays a tab I forget about.

Coding: A Real Jump, And Nearly Opus

Bar chart of SWE-bench Pro coding scores showing Opus 4.8 at 69.2, Sonnet 5 at 63.2, GLM 5.2 at 62.1, MiniMax M3 at 59.0, GPT-5.5 at 58.6, Sonnet 4.6 at 58.1, and Gemini 3.1 Pro at 54.2 SWE-bench Pro, the harder variant, from a third-party llm-stats aggregate. Read it as directional. Sonnet 5 lands just under Opus 4.8 and ahead of GPT-5.5.

On SWE-bench Pro, the harder coding benchmark, Sonnet 5 scores 63.2%. That's a real generational jump over Sonnet 4.6 at 58.1%, and it sits about six points under Opus 4.8 at 69.2%.

It also edges GPT-5.5 at 58.6% and noses ahead of the cheap-tier coders, GLM 5.2 and MiniMax M3. On the easier SWE-bench Verified the order reshuffles and GPT-5.5 and Opus lead, but the takeaway holds: this is near-Opus coding at Sonnet money.

I'll be straight about what I can and can't claim. I can't spin up five models and race them myself, so the scores above are published, cited, and directional. What I can vouch for is the Opus baseline they're measured against, because this very post was built by Opus 4.8 in Claude Code, holding research, file edits, and verification across one long session without losing the plot.

Copywriting: I'm Keeping Sonnet 4.6

There's no clean benchmark for copy, so this part is opinion, and mine hasn't moved. Sonnet 4.6 is still my writer.

Here's why. The newest models follow instructions more literally, which is great for structured work and a little flattening for voice. Sonnet 5 drafts clean copy, but it needs a tighter brief, with real examples, to stop sounding like a model.

Sonnet 4.6 gives me that human wobble more naturally. I split my writing and coding tools on purpose, the same way I split Claude and Gemini in my Gemini vs Claude review. For now Sonnet 5 is a capable drafter I have to steer, not the writer I reach for.

Tool Calling: This Is Where Sonnet 5 Shines

Tool use is the headline for me, because most of my real work is agentic. On Terminal-Bench, the agentic command-line test, Sonnet 5 scores 80.4%.

That's a 13-point jump over Sonnet 4.6 at 67.0%, and it's within striking distance of Opus 4.8 at 82.7% and GPT-5.5 at 83.4%. For a mid-tier model, that's the gap closing fast.

In Claude Code that jump means tighter, more reliable agent loops for everyday coding. In CoWork, where the work is MCP-heavy orchestration across business tools, Sonnet 5's tool gains plus its lower price make it a strong default driver, with Opus held back for the gnarliest multi-step jobs. The 1M token context window and 128K output ceiling give it room to run long without losing state.

Cost: Where It Gets Interesting

Grouped bar chart of API price per million tokens, input and output, for MiniMax M3 at 0.60 and 2.40, GLM 5.2 at 1.40 and 4.40, Gemini 3 Pro at 2 and 12, Sonnet 5 at 3 and 15, Sonnet 4.6 at 3 and 15, Opus 4.8 at 5 and 25, and GPT-5.5 at 5 and 30 Standard API pricing per million tokens. Sonnet 5 matches Sonnet 4.6's price for near-Opus quality. Intro pricing of $2/$10 runs through August 31, 2026.

Sonnet 5 lands at $3 per million input and $15 per million output, the exact price of Sonnet 4.6. You're paying Sonnet 4.6 money for a clear step up in coding and agentic work.

Against Opus 4.8 at $5/$25, that's roughly 40% cheaper for about six points less on SWE-bench Pro. Against GPT-5.5, it's half the output price, $15 against $30, while scoring higher on that same benchmark.

Then there's the cheap tier. GLM 5.2 at $1.40/$4.40 and MiniMax M3 at $0.60/$2.40 are a fraction of the price and now genuinely close on benchmarks, which is exactly why GLM stays my cost backup. Sonnet 5's real pitch isn't "cheapest," it's "best quality per dollar in the Claude lineup," and the intro price of $2/$10 through August 31 briefly undercuts even GLM on output.

Which Model For Which Job

Table matching jobs to models: everyday coding to Sonnet 5, hardest long-horizon builds to Opus 4.8, copywriting to Sonnet 4.6, tool calling and agents to Sonnet 5, and rock-bottom cost to GLM 5.2 or MiniMax M3 My working split after the stress test. It'll shift as the independent benchmarks fill in.

Here's where I've landed after a day with it. Sonnet 5 is my new default for everyday coding and for agentic tool work, because it's close enough to Opus and costs a lot less.

Opus 4.8 stays for the hardest long-horizon builds, where its follow-through and SWE-bench Pro lead earn the premium. Sonnet 4.6 keeps the copywriting seat. And when a side project can't justify Claude-grade spend, GLM 5.2 or MiniMax M3 still win on raw price.

The Verdict

Sonnet 5 is the most exciting Sonnet release in a while, because it moves the value line. Near-Opus coding and agentic ability at Sonnet 4.6 pricing is a genuinely good deal, and the tool-calling jump is the part I'll feel every day.

It didn't take the copywriting crown, and it's not the cheapest option on the board. But for the bulk of what I do in Claude Code and CoWork, it's about to become the model I reach for first.

A fair last word. It's a day old, so I'll revisit this once the independent leaderboards catch up and I've run it hard for a few weeks. Testing Sonnet 5 against your own stack? I'd like to hear where it lands for you. Find me on LinkedIn.

Marketing that actually moves the needle

Occasional notes on SEO, paid ads, and growth, plus every new post, straight to your inbox. Written for operators, not skimmers.

No spam. Unsubscribe anytime.