Skip to content
← Back to Blog
AISeptember 30, 2026•6 min read

Medium Effort, Cheaper Model: What 270 Test Runs Taught Us About Choosing Claude Models

Anthropic's advice for its newest models is simple: start at medium effort, and pay for more only where you have measured a gain. We measured it on our own work. For well-specified coding, Claude Sonnet 5.5 at medium effort matched Claude Opus 5.5 at about 60% lower API cost. For reviewing code, it did not come close.

By Ivaylo Tsvetkov, Co-Founder

Medium Effort, Cheaper Model: What 270 Test Runs Taught Us About Choosing Claude Models featured image

What Anthropic recommends

Anthropic's prompting guides for Opus 5.5 and Sonnet 5.5 both treat effort as a dial you tune, not a quality setting you max out.

  • Opus 5.5 defaults to medium effort. Anthropic advises keeping xhigh and max for work where you have measured a gain.
  • Sonnet 5.5 defaults to high. For agentic coding, the guide says to "start at medium for well-specified tasks and move to high for harder or longer ones".
  • At low and medium effort, Sonnet 5.5 "is more likely to stop and check in with the user before it finishes". The guide suggests a short instruction telling it to keep working until the job is done.
  • On model choice, Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment", while Sonnet 5.5 is strongest at "well-scoped everyday tasks".

That is sound advice. It still leaves the question every team has to answer for itself: where exactly does my work sit on that line?

How we tested

We run a pipeline of AI agents that turns written specs into working code, with a separate reviewer checking each result. To test models on it, we replayed six real coding jobs from that pipeline with their original inputs.

  • Scored by real tests. Each job was graded by the acceptance tests it originally shipped against, as the share of checks passed.
  • Rules written first. Before any run, we fixed what would count as "as good" (within 3 points), the confidence level (90%) and the cost bar.
  • Paired and repeated. Both setups ran on the same jobs at the same time, 3 to 5 times each. Single runs mislead: the same job on the same model varied by up to 65 points.
  • Served model verified. We read which model actually answered each run, rather than trusting what was requested.

Across four experiments that came to about 270 runs.

Result 1: medium effort is enough

On Opus 5.5, medium effort coded as well as high across 48 paired runs, and used about a quarter less time and tokens. The average score gap was 0.1 points.

  • First round, 18 pairs: medium 3.7 points ahead, 24% less time, 27% fewer output tokens.
  • Second round, 30 pairs: high 2.1 points ahead, 21% less time and 23% fewer tokens at medium.
  • Combined, 48 pairs: a 0.1-point gap, well inside the noise.

This confirms Anthropic's guidance directly. An earlier comparison of xhigh against high showed the same pattern: more thinking did not mean better code on these jobs.

Result 2: Sonnet 5.5 codes like Opus 5.5

Sonnet 5.5 at medium effort matched Opus 5.5 on all six jobs, at about 60% lower API cost and in about 40% less time. Each setup ran 3 times on each job.

  • Opus 5.5, medium: average score 82.0, 1.28M output tokens.
  • Sonnet 5.5, medium: 87.8, 1.02M tokens, 60% lower API output cost.
  • Sonnet 5.5, high: 85.8, 1.60M tokens, 37% lower cost.
  • Sonnet 5.5, medium plus Anthropic's "keep working" instruction: 85.7, 0.92M tokens, 64% lower cost.

All three Sonnet setups cleared the "no more than 3 points worse" bar, and no job that Opus reliably passed was failed by Sonnet. Raising Sonnet to high made it slower and more expensive, not better. Costs use API list prices: Sonnet $2/$10 and Opus $4/$20 per million input/output tokens.

Result 3: reviewing code is a different job

As a code reviewer, Sonnet 5.5 caught 38 of 54 planted defects. The Opus model we use for review caught all 54. The review test hides known problems in finished coding runs, such as a report claiming a change the code doesn't contain, or a spec rule that allows two readings.

  • Opus 5 at high: 54 of 54.
  • Opus 5.5 at high: 49 of 54.
  • Opus 5.5 at medium: 46 of 54.
  • Sonnet 5.5 at high: 38 of 54.

Sonnet's weakest area was ambiguous spec rules, where it caught 1 of 9. That lines up with Anthropic's own positioning: judgment-heavy work is where Opus keeps its lead. Being good at writing code did not make a model good at judging it.

Two more things worth knowing

Sonnet at medium asks before it improvises. In one variant of the test, each brief named a folder outside the agent's workspace, while the same files sat in its working folder. Opus found the copies and carried on every time, and so did Sonnet at high. Sonnet at medium stopped to ask in 7 of 18 runs. This is the check-in behaviour Anthropic describes. For unattended agents, keep every brief consistent with the environment and add the keep-working instruction.

The spec is a bigger lever than the model. The gaps between jobs were far larger than the gaps between models or effort levels. Two jobs peaked at about 70% on every setup, and the lowest scores all came from one job whose spec left a command-line format open. Pinning interfaces and output formats is cheaper than any model upgrade.

Do we agree with Anthropic?

Yes, and our data sharpens their guidance in two places.

  • The Sonnet default. Sonnet 5.5 defaults to high effort. For well-specified coding, medium scored as well with about a third fewer tokens and a third less time. Anthropic's own advice for agentic coding says to start at medium; our numbers say to stay there unless a task proves harder.
  • "Coding" is not one job. Anthropic draws the line between well-scoped work and open-ended judgment. Our results show that line runs through coding itself: writing code from a clear spec is well-scoped, and there Sonnet matched Opus. Reviewing that code is judgment work, and there Opus stayed clearly ahead.

What we run now: Sonnet 5.5 at medium writes the code, Opus 5.5 at medium steps in when a check fails or the spec is open-ended, and Opus reviews. Their guidance gives the direction. Only your own evals show where your work sits on the line.

Want to discuss how this applies to your business? Book a free call.

Ready to add AI to your business?

We help businesses identify, design, and deploy AI systems that actually work. Book a free discovery call and see what's possible.

Book a free call →

Related Posts

View all posts →