LLM & Models

Claude Sonnet 5.5: what the benchmarks and pricing really mean

Anthropic launched Claude Sonnet 5.5 on September 28, six days after Opus 5.5. It's the second model in the 5.5 family, billed as 30%+ faster and up to 30% cheaper per task than Sonnet 5.

Pricing

| Per million tokens | Sonnet 5.5 | Opus 5.5 | |---|---|---| | Input | $2 | $4 | | Output | $10 | $20 | | Cache reads | $0.20 | $0.20 |

The sticker price doesn't drop: it's identical to Sonnet 5. The advertised saving comes from using fewer tokens per task, not from the list price. 1M-token context, model ID claude-sonnet-5-5, available on AWS, Google Cloud and Azure.

The headline scores

According to Anthropic: 70.6% on Terminal-Bench 4.0 (vs 10.3% for Sonnet 5 and 66.4% for Opus 5.5), 1,844 on GDPval-AA (Opus 5.5: 1,846) and 80.1% on OSWorld 2.1. On paper, Sonnet 5.5 nearly matches the flagship at half the price.

The caveats

  • The headline score is at maximum effort. Per Digital Applied's analysis, Terminal-Bench drops to about 43% at "high" effort, the API default, where Opus 5.5 scores 64.2%. The gap widens where most integrations actually run.
  • Cost per task depends on the setting. Artificial Analysis measured Sonnet 5.5 using more tokens per task at max effort than any model it has tested, about 50% more cost than Sonnet 5 in that case.
  • The gains are real at low effort. At "low" or "medium" it beats Sonnet 5's best score for roughly a tenth of the cost per task.
  • Anthropic itself says Opus 5.5 remains clearly stronger on complex work requiring sustained judgment.

In practice

For well-scoped agent work (coding, automation), test Sonnet 5.5 at medium or high and measure cost per completed task, not per token. Keep Opus for open-ended work. If you ran Sonnet with thinking off, the new between_tools setting is required before migrating.

Sources: Anthropic, Decrypt, VentureBeat, Digital Applied. Scores are published by Anthropic and not audited by NextOptimus.