Anthropic presented Opus 5 as a major technological leap, and the figures from the company’s own tests confirm this. Some developers saw a different picture in their day-to-day work: verbosity, slow responses, and a tendency to overcomplicate simple tasks. We examine where the gap between a benchmark record and the experience of using the model on a real project comes from.
What the tests showed
On the internal Frontier-Bench v0.1 benchmark, Opus 5 scored 43,3 points versus 18,9 for Opus 4.8—more than twice the predecessor’s result and significantly ahead of its nearest competitor. The model holds a strong position near the top of public programming rankings.
Its strengths are also confirmed outside testing: long-running autonomous tasks, research work, computer control, and complex visual projects. Here, the model works for a long time without human intervention and sees the task through to completion.
What was seen in working conditions
The code-checking service CodeRabbit ran the model on its controlled sample containing known defects. Opus 5 found fewer of these defects than the company’s baseline system while producing roughly four times as many low-significance comments.
Token usage proved to be the second problem. On the same code-review task, Opus 5 used about 60,5 thousand input tokens and 9,5 thousand output tokens per call, compared with approximately 40,5 thousand and 5,8 thousand for GPT-5.6. In other words, it reads about 50 percent more and writes 65 percent more for the same result.
Theo Brown, a developer and head of T3 Chat, described the model’s behavior as follows: it treats almost any comment in the code as a critical problem and suggests rewriting thousands of lines. After a day of intensive work, he estimated the actual savings from switching to Opus 5 at approximately 20 percent instead of the figures expected from the pricing—the result of the higher token usage per task.
Where the gap comes from
Benchmarks and workdays are structured differently, which explains most of the discrepancy.
- In a test, the task is fully specified, the success criterion is known in advance, and the context is limited. The model is optimized precisely for this.
- At work, the instructions are usually incomplete, requirements change along the way, and a small problem is expected to receive a small fix. A model inclined toward exhaustive answers loses out here even if it has greater intelligence.
- Thoroughness that earns points on a long autonomous task turns into noise in a short review: a fourfold stream of minor comments tires the reviewer more than it helps.
None of these points means that the model is weak. They mean that the metric and the use case have diverged.
Why an in-house benchmark is a weak argument
Frontier-Bench is Anthropic’s own test. That does not make the figures fabricated, but it does mean that the task set and evaluation criteria were chosen by a party interested in the outcome. CodeRabbit’s independent check was conducted on a different task set and produced the opposite result: fewer defects found alongside a larger stream of comments.
Both figures can be true at the same time. They simply measure different things, and neither answers the question of whether working will become easier for a particular team.
What to do about it
Look at the cost per task, not the price per million tokens. A model that reads half again as much and writes two-thirds more consumes the pricing difference—and conversely, it may pay for itself if it solves a complex task in fewer attempts.
Divide tasks by type. Long autonomous work, research, and analysis of a large repository are scenarios where thoroughness is justified. Routine reviews, minor fixes, and short refactorings are scenarios where a simpler, cheaper model is more advantageous.
Test it on your own code, not on someone else’s charts. Running it on a dozen typical tasks from your project while tracking tokens and time will tell you more than any public ranking. It will also show how many of the model’s comments you actually accept and how many you close without even looking.
Compare models before you start
The service sets its plans, limits and model catalog. If they differ from this article, contact us so we can update it and record a new review date.
Browse modelsAffiliate link: your price stays the same and the project earns a commission.