While the industry discusses the bold promises of autonomous agents, Anthropic has focused on solving the most down-to-earth engineering pain point: how to get a model to reliably execute commands in a real shell without falling into endless loops at the first syntax error. The release of Claude Sonnet 5.5 showed that the gap between synthetic coding tests (HumanEval) and the real-world practice of working in a terminal has finally begun to narrow.
Terminal Bench 4.0: A qualitative leap from 10% to 70%
Today, the main benchmark for the maturity of coding agents is not the generation of isolated functions, but Terminal Bench 4.0. In this test, models are given a real Linux environment with incomplete documentation, broken dependencies, outdated packages, and a task to locate and fix a problem in a large codebase.
| Model | Terminal Bench 4.0 (agentic command-line work) | Cursor Bench 4.0 (IDE refactoring) | OSWorld 2.0 (OS control) | Visual interface (“Cartographer”) |
|---|---|---|---|---|
| Claude Sonnet 5.5 | 70.6% | 55.5% | 80.1% | 61.6% |
| Claude Opus 5 | 66.4% | 35.8% | 72.0% | 38.2% |
| Claude Sonnet 5.0 (baseline) | 10.3% | 34.1% | 57.0% | 15.6% |
The difference between Sonnet 5's 10.3% and version 5.5's 70.6% is less about an increase in the number of parameters than a change in its self-reflection logic. When a build fails, earlier models tended to mechanically repeat the same command with slight variations in the flags. Sonnet 5.5 is trained to analyze the error stack trace, halt execution, check system logs via journalctl or dmesg and only then form a hypothesis about how to fix the problem.
Cursor Bench and Context Retention in Large Monorepos
The second critical test is Cursor Bench 4.0, which tests a model's ability to make coordinated changes across a dozen interrelated files at once: updating a method signature in the backend, adapting the database schema, rewriting the client SDK, and updating automated tests.
Sonnet 5.5 scored 55.5%, demonstrating a better understanding of the project's abstract syntax tree (AST). The model learned to filter out irrelevant code when searching with grep/ripgrep, keeping hundreds of unnecessary lines from filling the context window. This made it possible to reduce actual token consumption by about 30% per completed task: although the rates remained at $2 per 1M input tokens and $10 per 1M output tokens, development became noticeably cheaper thanks to fewer iterations.
Computer Use and OSWorld 2.0: An Agent at the Real Screen
Desktop control technology (Computer Use), first introduced by Anthropic in the fall of 2024, took a significant leap forward in version 5.5. In the OSWorld 2.0 test, the score rose to 80.1%.
The breakthrough is due to the introduction of a specialized spatial encoder (the “Cartographer” benchmark showed accuracy rising from 15.6% to 61.6%). The model has learned to:
- Precisely calculate click coordinates for nonstandard interface elements (canvas, WebGL, vector charts);
- Account for interface rendering delays (debouncing) and avoid trying to click elements before CSS animations have finished;
- Handle scenarios involving dynamic modal windows and operating system security confirmations.
Practical Limits of Applicability
Despite the impressive benchmark figures, engineers should take a clear-eyed view of the architecture's limitations:
- Architectural blindness: the model still struggles with high-level decisions that have no definitive technical answer. Choosing between microservices and a monolith, or designing a data sharding schema, still requires expert human intervention.
- Path hallucinations: in monorepos with deeply nested folder structures, Sonnet 5.5 sometimes generates nonexistent relative paths in imports.
- Getting stuck in complex GUIs: if a desktop application freezes or an unexpected OS modal appears, the agent may lose its sequence of actions and require a manual session reset.
Compare models before you start
The service sets its plans, limits and model catalog. If they differ from this article, contact us so we can update it and record a new review date.
Browse modelsAffiliate link: your price stays the same and the project earns a commission.