OSWorld 2.0: Top AI Agents Finish Just 20.6% of Tasks
July 31, 2026
OSWorld 2.0 tests AI agents on workflows that take humans 1.6 hours. The best model finishes 20.6% of them — after scoring 83.4% on the older OSWorld.
OSWorld 2.0 tests AI agents on workflows that take humans 1.6 hours. The best model finishes 20.6% of them — after scoring 83.4% on the older OSWorld.
Claude Sonnet 5 brings near-Opus 4.8 agentic coding at $2/$10 per million tokens through August. See the pricing, the tokenizer catch, and the system card.
Anthropic shipped Claude Opus 4.8 on May 28, 2026: 88.6% SWE-bench Verified, dynamic workflows with up to 1,000 subagents, and a 3x cheaper fast mode.