🎙️ Episode 34006:36July 31, 2026

OSWorld 2.0: Top AI Agents Finish Just 20.6% of Tasks

Listen to this episode

AI-generated discussion by Alex and Jamie

About this episode

Join hosts Alex and Jamie in this episode of the Nerd Level Tech AI Cast as they unpack the intriguing results of OSWorld 2.0, the latest benchmark challenging AI agents with complex, real-world tasks. Discover why even the top performers like Claude Opus 4.8 and GPT-5.5 are struggling to complete lengthy workflows, and what this reveals about the capabilities—and limitations—of AI in professional settings. Tune in for a lively discussion that blends tech insights with a dash of humor!

Transcript

[Alex]: Welcome back to the Nerd Level Tech AI Cast, your friendly neighborhood podcast where AI benchmarks meet meme culture and we still can’t get our smart fridge to stop sending us spam.

[Jamie]: Hey everyone! I’m Jamie, your slightly confused but very enthusiastic tech co-pilot. And as always, sitting across from me—virtually, of course—is Alex, our resident code whisperer and breaker-downer of all things complex.

[Alex]: That’s right. Today we’re diving into the latest AI agent showdown: OSWorld 2.0. Spoiler alert—the bots are struggling, but not in the way you might think.

[Jamie]: Yeah, I saw the headline and thought, “Wait, only 20.6 percent? Did they try turning it off and back on again?”

[Alex]: Tempting, but this one’s a bit deeper. OSWorld 2.0 is the new benchmark on the block, measuring how well AI agents can handle real, professional computer workflows—think the kind of stuff that takes a human about an hour and a half, not just clicking buttons or filling out one form.

[Jamie]: So, not just “open calculator and do math,” but actually, like, reconcile expense reports, file reimbursements, or build a CAD model from scratch?

[Alex]: Exactly! They designed 108 of these long, complicated workflows—spanning business, creative, engineering, healthcare, you name it. And here’s the wild part: the best AI agents, like Claude Opus 4.8 and GPT-5.5, finished just over 20 percent of them completely, end to end.

[Jamie]: Wait, but didn’t these same models ace the previous benchmark? Wasn’t it, like, 80-something percent?

[Alex]: Good memory. On the old OSWorld—let’s call it OSWorld 1.0—the top agents were hitting scores in the 80s. But here’s the catch: those tasks were much shorter. Like, 30 steps or less. Version 2.0 stretches each task over 250 steps on average. Think of it like running a sprint versus a marathon.

[Jamie]: So, the AI is basically passing out halfway through the long run?

[Alex]: Pretty much. There’s even a “horizon wall”—a point where completion rates just fall off a cliff. If a task takes over about 2 hours, none of the current agents can finish it.

[Jamie]: Ouch. So what’s tripping them up? Is it, like, they can’t click the right buttons, or is it something sneakier?

[Alex]: Great question. It’s not about basic computer skills. The paper says the issue isn’t GUI control or coding. The real problem is state management—they lose track of what’s changed, forget constraints, miss new information that pops up midway, or skip double-checking their own work.

[Jamie]: So they’re basically like me trying to follow a recipe while watching YouTube and answering Slack messages. Start off strong, then forget if I added the salt or not.

[Alex]: [laughs] Exactly! These “long-horizon” tasks force agents to juggle multiple sources, keep tabs on changing requirements, and remember what’s relevant at each step. That’s where things fall apart.

[Jamie]: Any examples? Like, how does an agent actually mess up mid-task?

[Alex]: Sure. One agent was piecing together a purchase order from team chats and spreadsheets. Mid-task, a late approval arrives. The agent notices it, files it as an exception, but then keeps working from its old, out-of-date spreadsheet. It basically verifies a document that’s missing the new approval—so it “checks its work,” but the work is wrong.

[Jamie]: So, they’re good at following instructions—until the instructions change, and then it’s “la la la, I can’t hear you!” [PAUSE] Oh, and I bet the more steps, the more chances to go off the rails.

[Alex]: Bingo. And the data backs that up. For tasks under 45 minutes, the top agents do pretty well. As the workflow gets longer, completion rates nosedive. By the time you hit two or three hours, nobody’s finishing.

[Jamie]: And what about those partial scores? I saw some vendors bragging about much higher numbers.

[Alex]: Ah, yes—the great “partial credit” debate. OSWorld 2.0 gives two scores: a binary completion rate—did you finish the whole job?—and a partial score, which tracks progress at checkpoints along the way. Vendors tend to quote the higher partial score, which sounds better—like, “Look, Mom, I got a 62!” But the reality is, nobody’s actually finishing more than about 20 percent of these workflows.

[Jamie]: So, it’s like saying, “I cleaned 60 percent of my room!” but the dirty laundry is still on the floor and you never made the bed.

[Alex]: [chuckles] Pretty much. The partial scores are legit—agents do make meaningful progress—but a lot of them stall out before the finish line.

[Jamie]: Any other surprises on the leaderboard? Is anybody crushing it with a secret model?

[Alex]: As of late July, the best verified result is still Claude Opus 4.8 at that 20.6 percent. There are some new models—like GPT-5.6 and Claude Opus 5—that claim higher partial scores, but their “full completion” numbers aren’t independently verified yet. And, fun fact, some models do better when allowed to batch their tool calls, but others, like GPT-5.5, don’t gain anything from having more steps or tokens. Super efficient, but it doesn’t translate to more completions.

[Jamie]: So, we’re not looking at a sudden leap in AI superpowers—just some clever number games and maybe a bit of marketing magic.

[Alex]: You got it. If you’re building or deploying long-running AI agents, OSWorld 2.0 is a reality check. The hard part isn’t clicking around—it’s keeping track of the bigger picture as tasks stretch on. The memory bottleneck is real.

[Jamie]: So, for now, if your AI assistant starts a big project, maybe don’t go too far from the “undo” button?

[Alex]: [laughs] Exactly. Or, at the very least, expect to help them over the finish line.

[Jamie]: Well, that’s a wrap for today’s deep dive into OSWorld 2.0. Huge thanks for tuning in to Nerd Level Tech AI Cast! If your agent finishes more than 20 percent of its to-do list, you’re officially ahead of the curve.

[Alex]: And if you enjoyed the episode—or have your own tales of AI state management gone wrong—drop us a note or share the podcast with your favorite tech friends. Until next time, keep your agents focused and your workflows manageable!

[Jamie]: See ya, nerds! [OUTRO MUSIC FADES OUT]