Claude Opus 5 Benchmarks: What the Numbers Actually Show
I went looking for the actual numbers behind those phrases. What I found says as much about how model launches get reported as it does about the model. The Numbers That Are Actually Comparable The cleanest row is SWE-bench Pro, which runs a model against real GitHub issues and checks whether the patch passes the repository's own tests.
- ▪I went looking for the actual numbers behind those phrases.
- ▪What I found says as much about how model launches get reported as it does about the model.
- ▪The Numbers That Are Actually Comparable The cleanest row is SWE-bench Pro, which runs a model against real GitHub issues and checks whether the patch passes the repository's own tests.
2 outlets in our directory ran this story, first to last over 25 hours. All of the coverage we found sits in one bucket: centre. That one-sidedness is itself worth noticing.
- ▪ Show HN: AI Toolbox supports Claude Opus 5 — AI Toolbox
DEV.to (Top) files mainly under programming. We currently carry 4,892 of its stories.
Opening excerpt (first ~120 words) tap to expand
try { if(localStorage) { let currentUser = localStorage.getItem('current_user'); if (currentUser) { currentUser = JSON.parse(currentUser); if (currentUser.id === 3848289) { document.getElementById('article-show-container').classList.add('current-user-is-article-author'); } } } } catch (e) { console.error(e); } RAXXO Studios Posted on Jul 26 • Originally published at raxxo.shop Claude Opus 5 Benchmarks: What the Numbers Actually Show #ai #productivity #claudecode #automation Opus 5 posts 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10 point jump with no change in per-token price Anthropic published most gains as ratios (three times ARC-AGI-3, more than double Frontier-Bench) rather than absolute scores On CursorBench 3.2 at max effort it lands within 0.5 percent of Fable 5's…
Excerpt limited to ~120 words for fair-use compliance. The full article is at DEV Community.