
Agents Don't Fail on Intelligence, They Fail on Execution
A recent benchmark report reveals that the main issue with agentic AI is not intelligence but execution. The study found that many models fail due to high rates of malformed outputs, leading to increased costs and latency. The report introduces the concept of the Agent Execution Tax, highlighting the importance of reliability in agent systems.
- ▪The benchmark involved 720 browser automation tasks across four large language models (LLMs).
- ▪The worst-performing model had an execution tax of 22.9%, resulting in significant wasted inference costs.
- ▪Reliability in structured output is more critical than raw reasoning ability for successful agent performance.
Hacker News (Newest) files mainly under programming. We currently carry 5,306 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | Fireworks AI |
| Canonical URL | https://fireworks.ai/blog/agent-execution-tax |
| Publication time | Thu, 21 May 2026 14:44:55 +0000 |
| Retrieval time | 2026-05-21T14:51:11.084Z |
| Last seen | 2026-05-21T14:51:11.084Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | rBMO08wk0SB4 |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
What 720 browser agent runs revealed about the real bottleneck in agentic AI.A Notte × Fireworks AI benchmark report.Foundation models keep getting smarter. They ace reasoning benchmarks, write fluent code, and pass professional exams. Yet when you put them inside an agent loop, where they must observe a webpage, decide what to do, and output a structured action ten times in a row, they fail roughly half the time.We ran 720 browser automation tasks across four LLMs to find out why. The answer was not intelligence. It was execution: one model wasted nearly 1 in 5 LLM calls on malformed JSON that had to be retried.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Fireworks AI.