TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
The article introduces MM-ToolBench, a benchmark designed for evaluating task-oriented omni-modal tool-using agents. It aims to bridge the gap between isolated evaluations of tool use and real-world applications by incorporating closed-loop multimodal verification. Experiments reveal that even advanced models struggle to meet human performance benchmarks, highlighting the challenge of the tasks presented.
- ▪MM-ToolBench includes 100 executable tasks across two macro task families: Customer Service and Intelligent Creation.
- ▪The benchmark is supported by 27 MCP servers and 324 tools, emphasizing its comprehensive design.
- ▪Experiments show that the best-performing coding-agent model achieves only 32.0% task success, significantly lower than the 94.0% success rate of humans.
arXiv cs.AI files mainly under ai research. We currently carry 1,128 of its stories.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Artificial Intelligence arXiv:2605.16909 (cs) [Submitted on 16 May 2026] Title:TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents Authors:Zhiqiang Liu, Wenhui Dong, Yilang Tan, Yuwen Qu, Haochen Yin, Chenyang Si View a PDF of the paper titled TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents, by Zhiqiang Liu and 5 other authors View PDF HTML (experimental) Abstract:Tool-using agents are increasingly expected to operate across realistic professional workflows, where they must interpret multimodal inputs, coordinate external tools, inspect intermediate artifacts, and revise their actions before producing a final result.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv cs.AI.