TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhiqiang, Dong, Wenhui, Tan, Yilang, Qu, Yuwen, Yin, Haochen, Si, Chenyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NGM: A Plug-and-Play Training-Free Memory Module for LLMs
von: Qu, Yuwen, et al.
Veröffentlicht: (2026)
von: Qu, Yuwen, et al.
Veröffentlicht: (2026)
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
von: Chu, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Chu, Zhaoyang, et al.
Veröffentlicht: (2026)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
von: Henry, Felix, et al.
Veröffentlicht: (2026)
von: Henry, Felix, et al.
Veröffentlicht: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
von: Cui, Fan, et al.
Veröffentlicht: (2026)
von: Cui, Fan, et al.
Veröffentlicht: (2026)
NEWSAGENT: Benchmarking Multimodal Agents as Journalists with Real-World Newswriting Tasks
von: Chien, Yen-Che, et al.
Veröffentlicht: (2025)
von: Chien, Yen-Che, et al.
Veröffentlicht: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
von: Chen, Zixin, et al.
Veröffentlicht: (2026)
von: Chen, Zixin, et al.
Veröffentlicht: (2026)
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
OmniGAIA: Towards Native Omni-Modal AI Agents
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
von: Wei, Shaohang, et al.
Veröffentlicht: (2025)
von: Wei, Shaohang, et al.
Veröffentlicht: (2025)
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
von: Chi, Yizhe, et al.
Veröffentlicht: (2026)
von: Chi, Yizhe, et al.
Veröffentlicht: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
von: Long, Xiang, et al.
Veröffentlicht: (2026)
von: Long, Xiang, et al.
Veröffentlicht: (2026)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Lu, Jiaxuan, et al.
Veröffentlicht: (2026)
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
von: Si, Shuzheng, et al.
Veröffentlicht: (2023)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
von: Liu, Yunqi, et al.
Veröffentlicht: (2026)
von: Liu, Yunqi, et al.
Veröffentlicht: (2026)
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
von: Li, Keyu, et al.
Veröffentlicht: (2026)
von: Li, Keyu, et al.
Veröffentlicht: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
Evaluating Privilege Usage of Agents with Real-World Tools
von: Zhang, Quan, et al.
Veröffentlicht: (2026)
von: Zhang, Quan, et al.
Veröffentlicht: (2026)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
von: Lei, Fei, et al.
Veröffentlicht: (2025)
von: Lei, Fei, et al.
Veröffentlicht: (2025)
RealDPO: Real or Not Real, that is the Preference
von: Cheng, Guo, et al.
Veröffentlicht: (2025)
von: Cheng, Guo, et al.
Veröffentlicht: (2025)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation
von: Sawarni, Ayush, et al.
Veröffentlicht: (2026)
von: Sawarni, Ayush, et al.
Veröffentlicht: (2026)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
von: Liu, Yibing, et al.
Veröffentlicht: (2026)
von: Liu, Yibing, et al.
Veröffentlicht: (2026)
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
von: Liu, Guohong, et al.
Veröffentlicht: (2026)
von: Liu, Guohong, et al.
Veröffentlicht: (2026)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
von: Li, Junlong, et al.
Veröffentlicht: (2025)
von: Li, Junlong, et al.
Veröffentlicht: (2025)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
von: Xie, Tianyu, et al.
Veröffentlicht: (2026)
von: Xie, Tianyu, et al.
Veröffentlicht: (2026)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System
von: Liu, Zhiwei, et al.
Veröffentlicht: (2024)
von: Liu, Zhiwei, et al.
Veröffentlicht: (2024)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
An Executable Benchmarking Suite for Tool-Using Agents
von: Zhong, Zhiqing, et al.
Veröffentlicht: (2026)
von: Zhong, Zhiqing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
NGM: A Plug-and-Play Training-Free Memory Module for LLMs
von: Qu, Yuwen, et al.
Veröffentlicht: (2026) -
Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks
von: Cheng, Xiang, et al.
Veröffentlicht: (2025) -
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
von: Bie, Fuqing, et al.
Veröffentlicht: (2025) -
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
von: Chu, Zhaoyang, et al.
Veröffentlicht: (2026) -
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
von: Henry, Felix, et al.
Veröffentlicht: (2026)