DeliveryBench: Can Agents Earn Profit in Real World?
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Lingjun, Ren, Jiawei, Zhou, Kun, Chen, Jixuan, Ma, Ziqiao, Qin, Lianhui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning
by: Kang, Haoqiang, et al.
Published: (2026)
by: Kang, Haoqiang, et al.
Published: (2026)
C-World: A Computer Use Agent Environment Creator
by: Xi, Ziqiao, et al.
Published: (2026)
by: Xi, Ziqiao, et al.
Published: (2026)
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
by: Ren, Jiawei, et al.
Published: (2025)
by: Ren, Jiawei, et al.
Published: (2025)
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration
by: Zhuang, Yan, et al.
Published: (2025)
by: Zhuang, Yan, et al.
Published: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
Evaluating the Search Agent in a Parallel World
by: Chen, Jiawei, et al.
Published: (2026)
by: Chen, Jiawei, et al.
Published: (2026)
CocoaBench: Evaluating Unified Digital Agents in the Wild
by: CocoaBench Team, et al.
Published: (2026)
by: CocoaBench Team, et al.
Published: (2026)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
by: Liu, Ruoqi, et al.
Published: (2026)
by: Liu, Ruoqi, et al.
Published: (2026)
PrivacyReasoner: Can LLM Emulate a Human-like Privacy Mind?
by: Tu, Yiwen, et al.
Published: (2026)
by: Tu, Yiwen, et al.
Published: (2026)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
by: Song, Zhiheng, et al.
Published: (2026)
by: Song, Zhiheng, et al.
Published: (2026)
MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
by: Guo, Zikang, et al.
Published: (2025)
by: Guo, Zikang, et al.
Published: (2025)
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
by: Zhang, Qiaohong, et al.
Published: (2026)
by: Zhang, Qiaohong, et al.
Published: (2026)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
by: Lu, Jiaxuan, et al.
Published: (2026)
by: Lu, Jiaxuan, et al.
Published: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
EduAgent: Generative Student Agents in Learning
by: Xu, Songlin, et al.
Published: (2024)
by: Xu, Songlin, et al.
Published: (2024)
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
by: Han, Kevin, et al.
Published: (2026)
by: Han, Kevin, et al.
Published: (2026)
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
by: Workman, Kenny, et al.
Published: (2025)
by: Workman, Kenny, et al.
Published: (2025)
GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
by: Yang, Jingqi, et al.
Published: (2025)
by: Yang, Jingqi, et al.
Published: (2025)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
by: Cui, Fan, et al.
Published: (2026)
by: Cui, Fan, et al.
Published: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
by: Yang, Tianzhuo, et al.
Published: (2026)
by: Yang, Tianzhuo, et al.
Published: (2026)
LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
by: Li, Keyu, et al.
Published: (2026)
by: Li, Keyu, et al.
Published: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
by: Han, Tingxu, et al.
Published: (2026)
by: Han, Tingxu, et al.
Published: (2026)
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
by: Pereira, Kristen, et al.
Published: (2026)
by: Pereira, Kristen, et al.
Published: (2026)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)
by: Zhang, Zehua, et al.
Published: (2025)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
by: Men, Tianyi, et al.
Published: (2025)
by: Men, Tianyi, et al.
Published: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment
by: Zheng, Jingnan, et al.
Published: (2026)
by: Zheng, Jingnan, et al.
Published: (2026)
Towards Evaluation for Real-World LLM Unlearning
by: Miao, Ke, et al.
Published: (2025)
by: Miao, Ke, et al.
Published: (2025)
Smart Ride and Delivery Services with Electric Vehicles: Leveraging Bidirectional Charging for Profit Optimisation
by: Du, Jinchun, et al.
Published: (2025)
by: Du, Jinchun, et al.
Published: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
Similar Items
-
SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning
by: Kang, Haoqiang, et al.
Published: (2026) -
C-World: A Computer Use Agent Environment Creator
by: Xi, Ziqiao, et al.
Published: (2026) -
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
by: Ren, Jiawei, et al.
Published: (2025) -
SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration
by: Zhuang, Yan, et al.
Published: (2025) -
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)