ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yuanyi, Huang, Heyuan, Lin, Qiqiang, Zhao, Yin, Qu, Xiangmou, Wang, Jun, Lou, Xingyu, Liu, Weiwen, Zhang, Zhuosheng, Yu, Yong, Zhang, Weinan, Wang, Zhaoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
by: Wang, Jihong, et al.
Published: (2026)
by: Wang, Jihong, et al.
Published: (2026)
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Adaptive Milestone Reward for GUI Agents
by: Zheng, Congmin, et al.
Published: (2026)
by: Zheng, Congmin, et al.
Published: (2026)
MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning
by: Wu, Zheng, et al.
Published: (2026)
by: Wu, Zheng, et al.
Published: (2026)
ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
by: Wu, Fangwen, et al.
Published: (2025)
by: Wu, Fangwen, et al.
Published: (2025)
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
by: Cai, Muzhen, et al.
Published: (2025)
by: Cai, Muzhen, et al.
Published: (2025)
TopoClaw: A Human-Centric and Topology-Aware Agent Operating System
by: Huang, Heyuan, et al.
Published: (2026)
by: Huang, Heyuan, et al.
Published: (2026)
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
by: Liang, Yijun, et al.
Published: (2025)
by: Liang, Yijun, et al.
Published: (2025)
Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation
by: Zhang, Weiming, et al.
Published: (2026)
by: Zhang, Weiming, et al.
Published: (2026)
SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
by: Chai, Huacan, et al.
Published: (2026)
by: Chai, Huacan, et al.
Published: (2026)
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
by: Zhang, Linghua, et al.
Published: (2026)
by: Zhang, Linghua, et al.
Published: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
by: Zhou, Qixing, et al.
Published: (2026)
by: Zhou, Qixing, et al.
Published: (2026)
ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents
by: Guo, Shuhan, et al.
Published: (2026)
by: Guo, Shuhan, et al.
Published: (2026)
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
by: Zhu, Jiachen, et al.
Published: (2026)
by: Zhu, Jiachen, et al.
Published: (2026)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
by: Cheng, Zihao, et al.
Published: (2026)
by: Cheng, Zihao, et al.
Published: (2026)
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
by: Wang, Mingze, et al.
Published: (2025)
by: Wang, Mingze, et al.
Published: (2025)
Voltran: Unlocking Trust and Confidentiality in Decentralized Federated Learning Aggregation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
HMDN: Hierarchical Multi-Distribution Network for Click-Through Rate Prediction
by: Lou, Xingyu, et al.
Published: (2024)
by: Lou, Xingyu, et al.
Published: (2024)
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
by: Chen, Xueyao, et al.
Published: (2026)
by: Chen, Xueyao, et al.
Published: (2026)
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
by: Lin, Zuyao, et al.
Published: (2026)
by: Lin, Zuyao, et al.
Published: (2026)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
Hammer: Robust Function-Calling for On-Device Language Models via Function Masking
by: Lin, Qiqiang, et al.
Published: (2024)
by: Lin, Qiqiang, et al.
Published: (2024)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
by: Diao, Lingxiao, et al.
Published: (2025)
by: Diao, Lingxiao, et al.
Published: (2025)
DIIT: A Domain-Invariant Information Transfer Method for Industrial Cross-Domain Recommendation
by: Huang, Heyuan, et al.
Published: (2024)
by: Huang, Heyuan, et al.
Published: (2024)
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
by: Wang, Weixuan, et al.
Published: (2025)
by: Wang, Weixuan, et al.
Published: (2025)
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
by: Orlanski, Gabriel, et al.
Published: (2026)
by: Orlanski, Gabriel, et al.
Published: (2026)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search
by: Li, Qingyao, et al.
Published: (2026)
by: Li, Qingyao, et al.
Published: (2026)
HorizonBench: Long-Horizon Personalization with Evolving Preferences
by: Li, Shuyue Stella, et al.
Published: (2026)
by: Li, Shuyue Stella, et al.
Published: (2026)
Similar Items
-
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
by: Wu, Zheng, et al.
Published: (2025) -
Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
by: Wu, Zheng, et al.
Published: (2025) -
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
by: Wang, Jihong, et al.
Published: (2026) -
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
by: Li, Ning, et al.
Published: (2025) -
Adaptive Milestone Reward for GUI Agents
by: Zheng, Congmin, et al.
Published: (2026)