What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Bu, Wendong, Wu, Yang, Yu, Qifan, Gao, Minghe, Miao, Bingchen, Zhang, Zhenkui, Pan, Kaihang, Li, Yunfei, Li, Mengze, Ji, Wei, Li, Juncheng, Tang, Siliang, Zhuang, Yueting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
by: Gao, Minghe, et al.
Published: (2024)
by: Gao, Minghe, et al.
Published: (2024)
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
by: Miao, Bingchen, et al.
Published: (2025)
by: Miao, Bingchen, et al.
Published: (2025)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
by: Wang, Keyu, et al.
Published: (2026)
by: Wang, Keyu, et al.
Published: (2026)
Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
by: Chen, Weile, et al.
Published: (2026)
by: Chen, Weile, et al.
Published: (2026)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
OmniBench: Towards The Future of Universal Omni-Language Models
by: Li, Yizhi, et al.
Published: (2024)
by: Li, Yizhi, et al.
Published: (2024)
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
by: Qiu, Haiyi, et al.
Published: (2026)
by: Qiu, Haiyi, et al.
Published: (2026)
OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
by: Bu, Wendong, et al.
Published: (2025)
by: Bu, Wendong, et al.
Published: (2025)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
by: Qiu, Haiyi, et al.
Published: (2024)
by: Qiu, Haiyi, et al.
Published: (2024)
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
by: Miao, Bingchen, et al.
Published: (2024)
by: Miao, Bingchen, et al.
Published: (2024)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
OmniBench-RAG: A Multi-Domain Evaluation Platform for Retrieval-Augmented Generation Tools
by: Liang, Jiaxuan, et al.
Published: (2025)
by: Liang, Jiaxuan, et al.
Published: (2025)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
Auto-Encoding Morph-Tokens for Multimodal LLM
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
by: Yu, Qifan, et al.
Published: (2023)
by: Yu, Qifan, et al.
Published: (2023)
Bridging Local Details and Global Context in Text-Attributed Graphs
by: Wang, Yaoke, et al.
Published: (2024)
by: Wang, Yaoke, et al.
Published: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
by: Gao, Minghe, et al.
Published: (2023)
by: Gao, Minghe, et al.
Published: (2023)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
by: Gao, Minghe, et al.
Published: (2025)
by: Gao, Minghe, et al.
Published: (2025)
Fact :Teaching MLLMs with Faithful, Concise and Transferable Rationales
by: Gao, Minghe, et al.
Published: (2024)
by: Gao, Minghe, et al.
Published: (2024)
WorldGPT: Empowering LLM as Multimodal World Model
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models
by: Dai, Guangyu, et al.
Published: (2025)
by: Dai, Guangyu, et al.
Published: (2025)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
by: De Brouwer, Edward, et al.
Published: (2026)
by: De Brouwer, Edward, et al.
Published: (2026)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
by: Yin, Aoxiong, et al.
Published: (2024)
by: Yin, Aoxiong, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
OmniVTON: Training-Free Universal Virtual Try-On
by: Yang, Zhaotong, et al.
Published: (2025)
by: Yang, Zhaotong, et al.
Published: (2025)
Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
by: Chen, Xiangnan, et al.
Published: (2025)
by: Chen, Xiangnan, et al.
Published: (2025)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
by: Li, Bingchen, et al.
Published: (2024)
by: Li, Bingchen, et al.
Published: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
by: Dai, Guangyu, et al.
Published: (2025)
by: Dai, Guangyu, et al.
Published: (2025)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
by: Qian, Long, et al.
Published: (2024)
by: Qian, Long, et al.
Published: (2024)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025)
by: Jiang, Yixing, et al.
Published: (2025)
Similar Items
-
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
by: Gao, Minghe, et al.
Published: (2024) -
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
by: Miao, Bingchen, et al.
Published: (2025) -
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
by: Wang, Keyu, et al.
Published: (2026) -
Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
by: Chen, Weile, et al.
Published: (2026) -
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
by: Pan, Kaihang, et al.
Published: (2025)