Saved in:
| Main Authors: | Agashe, Saaket, Wong, Kyle, Tu, Vincent, Yang, Jiachen, Li, Ang, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.00906 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agent S: An Open Agentic Framework that Uses Computers Like a Human
by: Agashe, Saaket, et al.
Published: (2024)
by: Agashe, Saaket, et al.
Published: (2024)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
On the Reliability of Computer Use Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent
by: Yang, Bowen, et al.
Published: (2026)
by: Yang, Bowen, et al.
Published: (2026)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
by: Hu, Xueyu, et al.
Published: (2025)
by: Hu, Xueyu, et al.
Published: (2025)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
by: Hu, Siyuan, et al.
Published: (2024)
by: Hu, Siyuan, et al.
Published: (2024)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
by: Zheng, Kaizhi, et al.
Published: (2022)
by: Zheng, Kaizhi, et al.
Published: (2022)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
by: Kapoor, Raghav, et al.
Published: (2024)
by: Kapoor, Raghav, et al.
Published: (2024)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
Self-Resource Allocation in Multi-Agent LLM Systems
by: Amayuelas, Alfonso, et al.
Published: (2025)
by: Amayuelas, Alfonso, et al.
Published: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
by: Hu, Wenbo, et al.
Published: (2026)
by: Hu, Wenbo, et al.
Published: (2026)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
by: Saha, Aranya, et al.
Published: (2025)
by: Saha, Aranya, et al.
Published: (2025)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
by: InternAgent Team, et al.
Published: (2025)
by: InternAgent Team, et al.
Published: (2025)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
by: Li, Bingxuan, et al.
Published: (2025)
by: Li, Bingxuan, et al.
Published: (2025)
Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning
by: Zhang, Yunchao, et al.
Published: (2022)
by: Zhang, Yunchao, et al.
Published: (2022)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
by: LASA Team, et al.
Published: (2025)
by: LASA Team, et al.
Published: (2025)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
by: Luo, Man, et al.
Published: (2025)
by: Luo, Man, et al.
Published: (2025)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
by: Chen, Haoyu, et al.
Published: (2024)
by: Chen, Haoyu, et al.
Published: (2024)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
by: Xi, Jiajun, et al.
Published: (2024)
by: Xi, Jiajun, et al.
Published: (2024)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
by: Jin, Zhuoran, et al.
Published: (2025)
by: Jin, Zhuoran, et al.
Published: (2025)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
by: Buettner, Kyle, et al.
Published: (2025)
by: Buettner, Kyle, et al.
Published: (2025)
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning
by: Gu, Zishan, et al.
Published: (2024)
by: Gu, Zishan, et al.
Published: (2024)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
A Multimodal Automated Interpretability Agent
by: Shaham, Tamar Rott, et al.
Published: (2024)
by: Shaham, Tamar Rott, et al.
Published: (2024)
Large Multimodal Agents: A Survey
by: Xie, Junlin, et al.
Published: (2024)
by: Xie, Junlin, et al.
Published: (2024)
OpenCUA: Open Foundations for Computer-Use Agents
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
by: Liu, Dongrui, et al.
Published: (2026)
by: Liu, Dongrui, et al.
Published: (2026)
Fara-7B: An Efficient Agentic Model for Computer Use
by: Awadallah, Ahmed, et al.
Published: (2025)
by: Awadallah, Ahmed, et al.
Published: (2025)
From Specialist to Generalist: Unlocking SAM's Learning Potential on Unlabeled Medical Images
by: Vu, Vi, et al.
Published: (2026)
by: Vu, Vi, et al.
Published: (2026)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)
by: Zheng, Xinyue, et al.
Published: (2023)
Similar Items
-
Agent S: An Open Agentic Framework that Uses Computers Like a Human
by: Agashe, Saaket, et al.
Published: (2024) -
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025) -
On the Reliability of Computer Use Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2026) -
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023) -
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent
by: Yang, Bowen, et al.
Published: (2026)