Gespeichert in:
| Hauptverfasser: | Agashe, Saaket, Wong, Kyle, Tu, Vincent, Yang, Jiachen, Li, Ang, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.00906 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent S: An Open Agentic Framework that Uses Computers Like a Human
von: Agashe, Saaket, et al.
Veröffentlicht: (2024)
von: Agashe, Saaket, et al.
Veröffentlicht: (2024)
Scaling Agents for Computer Use
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
On the Reliability of Computer Use Agents
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2026)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2026)
An Embodied Generalist Agent in 3D World
von: Huang, Jiangyong, et al.
Veröffentlicht: (2023)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2023)
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent
von: Yang, Bowen, et al.
Veröffentlicht: (2026)
von: Yang, Bowen, et al.
Veröffentlicht: (2026)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
von: Hu, Siyuan, et al.
Veröffentlicht: (2024)
von: Hu, Siyuan, et al.
Veröffentlicht: (2024)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
Self-Resource Allocation in Multi-Agent LLM Systems
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2025)
von: Amayuelas, Alfonso, et al.
Veröffentlicht: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
von: Feng, Weixi, et al.
Veröffentlicht: (2024)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
von: Saha, Aranya, et al.
Veröffentlicht: (2025)
von: Saha, Aranya, et al.
Veröffentlicht: (2025)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
von: InternAgent Team, et al.
Veröffentlicht: (2025)
von: InternAgent Team, et al.
Veröffentlicht: (2025)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning
von: Zhang, Yunchao, et al.
Veröffentlicht: (2022)
von: Zhang, Yunchao, et al.
Veröffentlicht: (2022)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
von: LASA Team, et al.
Veröffentlicht: (2025)
von: LASA Team, et al.
Veröffentlicht: (2025)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
von: Luo, Man, et al.
Veröffentlicht: (2025)
von: Luo, Man, et al.
Veröffentlicht: (2025)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
von: Xi, Jiajun, et al.
Veröffentlicht: (2024)
von: Xi, Jiajun, et al.
Veröffentlicht: (2024)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
von: Jin, Zhuoran, et al.
Veröffentlicht: (2025)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2025)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
von: Buettner, Kyle, et al.
Veröffentlicht: (2025)
von: Buettner, Kyle, et al.
Veröffentlicht: (2025)
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
A Multimodal Automated Interpretability Agent
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024)
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024)
Large Multimodal Agents: A Survey
von: Xie, Junlin, et al.
Veröffentlicht: (2024)
von: Xie, Junlin, et al.
Veröffentlicht: (2024)
OpenCUA: Open Foundations for Computer-Use Agents
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2025)
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
von: Liu, Dongrui, et al.
Veröffentlicht: (2026)
Fara-7B: An Efficient Agentic Model for Computer Use
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
From Specialist to Generalist: Unlocking SAM's Learning Potential on Unlabeled Medical Images
von: Vu, Vi, et al.
Veröffentlicht: (2026)
von: Vu, Vi, et al.
Veröffentlicht: (2026)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
von: Wei, Cong, et al.
Veröffentlicht: (2024)
von: Wei, Cong, et al.
Veröffentlicht: (2024)
MCU: An Evaluation Framework for Open-Ended Game Agents
von: Zheng, Xinyue, et al.
Veröffentlicht: (2023)
von: Zheng, Xinyue, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Agent S: An Open Agentic Framework that Uses Computers Like a Human
von: Agashe, Saaket, et al.
Veröffentlicht: (2024) -
Scaling Agents for Computer Use
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025) -
On the Reliability of Computer Use Agents
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2026) -
An Embodied Generalist Agent in 3D World
von: Huang, Jiangyong, et al.
Veröffentlicht: (2023) -
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent
von: Yang, Bowen, et al.
Veröffentlicht: (2026)