OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Fuente:
arXiv
Saved in:
| Main Authors: | Kapoor, Raghav, Butala, Yash Parag, Russak, Melisa, Koh, Jing Yu, Kamble, Kiran, Alshikh, Waseem, Salakhutdinov, Ruslan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
by: Vasudev, Rakshith, et al.
Published: (2026)
by: Vasudev, Rakshith, et al.
Published: (2026)
Expect the Unexpected: FailSafe Long Context QA for Finance
by: Kamble, Kiran, et al.
Published: (2025)
by: Kamble, Kiran, et al.
Published: (2025)
Writing in the Margins: Better Inference Pattern for Long Context Retrieval
by: Russak, Melisa, et al.
Published: (2024)
by: Russak, Melisa, et al.
Published: (2024)
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
by: Bensal, Shelly, et al.
Published: (2025)
by: Bensal, Shelly, et al.
Published: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
REVEAL -- Reasoning and Evaluation of Visual Evidence through Aligned Language
by: Praharaj, Ipsita, et al.
Published: (2025)
by: Praharaj, Ipsita, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Multi-Agent Computer Use
by: Koh, Jing Yu, et al.
Published: (2026)
by: Koh, Jing Yu, et al.
Published: (2026)
Dissecting Adversarial Robustness of Multimodal LM Agents
by: Wu, Chen Henry, et al.
Published: (2024)
by: Wu, Chen Henry, et al.
Published: (2024)
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
by: Zhao, Zhenyu, et al.
Published: (2026)
by: Zhao, Zhenyu, et al.
Published: (2026)
Neural MP: A Generalist Neural Motion Planner
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
Tree Search for Language Model Agents
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts
by: Wu, Jialin, et al.
Published: (2023)
by: Wu, Jialin, et al.
Published: (2023)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
by: Yu, Haofei, et al.
Published: (2023)
by: Yu, Haofei, et al.
Published: (2023)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
by: Jin, Zhuoran, et al.
Published: (2025)
by: Jin, Zhuoran, et al.
Published: (2025)
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
by: Nandy, Abhilash, et al.
Published: (2023)
by: Nandy, Abhilash, et al.
Published: (2023)
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
by: Zheng, Lianqing, et al.
Published: (2024)
by: Zheng, Lianqing, et al.
Published: (2024)
Pages from the Desktop: Desktop Publishing Today.
by: Crawford, Walt
Published: (1994)
by: Crawford, Walt
Published: (1994)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
Effective Data Augmentation With Diffusion Models
by: Trabucco, Brandon, et al.
Published: (2023)
by: Trabucco, Brandon, et al.
Published: (2023)
Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
by: Liang, Paul Pu, et al.
Published: (2023)
by: Liang, Paul Pu, et al.
Published: (2023)
OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
by: Tang, Tao, et al.
Published: (2025)
by: Tang, Tao, et al.
Published: (2025)
Local Policies Enable Zero-shot Long-horizon Manipulation
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
Understanding Visual Concepts Across Models
by: Trabucco, Brandon, et al.
Published: (2024)
by: Trabucco, Brandon, et al.
Published: (2024)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025)
by: Gunawan, Agus, et al.
Published: (2025)
Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Libraries and Desktop Storage Options: Results of a Web-Based Survey.
by: Hendricks, Arthur, et al.
Published: (2002)
by: Hendricks, Arthur, et al.
Published: (2002)
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
by: Hamamci, Ibrahim Ethem, et al.
Published: (2024)
by: Hamamci, Ibrahim Ethem, et al.
Published: (2024)
Leafy Spurge Dataset: Real-world Weed Classification Within Aerial Drone Imagery
by: Doherty, Kyle, et al.
Published: (2024)
by: Doherty, Kyle, et al.
Published: (2024)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
by: Qu, Yuxiao, et al.
Published: (2026)
by: Qu, Yuxiao, et al.
Published: (2026)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2024)
by: Wang, Shihao, et al.
Published: (2024)
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
Automatic Question-Answer Generation for Long-Tail Knowledge
by: Kumar, Rohan, et al.
Published: (2024)
by: Kumar, Rohan, et al.
Published: (2024)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
by: Soni, Aditya Bharat, et al.
Published: (2025)
by: Soni, Aditya Bharat, et al.
Published: (2025)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
by: Yang, Zhengwei, et al.
Published: (2026)
by: Yang, Zhengwei, et al.
Published: (2026)
Similar Items
-
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
by: Vasudev, Rakshith, et al.
Published: (2026) -
Expect the Unexpected: FailSafe Long Context QA for Finance
by: Kamble, Kiran, et al.
Published: (2025) -
Writing in the Margins: Better Inference Pattern for Long Context Retrieval
by: Russak, Melisa, et al.
Published: (2024) -
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
by: Bensal, Shelly, et al.
Published: (2025) -
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)