CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Jian, Xiangru, Nayak, Shravan, Lin, Kevin Qinghong, Feizi, Aarash, Li, Kaixin, Bechard, Patrice, Gella, Spandana, Rajeswar, Sai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
by: Sheikholeslami, Nima, et al.
Published: (2025)
by: Sheikholeslami, Nima, et al.
Published: (2025)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Terminal Agents Suffice for Enterprise Automation
by: Bechard, Patrice, et al.
Published: (2026)
by: Bechard, Patrice, et al.
Published: (2026)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
OpenCUA: Open Foundations for Computer-Use Agents
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
PRO-CUA: Process-Reward Optimization for Computer Use Agents
by: He, Yifei, et al.
Published: (2026)
by: He, Yifei, et al.
Published: (2026)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
by: Maheshwary, Rishabh, et al.
Published: (2025)
by: Maheshwary, Rishabh, et al.
Published: (2025)
Grammar Search for Multi-Agent Systems
by: Singh, Mayank, et al.
Published: (2025)
by: Singh, Mayank, et al.
Published: (2025)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
Reducing hallucination in structured outputs via Retrieval-Augmented Generation
by: Béchard, Patrice, et al.
Published: (2024)
by: Béchard, Patrice, et al.
Published: (2024)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
by: Béchard, Patrice, et al.
Published: (2025)
by: Béchard, Patrice, et al.
Published: (2025)
Generating a Low-code Complete Workflow via Task Decomposition and RAG
by: Ayala, Orlando Marquez, et al.
Published: (2024)
by: Ayala, Orlando Marquez, et al.
Published: (2024)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
by: Liu, Zhaoyang, et al.
Published: (2025)
by: Liu, Zhaoyang, et al.
Published: (2025)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
by: Hu, Xuhao, et al.
Published: (2026)
by: Hu, Xuhao, et al.
Published: (2026)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
CUA-Skill: Develop Skills for Computer Using Agent
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026)
by: Xue, Taofeng, et al.
Published: (2026)
IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use Agents
by: Lee, Seoyoung, et al.
Published: (2026)
by: Lee, Seoyoung, et al.
Published: (2026)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
by: Feizi, Aarash, et al.
Published: (2024)
by: Feizi, Aarash, et al.
Published: (2024)
FairLoRA: Unpacking Bias Mitigation in Vision Models with Fairness-Driven Low-Rank Adaptation
by: Sukumaran, Rohan, et al.
Published: (2024)
by: Sukumaran, Rohan, et al.
Published: (2024)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
by: Malay, Shiva Krishna Reddy, et al.
Published: (2026)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
by: Pang, Wei, et al.
Published: (2025)
by: Pang, Wei, et al.
Published: (2025)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
by: Wang, Suyuchen, et al.
Published: (2025)
by: Wang, Suyuchen, et al.
Published: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining
by: Abro, Aarash, et al.
Published: (2026)
by: Abro, Aarash, et al.
Published: (2026)
Similar Items
-
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025) -
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025) -
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025) -
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025) -
Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
by: Sheikholeslami, Nima, et al.
Published: (2025)