Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Tianbao, Deng, Jiaqi, Li, Xiaochuan, Yang, Junlin, Wu, Haoyuan, Chen, Jixuan, Hu, Wenjing, Wang, Xinyuan, Xu, Yuhui, Wang, Zekun, Xu, Yiheng, Wang, Junli, Sahoo, Doyen, Yu, Tao, Xiong, Caiming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
by: Xu, Yiheng, et al.
Published: (2024)
by: Xu, Yiheng, et al.
Published: (2024)
Scalable Chain of Thoughts via Elastic Reasoning
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
by: Lu, Dunjie, et al.
Published: (2025)
by: Lu, Dunjie, et al.
Published: (2025)
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
by: Xu, Yiheng, et al.
Published: (2024)
by: Xu, Yiheng, et al.
Published: (2024)
Entropy-Based Block Pruning for Efficient Large Language Models
by: Yang, Liangwei, et al.
Published: (2025)
by: Yang, Liangwei, et al.
Published: (2025)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
ThinK: Thinner Key Cache by Query-Driven Pruning
by: Xu, Yuhui, et al.
Published: (2024)
by: Xu, Yuhui, et al.
Published: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
by: Zhao, Zirui, et al.
Published: (2024)
by: Zhao, Zirui, et al.
Published: (2024)
XForecast: Evaluating Natural Language Explanations for Time Series Forecasting
by: Aksu, Taha, et al.
Published: (2024)
by: Aksu, Taha, et al.
Published: (2024)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
Reward Models Identify Consistency, Not Causality
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions
by: Zhang, David Junhao, et al.
Published: (2024)
by: Zhang, David Junhao, et al.
Published: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
OpenCUA: Open Foundations for Computer-Use Agents
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
by: Hu, Zhiyuan, et al.
Published: (2025)
by: Hu, Zhiyuan, et al.
Published: (2025)
Direct Judgement Preference Optimization
by: Wang, Peifeng, et al.
Published: (2024)
by: Wang, Peifeng, et al.
Published: (2024)
A Hybrid Attention Framework for Fake News Detection with Large Language Models
by: Xu, Xiaochuan, et al.
Published: (2025)
by: Xu, Xiaochuan, et al.
Published: (2025)
Hierarchical Multi-Stage BERT Fusion Framework with Dual Attention for Enhanced Cyberbullying Detection in Social Media
by: Wang, Jiani, et al.
Published: (2025)
by: Wang, Jiani, et al.
Published: (2025)
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
by: Cao, Ruisheng, et al.
Published: (2024)
by: Cao, Ruisheng, et al.
Published: (2024)
Navigating Transitions: Envisioning Conversational User Interfaces to Support International Students
by: Xu, Yuhui, et al.
Published: (2026)
by: Xu, Yuhui, et al.
Published: (2026)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
by: Zhu, Chiwei, et al.
Published: (2025)
by: Zhu, Chiwei, et al.
Published: (2025)
Lemur: Harmonizing Natural Language and Code for Language Agents
by: Xu, Yiheng, et al.
Published: (2023)
by: Xu, Yiheng, et al.
Published: (2023)
Select, Read, and Write: A Multi-Agent Framework of Full-Text-based Related Work Generation
by: Liu, Xiaochuan, et al.
Published: (2025)
by: Liu, Xiaochuan, et al.
Published: (2025)
Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models
by: Roy, Aryan, et al.
Published: (2026)
by: Roy, Aryan, et al.
Published: (2026)
Can AI Write Classical Chinese Poetry like Humans? An Empirical Study Inspired by Turing Test
by: Deng, Zekun, et al.
Published: (2024)
by: Deng, Zekun, et al.
Published: (2024)
A Novel Method to Metigate Demographic and Expert Bias in ICD Coding with Causal Inference
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
A Novel ICD Coding Method Based on Associated and Hierarchical Code Description Distillation
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
Algorithm and Strategy Construction for Sure-Almost-Sure Stochastic Parity Games
by: Doyen, Laurent, et al.
Published: (2026)
by: Doyen, Laurent, et al.
Published: (2026)
Regular Games with Imperfect Information Are Not That Regular
by: Doyen, Laurent, et al.
Published: (2024)
by: Doyen, Laurent, et al.
Published: (2024)
Similar Items
-
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
by: Xu, Yiheng, et al.
Published: (2024) -
Scalable Chain of Thoughts via Elastic Reasoning
by: Xu, Yuhui, et al.
Published: (2025) -
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
by: Lu, Dunjie, et al.
Published: (2025) -
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
by: Xu, Yiheng, et al.
Published: (2024) -
Entropy-Based Block Pruning for Efficient Large Language Models
by: Yang, Liangwei, et al.
Published: (2025)