TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Deyang, Huang, Jing, Zhao, Xuanle, Chen, Lei, Zheng, Liming, Liu, Fanfan, Qiu, Haibo, Shi, Peng, Zeng, Zhixiong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
ScaleTrack: Scaling and back-tracking Automated GUI Agents
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
by: Han, Wenkang, et al.
Published: (2025)
by: Han, Wenkang, et al.
Published: (2025)
Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction
by: Cui, Chaoqun, et al.
Published: (2026)
by: Cui, Chaoqun, et al.
Published: (2026)
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)
by: Cao, Yilin, et al.
Published: (2026)
Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
by: Zhong, Yufeng, et al.
Published: (2026)
by: Zhong, Yufeng, et al.
Published: (2026)
Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
by: Liu, Fanfan, et al.
Published: (2026)
by: Liu, Fanfan, et al.
Published: (2026)
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
by: Yang, Longrong, et al.
Published: (2025)
by: Yang, Longrong, et al.
Published: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
by: Chen, Lei, et al.
Published: (2025)
by: Chen, Lei, et al.
Published: (2025)
UItron: Foundational GUI Agent with Advanced Perception and Planning
by: Zeng, Zhixiong, et al.
Published: (2025)
by: Zeng, Zhixiong, et al.
Published: (2025)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025)
by: Yang, Siqi, et al.
Published: (2025)
Context Cascade Compression: Exploring the Upper Limits of Text Compression
by: Liu, Fanfan, et al.
Published: (2025)
by: Liu, Fanfan, et al.
Published: (2025)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
by: Hu, Xuhao, et al.
Published: (2026)
by: Hu, Xuhao, et al.
Published: (2026)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
RETRACTED: Effect of angiogenesis inhibitors on wound healing in patients with ovarian cancer: A meta‐analysis
by: Xin Li, et al.
Published: (2024)
by: Xin Li, et al.
Published: (2024)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025)
by: Chen, Kun, et al.
Published: (2025)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
by: Liu, Zhaoyang, et al.
Published: (2025)
by: Liu, Zhaoyang, et al.
Published: (2025)
On the Stability of Contact Discontinuity for the 1D Compressible Full Navier‐Stokes‐Allen‐Cahn System With Free Boundary
by: Dan Lei, et al.
Published: (2026)
by: Dan Lei, et al.
Published: (2026)
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
by: Lan, Xiaohan, et al.
Published: (2025)
by: Lan, Xiaohan, et al.
Published: (2025)
CUA-Skill: Develop Skills for Computer Using Agent
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026)
by: Xue, Taofeng, et al.
Published: (2026)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
OpenCUA: Open Foundations for Computer-Use Agents
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
AXON: An Automated Netlist Optimization Framework for High-Speed Adders
by: Yang, Tiantian, et al.
Published: (2026)
by: Yang, Tiantian, et al.
Published: (2026)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
by: Liu, Fanfan, et al.
Published: (2024)
by: Liu, Fanfan, et al.
Published: (2024)
Scaling Agentic Verifier for Competitive Coding
by: Ma, Zeyao, et al.
Published: (2026)
by: Ma, Zeyao, et al.
Published: (2026)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
by: Zhao, Xuanle, et al.
Published: (2025)
by: Zhao, Xuanle, et al.
Published: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
by: Yang, Chenyu, et al.
Published: (2025)
by: Yang, Chenyu, et al.
Published: (2025)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
by: Mo, Ying, et al.
Published: (2026)
by: Mo, Ying, et al.
Published: (2026)
Similar Items
-
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025) -
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
by: Zhao, Xuanle, et al.
Published: (2025) -
ScaleTrack: Scaling and back-tracking Automated GUI Agents
by: Huang, Jing, et al.
Published: (2025) -
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
by: Han, Wenkang, et al.
Published: (2025) -
Agentic Reward Modeling: Verifying GUI Agent via Online Proactive Interaction
by: Cui, Chaoqun, et al.
Published: (2026)