Visual Test-time Scaling for GUI Agent Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Tiange, Logeswaran, Lajanugen, Johnson, Justin, Lee, Honglak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025)
by: Jang, Yunseok, et al.
Published: (2025)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
View Selection for 3D Captioning via Diffusion Ranking
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
by: Xiong, Weimin, et al.
Published: (2026)
by: Xiong, Weimin, et al.
Published: (2026)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
by: Wang, Wenkai, et al.
Published: (2026)
by: Wang, Wenkai, et al.
Published: (2026)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Subtask-Aware Visual Reward Learning from Segmented Demonstrations
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
by: Tang, Fei, et al.
Published: (2026)
by: Tang, Fei, et al.
Published: (2026)
CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image Manipulation
by: Xu, Sihan, et al.
Published: (2023)
by: Xu, Sihan, et al.
Published: (2023)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
by: Tang, Jiaqi, et al.
Published: (2025)
by: Tang, Jiaqi, et al.
Published: (2025)
Efficient Long-Horizon GUI Agents via Training-Free KV Cache Compression
by: Zhou, Bowen, et al.
Published: (2026)
by: Zhou, Bowen, et al.
Published: (2026)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
by: Lin, Zichuan, et al.
Published: (2026)
by: Lin, Zichuan, et al.
Published: (2026)
OmniParser for Pure Vision Based GUI Agent
by: Lu, Yadong, et al.
Published: (2024)
by: Lu, Yadong, et al.
Published: (2024)
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents
by: Singh, Kunal, et al.
Published: (2025)
by: Singh, Kunal, et al.
Published: (2025)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
A self-supervised framework for learning whole slide representations
by: Hou, Xinhai, et al.
Published: (2024)
by: Hou, Xinhai, et al.
Published: (2024)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
Auto-scaling Continuous Memory for GUI Agent
by: Wu, Wenyi, et al.
Published: (2025)
by: Wu, Wenyi, et al.
Published: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
by: Zhou, Yuqi, et al.
Published: (2025)
by: Zhou, Yuqi, et al.
Published: (2025)
ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
Less is More: Empowering GUI Agent with Context-Aware Simplification
by: Chen, Gongwei, et al.
Published: (2025)
by: Chen, Gongwei, et al.
Published: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
by: Yoon, Hyungjun, et al.
Published: (2024)
by: Yoon, Hyungjun, et al.
Published: (2024)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
by: Du, Zilin, et al.
Published: (2024)
by: Du, Zilin, et al.
Published: (2024)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025)
by: Zhang, Miaosen, et al.
Published: (2025)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
by: Gou, Boyu, et al.
Published: (2024)
by: Gou, Boyu, et al.
Published: (2024)
Similar Items
-
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025) -
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025) -
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024) -
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024) -
View Selection for 3D Captioning via Diffusion Ranking
by: Luo, Tiange, et al.
Published: (2024)