Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Aiden Yiliu, Yu, Bizhi, Lei, Daoan, Ren, Tianhe, Liu, Shilong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improved GUI Grounding via Iterative Narrowing
von: Nguyen, Anthony
Veröffentlicht: (2024)
von: Nguyen, Anthony
Veröffentlicht: (2024)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
von: Du, Yong, et al.
Veröffentlicht: (2025)
von: Du, Yong, et al.
Veröffentlicht: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
von: Acuna, David, et al.
Veröffentlicht: (2025)
von: Acuna, David, et al.
Veröffentlicht: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
VGR: Visual Grounded Reasoning
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
von: Lian, Shuquan, et al.
Veröffentlicht: (2025)
von: Lian, Shuquan, et al.
Veröffentlicht: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
Exploring Spatial Language Grounding Through Referring Expressions
von: Tumu, Akshar, et al.
Veröffentlicht: (2025)
von: Tumu, Akshar, et al.
Veröffentlicht: (2025)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
von: Tumu, Akshar, et al.
Veröffentlicht: (2025)
von: Tumu, Akshar, et al.
Veröffentlicht: (2025)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
von: Willemsen, Bram, et al.
Veröffentlicht: (2024)
von: Willemsen, Bram, et al.
Veröffentlicht: (2024)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
von: Vaishnav, Mohit, et al.
Veröffentlicht: (2026)
von: Vaishnav, Mohit, et al.
Veröffentlicht: (2026)
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improved GUI Grounding via Iterative Narrowing
von: Nguyen, Anthony
Veröffentlicht: (2024) -
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025) -
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025) -
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025) -
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)