POINTS-GUI-G: GUI-Grounding Journey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Zhongyin, Liu, Yuan, Liu, Yikun, Wang, Haicheng, Tian, Le, Zhou, Xiao, You, Yangxiu, Yu, Zilin, Yu, Yang, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
von: Liu, Yuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuan, et al.
Veröffentlicht: (2025)
POINTS: Improving Your Vision-language Model with Affordable Strategies
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
von: Liu, Tao, et al.
Veröffentlicht: (2025)
von: Liu, Tao, et al.
Veröffentlicht: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
On the Robustness of GUI Grounding Models Against Image Attacks
von: Zhao, Haoren, et al.
Veröffentlicht: (2025)
von: Zhao, Haoren, et al.
Veröffentlicht: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
Continual GUI Agents
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
Step-GUI Technical Report
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2025)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2025)
HiconAgent: History Context-aware Policy Optimization for GUI Agents
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
von: Zhou, Xurui, et al.
Veröffentlicht: (2025)
What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs
von: Lin, Jiaping, et al.
Veröffentlicht: (2026)
von: Lin, Jiaping, et al.
Veröffentlicht: (2026)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
MVP: Multiple View Prediction Improves GUI Grounding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026) -
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
von: Liu, Yuan, et al.
Veröffentlicht: (2025) -
POINTS: Improving Your Vision-language Model with Affordable Strategies
von: Liu, Yuan, et al.
Veröffentlicht: (2024) -
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
von: Liu, Yikun, et al.
Veröffentlicht: (2026) -
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
von: Liu, Yikun, et al.
Veröffentlicht: (2026)