AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pei, Siqi, Tang, Liang, Duan, Tiaonan, Chen, Long, Li, Shuxian, Huang, Kaer, Jing, Yanzhe, Yan, Yiqiang, Zhang, Bo, Jiang, Chenghao, Zhang, Borui, Lu, Jiwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025)
Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent
von: Cheng, Yuhao, et al.
Veröffentlicht: (2025)
von: Cheng, Yuhao, et al.
Veröffentlicht: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
Nuanced Emotion Recognition Based on a Segment-based MLLM Framework Leveraging Qwen3-Omni for AH Detection
von: Tang, Liang, et al.
Veröffentlicht: (2026)
von: Tang, Liang, et al.
Veröffentlicht: (2026)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
von: Yu, Xuan, et al.
Veröffentlicht: (2025)
von: Yu, Xuan, et al.
Veröffentlicht: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
von: Ge, Zhiqi, et al.
Veröffentlicht: (2024)
Aria-UI: Visual Grounding for GUI Instructions
von: Yang, Yuhao, et al.
Veröffentlicht: (2024)
von: Yang, Yuhao, et al.
Veröffentlicht: (2024)
POINTS-GUI-G: GUI-Grounding Journey
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
AttZoom: Attention Zoom for Better Visual Features
von: DeAlcala, Daniel, et al.
Veröffentlicht: (2025)
von: DeAlcala, Daniel, et al.
Veröffentlicht: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
Continual GUI Agents
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
von: Wei, Lai, et al.
Veröffentlicht: (2026)
von: Wei, Lai, et al.
Veröffentlicht: (2026)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
FIZZ: Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document
von: Yang, Joonho, et al.
Veröffentlicht: (2024)
von: Yang, Joonho, et al.
Veröffentlicht: (2024)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025)
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025)
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
von: Liu, Yichao, et al.
Veröffentlicht: (2026)
von: Liu, Yichao, et al.
Veröffentlicht: (2026)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
Zoom Conference Dataset
von: Gudaparthi, Hemanth
Veröffentlicht: (2025)
von: Gudaparthi, Hemanth
Veröffentlicht: (2025)
Zooming in on discrete space
von: Vanzella, Daniel A. Turolla
Veröffentlicht: (2024)
von: Vanzella, Daniel A. Turolla
Veröffentlicht: (2024)
Code Semantic Zooming
von: Ba, Jinsheng, et al.
Veröffentlicht: (2025)
von: Ba, Jinsheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025) -
Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements
von: Liu, Ziwei, et al.
Veröffentlicht: (2026) -
Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent
von: Cheng, Yuhao, et al.
Veröffentlicht: (2025) -
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026) -
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)