AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Ruilin, Xiong, Shegnwu, Zou, Tianyu, Xiong, Shili, Rong, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot Segmentation
von: Zou, Tianyu, et al.
Veröffentlicht: (2025)
von: Zou, Tianyu, et al.
Veröffentlicht: (2025)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
von: Gong, Litian, et al.
Veröffentlicht: (2025)
von: Gong, Litian, et al.
Veröffentlicht: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
Content-Style Decoupling for Unsupervised Makeup Transfer without Generating Pseudo Ground Truth
von: Sun, Zhaoyang, et al.
Veröffentlicht: (2024)
von: Sun, Zhaoyang, et al.
Veröffentlicht: (2024)
POINTS-GUI-G: GUI-Grounding Journey
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
von: Shakeel, Rozain, et al.
Veröffentlicht: (2026)
von: Shakeel, Rozain, et al.
Veröffentlicht: (2026)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
GroundingBooth: Grounding Text-to-Image Customization
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2024)
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2024)
Continuous Normalizing Flows for Uncertainty-Aware Human Pose Estimation
von: Liu, Shipeng, et al.
Veröffentlicht: (2025)
von: Liu, Shipeng, et al.
Veröffentlicht: (2025)
Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
von: Wang, Wanfu, et al.
Veröffentlicht: (2025)
von: Wang, Wanfu, et al.
Veröffentlicht: (2025)
Instance-Aware Pseudo-Labeling and Class-Focused Contrastive Learning for Weakly Supervised Domain Adaptive Segmentation of Electron Microscopy
von: Xiong, Shan, et al.
Veröffentlicht: (2025)
von: Xiong, Shan, et al.
Veröffentlicht: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
UGround: Towards Unified Visual Grounding with Unrolled Transformers
von: Qian, Rui, et al.
Veröffentlicht: (2025)
von: Qian, Rui, et al.
Veröffentlicht: (2025)
VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
von: Zhou, Beitong, et al.
Veröffentlicht: (2025)
von: Zhou, Beitong, et al.
Veröffentlicht: (2025)
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
On the Robustness of GUI Grounding Models Against Image Attacks
von: Zhao, Haoren, et al.
Veröffentlicht: (2025)
von: Zhao, Haoren, et al.
Veröffentlicht: (2025)
MVP: Multiple View Prediction Improves GUI Grounding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
Visual Test-time Scaling for GUI Agent Grounding
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
von: Zhang, Shilan, et al.
Veröffentlicht: (2025)
von: Zhang, Shilan, et al.
Veröffentlicht: (2025)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
von: Shu, Yan, et al.
Veröffentlicht: (2026)
von: Shu, Yan, et al.
Veröffentlicht: (2026)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
von: Xiong, Tianyu, et al.
Veröffentlicht: (2025)
von: Xiong, Tianyu, et al.
Veröffentlicht: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
UncTrack: Reliable Visual Object Tracking with Uncertainty-Aware Prototype Memory Network
von: Yao, Siyuan, et al.
Veröffentlicht: (2025)
von: Yao, Siyuan, et al.
Veröffentlicht: (2025)
EGM: Efficient Visual Grounding Language Models
von: Zhan, Guanqi, et al.
Veröffentlicht: (2026)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Balancing Conservatism and Aggressiveness: Prototype-Affinity Hybrid Network for Few-Shot Segmentation
von: Zou, Tianyu, et al.
Veröffentlicht: (2025) -
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024) -
AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
von: Gong, Litian, et al.
Veröffentlicht: (2025) -
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025) -
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)