Gespeichert in:
| Hauptverfasser: | Zhao, Haoren, Chen, Tianyi, Wang, Zhen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.04716 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
POINTS-GUI-G: GUI-Grounding Journey
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
von: Xie, Peng, et al.
Veröffentlicht: (2024)
von: Xie, Peng, et al.
Veröffentlicht: (2024)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation
von: Chu, Tianyi, et al.
Veröffentlicht: (2024)
von: Chu, Tianyi, et al.
Veröffentlicht: (2024)
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Robust Identity Perceptual Watermark Against Deepfake Face Swapping
von: Wang, Tianyi, et al.
Veröffentlicht: (2023)
von: Wang, Tianyi, et al.
Veröffentlicht: (2023)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
Robust Spiking Neural Networks Against Adversarial Attacks
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning
von: Xu, Hai-Ming, et al.
Veröffentlicht: (2024)
von: Xu, Hai-Ming, et al.
Veröffentlicht: (2024)
Unified Prompt Attack Against Text-to-Image Generation Models
von: Peng, Duo, et al.
Veröffentlicht: (2025)
von: Peng, Duo, et al.
Veröffentlicht: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
MVP: Multiple View Prediction Improves GUI Grounding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
SIFT-Graph: Benchmarking Multimodal Defense Against Image Adversarial Attacks With Robust Feature Graph
von: He, Jingjie, et al.
Veröffentlicht: (2025)
von: He, Jingjie, et al.
Veröffentlicht: (2025)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
ResGuard: Enhancing Robustness Against Known Original Attacks in Deep Watermarking
von: Wang, Hanyi, et al.
Veröffentlicht: (2026)
von: Wang, Hanyi, et al.
Veröffentlicht: (2026)
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
von: Zhao, Ruisi, et al.
Veröffentlicht: (2026)
von: Zhao, Ruisi, et al.
Veröffentlicht: (2026)
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
Towards Unified Robustness Against Both Backdoor and Adversarial Attacks
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
Reproducibility Study on Adversarial Attacks Against Robust Transformer Trackers
von: Nokabadi, Fatemeh Nourilenjan, et al.
Veröffentlicht: (2024)
von: Nokabadi, Fatemeh Nourilenjan, et al.
Veröffentlicht: (2024)
Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency Prior
von: Yu, Yi, et al.
Veröffentlicht: (2024)
von: Yu, Yi, et al.
Veröffentlicht: (2024)
Membership Inference Attack Against Masked Image Modeling
von: Li, Zheng, et al.
Veröffentlicht: (2024)
von: Li, Zheng, et al.
Veröffentlicht: (2024)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
Towards Robust Content Watermarking Against Removal and Forgery Attacks
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
von: Zhao, Haoren, et al.
Veröffentlicht: (2026) -
POINTS-GUI-G: GUI-Grounding Journey
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026) -
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025) -
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
von: Li, Zhecheng, et al.
Veröffentlicht: (2025) -
Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
von: Xie, Peng, et al.
Veröffentlicht: (2024)