Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Yue, Ding, Lei, Kuo, Ching-Chen, Jiang, Shan, Zhao, Yang, Guan, Xinze, Yang, Jie, Zhang, Yi, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
by: Yan, Qianqi, et al.
Published: (2025)
by: Yan, Qianqi, et al.
Published: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026)
by: Zhao, Zhongyin, et al.
Published: (2026)
Bodega: Serving Linearizable Reads Locally from Anywhere at Anytime via Roster Leases
by: Hu, Guanzhou, et al.
Published: (2025)
by: Hu, Guanzhou, et al.
Published: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents
by: Karmanov, Ilia, et al.
Published: (2025)
by: Karmanov, Ilia, et al.
Published: (2025)
Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
Transferring to Real-World Layouts: A Depth-aware Framework for Scene Adaptation
by: Chen, Mu, et al.
Published: (2023)
by: Chen, Mu, et al.
Published: (2023)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
by: Wang, Tianxu, et al.
Published: (2025)
by: Wang, Tianxu, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Supporting Students' Reading and Cognition with AI
by: Fu, Yue, et al.
Published: (2025)
by: Fu, Yue, et al.
Published: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
by: Li, Aiden Yiliu, et al.
Published: (2025)
by: Li, Aiden Yiliu, et al.
Published: (2025)
ROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction
by: Xie, Tingwei, et al.
Published: (2026)
by: Xie, Tingwei, et al.
Published: (2026)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
SMT-Layout: A MaxSMT-based Approach Supporting Real-time Interaction of Real-world GUI Layout
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
by: Garg, Aryan, et al.
Published: (2025)
by: Garg, Aryan, et al.
Published: (2025)
Basic Reading Distillation
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
by: Cao, Yue, et al.
Published: (2025)
by: Cao, Yue, et al.
Published: (2025)
Screen Design for Children's Reading: Some Key Issues.
by: Walker, Sue, et al.
Published: (2000)
by: Walker, Sue, et al.
Published: (2000)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
by: Jiang, Zhiyuan, et al.
Published: (2025)
by: Jiang, Zhiyuan, et al.
Published: (2025)
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
by: Li, Kaixin, et al.
Published: (2025)
by: Li, Kaixin, et al.
Published: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
by: Yan, Qianqi, et al.
Published: (2025)
by: Yan, Qianqi, et al.
Published: (2025)
Effects of Digital Reading With On‐Screen Distractions: An Eye‐Tracking Study
by: Angelica Ronconi, et al.
Published: (2024)
by: Angelica Ronconi, et al.
Published: (2024)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
by: Wang, Xuehui, et al.
Published: (2025)
by: Wang, Xuehui, et al.
Published: (2025)
InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Equipping Transformer with Random-Access Reading for Long-Context Understanding
by: Yang, Chenghao, et al.
Published: (2024)
by: Yang, Chenghao, et al.
Published: (2024)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
ReLayout: Integrating Relation Reasoning for Content-aware Layout Generation with Multi-modal Large Language Models
by: Tian, Jiaxu, et al.
Published: (2025)
by: Tian, Jiaxu, et al.
Published: (2025)
Similar Items
-
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
by: Yan, Qianqi, et al.
Published: (2025) -
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024) -
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025) -
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025) -
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
by: Zhang, Chong, et al.
Published: (2024)