PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Jingxuan, Bai, Xi, Liu, Shan, Jia, Caijun, Sun, Zheng, Xu, Xinglong, Li, Siyuan, Sun, Linzhuang, Yu, Bihui, He, Conghui, Tan, Cheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
by: Zhang, Xiangxiang, et al.
Published: (2026)
by: Zhang, Xiangxiang, et al.
Published: (2026)
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
by: Yu, Bihui, et al.
Published: (2026)
by: Yu, Bihui, et al.
Published: (2026)
Thinking with Drafting: Optical Decompression via Logical Reconstruction
by: Wei, Jingxuan, et al.
Published: (2026)
by: Wei, Jingxuan, et al.
Published: (2026)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
by: Sun, Linzhuang, et al.
Published: (2024)
by: Sun, Linzhuang, et al.
Published: (2024)
Synth-Empathy: Towards High-Quality Synthetic Empathy Data
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
by: Wei, Jingxuan, et al.
Published: (2024)
by: Wei, Jingxuan, et al.
Published: (2024)
Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training
by: Tan, Cheng, et al.
Published: (2023)
by: Tan, Cheng, et al.
Published: (2023)
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
by: Tan, Cheng, et al.
Published: (2024)
by: Tan, Cheng, et al.
Published: (2024)
From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing
by: Wei, Jingxuan, et al.
Published: (2024)
by: Wei, Jingxuan, et al.
Published: (2024)
ChartReasoner: Code-Driven Modality Bridging for Long-Chain Reasoning in Chart Question Answering
by: Jia, Caijun, et al.
Published: (2025)
by: Jia, Caijun, et al.
Published: (2025)
Rational Sensibility: LLM Enhanced Empathetic Response Generation Guided by Self-presentation Theory
by: Sun, Linzhuang, et al.
Published: (2023)
by: Sun, Linzhuang, et al.
Published: (2023)
Brain-inspired Computing Based on Deep Learning for Human-computer Interaction: A Review
by: Yu, Bihui, et al.
Published: (2023)
by: Yu, Bihui, et al.
Published: (2023)
Efficient-Empathy: Towards Efficient and Effective Selection of Empathy Data
by: Sun, Linzhuang, et al.
Published: (2024)
by: Sun, Linzhuang, et al.
Published: (2024)
SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis
by: Wang, Xuan, et al.
Published: (2026)
by: Wang, Xuan, et al.
Published: (2026)
Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
by: Li, Tingyu, et al.
Published: (2025)
by: Li, Tingyu, et al.
Published: (2025)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
by: Sun, Linzhuang, et al.
Published: (2025)
by: Sun, Linzhuang, et al.
Published: (2025)
Canvas-of-Thought: Grounding Reasoning via Mutable Structured States
by: Sun, Lingzhuang, et al.
Published: (2026)
by: Sun, Lingzhuang, et al.
Published: (2026)
TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents
by: Yu, Bihui, et al.
Published: (2026)
by: Yu, Bihui, et al.
Published: (2026)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
by: Pan, Chenkai, et al.
Published: (2026)
by: Pan, Chenkai, et al.
Published: (2026)
GenProve: Learning to Generate Text with Fine-Grained Provenance
by: Wei, Jingxuan, et al.
Published: (2026)
by: Wei, Jingxuan, et al.
Published: (2026)
Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration
by: Sun, Linzhuang, et al.
Published: (2025)
by: Sun, Linzhuang, et al.
Published: (2025)
The Trinity of Consistency as a Defining Principle for General World Models
by: Wei, Jingxuan, et al.
Published: (2026)
by: Wei, Jingxuan, et al.
Published: (2026)
A Survey on Image-text Multimodal Models
by: Guo, Ruifeng, et al.
Published: (2023)
by: Guo, Ruifeng, et al.
Published: (2023)
Towards Bridging the Cross-modal Semantic Gap for Multi-modal Recommendation
by: Wu, Xinglong, et al.
Published: (2024)
by: Wu, Xinglong, et al.
Published: (2024)
PAGER: A Framework for Failure Analysis of Deep Regression Models
by: Thiagarajan, Jayaraman J., et al.
Published: (2023)
by: Thiagarajan, Jayaraman J., et al.
Published: (2023)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
SecAgent: Efficient Mobile GUI Agent with Semantic Context
by: Xie, Yiping, et al.
Published: (2026)
by: Xie, Yiping, et al.
Published: (2026)
Bridging Granularity Gaps: Hierarchical Semantic Learning for Cross-domain Few-shot Segmentation
by: Sun, Sujun, et al.
Published: (2025)
by: Sun, Sujun, et al.
Published: (2025)
SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
by: Tan, Cheng, et al.
Published: (2025)
by: Tan, Cheng, et al.
Published: (2025)
Bridging the Programming Language Gap: Constructing a Multilingual Shared Semantic Space through AST Unification and Graph Matching
by: Chen, Junhao, et al.
Published: (2026)
by: Chen, Junhao, et al.
Published: (2026)
Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics
by: Hoss, Jonathan, et al.
Published: (2026)
by: Hoss, Jonathan, et al.
Published: (2026)
Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment
by: Sun, Yuchen, et al.
Published: (2026)
by: Sun, Yuchen, et al.
Published: (2026)
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
by: Shi, Chenrui, et al.
Published: (2025)
by: Shi, Chenrui, et al.
Published: (2025)
Diffusion Features to Bridge Domain Gap for Semantic Segmentation
by: Ji, Yuxiang, et al.
Published: (2024)
by: Ji, Yuxiang, et al.
Published: (2024)
Training-Free Point Cloud Recognition Based on Geometric and Semantic Information Fusion
by: Chen, Yan, et al.
Published: (2024)
by: Chen, Yan, et al.
Published: (2024)
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
by: Sun, Libo, et al.
Published: (2026)
by: Sun, Libo, et al.
Published: (2026)
Bridging Domain Gap of Point Cloud Representations via Self-Supervised Geometric Augmentation
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
Similar Items
-
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
by: Wei, Jingxuan, et al.
Published: (2025) -
Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
by: Wei, Jingxuan, et al.
Published: (2025) -
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
by: Zhang, Xiangxiang, et al.
Published: (2026) -
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
by: Yu, Bihui, et al.
Published: (2026) -
Thinking with Drafting: Optical Decompression via Logical Reconstruction
by: Wei, Jingxuan, et al.
Published: (2026)