DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Hang, Chen, Hongkai, Cai, Yujun, Liu, Chang, Ye, Qingwen, Yang, Ming-Hsuan, Wang, Yiwei |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
par: Cheng, Kanzhi, et autres
Publié: (2024)
par: Cheng, Kanzhi, et autres
Publié: (2024)
Aria-UI: Visual Grounding for GUI Instructions
par: Yang, Yuhao, et autres
Publié: (2024)
par: Yang, Yuhao, et autres
Publié: (2024)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
par: Henry, Felix, et autres
Publié: (2026)
par: Henry, Felix, et autres
Publié: (2026)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
par: Wang, Ziwei, et autres
Publié: (2025)
par: Wang, Ziwei, et autres
Publié: (2025)
History-Aware Reasoning for GUI Agents
par: Wang, Ziwei, et autres
Publié: (2025)
par: Wang, Ziwei, et autres
Publié: (2025)
WinClick: GUI Grounding with Multimodal Large Language Models
par: Hui, Zheng, et autres
Publié: (2025)
par: Hui, Zheng, et autres
Publié: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
par: Liu, Xinyi, et autres
Publié: (2025)
par: Liu, Xinyi, et autres
Publié: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
par: Nong, Songqin, et autres
Publié: (2025)
par: Nong, Songqin, et autres
Publié: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
par: Tang, Fei, et autres
Publié: (2025)
par: Tang, Fei, et autres
Publié: (2025)
ViMo: A Generative Visual GUI World Model for App Agents
par: Luo, Dezhao, et autres
Publié: (2025)
par: Luo, Dezhao, et autres
Publié: (2025)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
par: Garg, Aryan, et autres
Publié: (2025)
par: Garg, Aryan, et autres
Publié: (2025)
GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph
par: Yu, Mingxian, et autres
Publié: (2026)
par: Yu, Mingxian, et autres
Publié: (2026)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
par: Zhou, Shijie, et autres
Publié: (2025)
par: Zhou, Shijie, et autres
Publié: (2025)
GUI Agents: A Survey
par: Nguyen, Dang, et autres
Publié: (2024)
par: Nguyen, Dang, et autres
Publié: (2024)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
par: Kim, Jeonghyeon, et autres
Publié: (2026)
par: Kim, Jeonghyeon, et autres
Publié: (2026)
SMT-Layout: A MaxSMT-based Approach Supporting Real-time Interaction of Real-world GUI Layout
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
par: Wu, Zongru, et autres
Publié: (2025)
par: Wu, Zongru, et autres
Publié: (2025)
Advancing GUI for Generative AI: Charting the Design Space of Human-AI Interactions through Task Creativity and Complexity
par: Ding, Zijian
Publié: (2024)
par: Ding, Zijian
Publié: (2024)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
par: Tang, Liujian, et autres
Publié: (2025)
par: Tang, Liujian, et autres
Publié: (2025)
SpiritSight Agent: Advanced GUI Agent with One Look
par: Huang, Zhiyuan, et autres
Publié: (2025)
par: Huang, Zhiyuan, et autres
Publié: (2025)
Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
par: Liao, Zeyi, et autres
Publié: (2025)
par: Liao, Zeyi, et autres
Publié: (2025)
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
par: Zhang, Shuning, et autres
Publié: (2025)
par: Zhang, Shuning, et autres
Publié: (2025)
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
par: Liu, Guangyi, et autres
Publié: (2025)
par: Liu, Guangyi, et autres
Publié: (2025)
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
par: Gao, Longxi, et autres
Publié: (2024)
par: Gao, Longxi, et autres
Publié: (2024)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
par: Liu, Yuhang, et autres
Publié: (2025)
par: Liu, Yuhang, et autres
Publié: (2025)
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
par: Liu, Zhilin, et autres
Publié: (2026)
par: Liu, Zhilin, et autres
Publié: (2026)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
par: Gebreegziabher, Simret Araya, et autres
Publié: (2026)
par: Gebreegziabher, Simret Araya, et autres
Publié: (2026)
Establishing Heuristics for Improving the Usability of GUI Machine Learning Tools for Novice Users
par: Yamani, Asma, et autres
Publié: (2024)
par: Yamani, Asma, et autres
Publié: (2024)
Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding
par: Tan, Wee Joe, et autres
Publié: (2026)
par: Tan, Wee Joe, et autres
Publié: (2026)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
par: Luo, Run, et autres
Publié: (2025)
par: Luo, Run, et autres
Publié: (2025)
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
par: Liu, Guangyi, et autres
Publié: (2025)
par: Liu, Guangyi, et autres
Publié: (2025)
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation
par: Hao, Pan, et autres
Publié: (2026)
par: Hao, Pan, et autres
Publié: (2026)
Improving Data Quality via Pre-Task Participant Screening in Crowdsourced GUI Experiments
par: Miyama, Takaya, et autres
Publié: (2026)
par: Miyama, Takaya, et autres
Publié: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
par: Jing, Hongyi, et autres
Publié: (2025)
par: Jing, Hongyi, et autres
Publié: (2025)
E-ANT: A Large-Scale Dataset for Efficient Automatic GUI NavigaTion
par: Wang, Ke, et autres
Publié: (2024)
par: Wang, Ke, et autres
Publié: (2024)
GUI Agents with Foundation Models: A Comprehensive Survey
par: Wang, Shuai, et autres
Publié: (2024)
par: Wang, Shuai, et autres
Publié: (2024)
API Agents vs. GUI Agents: Divergence and Convergence
par: Zhang, Chaoyun, et autres
Publié: (2025)
par: Zhang, Chaoyun, et autres
Publié: (2025)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
par: Li, Hongxin, et autres
Publié: (2025)
par: Li, Hongxin, et autres
Publié: (2025)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
par: Cheng, Pengzhou, et autres
Publié: (2025)
par: Cheng, Pengzhou, et autres
Publié: (2025)
TinyClick: Single-Turn Agent for Empowering GUI Automation
par: Pawlowski, Pawel, et autres
Publié: (2024)
par: Pawlowski, Pawel, et autres
Publié: (2024)
Documents similaires
-
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
par: Cheng, Kanzhi, et autres
Publié: (2024) -
Aria-UI: Visual Grounding for GUI Instructions
par: Yang, Yuhao, et autres
Publié: (2024) -
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
par: Henry, Felix, et autres
Publié: (2026) -
MP-GUI: Modality Perception with MLLMs for GUI Understanding
par: Wang, Ziwei, et autres
Publié: (2025) -
History-Aware Reasoning for GUI Agents
par: Wang, Ziwei, et autres
Publié: (2025)