MP-GUI: Modality Perception with MLLMs for GUI Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ziwei, Chen, Weizhi, Yang, Leyang, Zhou, Sheng, Zhao, Shengchu, Zhan, Hanbei, Jin, Jiongchao, Li, Liangcheng, Shao, Zirui, Bu, Jiajun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
by: Gu, Ming, et al.
Published: (2025)
by: Gu, Ming, et al.
Published: (2025)
History-Aware Reasoning for GUI Agents
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
PG-Agent: An Agent Powered by Page Graph
by: Chen, Weizhi, et al.
Published: (2025)
by: Chen, Weizhi, et al.
Published: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
by: Gebreegziabher, Simret Araya, et al.
Published: (2026)
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
by: Liu, Guangyi, et al.
Published: (2025)
by: Liu, Guangyi, et al.
Published: (2025)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
SpiritSight Agent: Advanced GUI Agent with One Look
by: Huang, Zhiyuan, et al.
Published: (2025)
by: Huang, Zhiyuan, et al.
Published: (2025)
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
by: Liu, Guangyi, et al.
Published: (2025)
by: Liu, Guangyi, et al.
Published: (2025)
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
by: Zhang, Shuning, et al.
Published: (2025)
by: Zhang, Shuning, et al.
Published: (2025)
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
by: Gao, Longxi, et al.
Published: (2024)
by: Gao, Longxi, et al.
Published: (2024)
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
by: Cheng, Ziming, et al.
Published: (2025)
by: Cheng, Ziming, et al.
Published: (2025)
WinClick: GUI Grounding with Multimodal Large Language Models
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
by: Jing, Hongyi, et al.
Published: (2025)
by: Jing, Hongyi, et al.
Published: (2025)
Establishing Heuristics for Improving the Usability of GUI Machine Learning Tools for Novice Users
by: Yamani, Asma, et al.
Published: (2024)
by: Yamani, Asma, et al.
Published: (2024)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
by: Li, Chentao, et al.
Published: (2026)
by: Li, Chentao, et al.
Published: (2026)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation
by: Hao, Pan, et al.
Published: (2026)
by: Hao, Pan, et al.
Published: (2026)
Improving Data Quality via Pre-Task Participant Screening in Crowdsourced GUI Experiments
by: Miyama, Takaya, et al.
Published: (2026)
by: Miyama, Takaya, et al.
Published: (2026)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
by: Garg, Aryan, et al.
Published: (2025)
by: Garg, Aryan, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph
by: Yu, Mingxian, et al.
Published: (2026)
by: Yu, Mingxian, et al.
Published: (2026)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Exploring MLLMs Perception of Network Visualization Principles
by: Miller, Jacob, et al.
Published: (2025)
by: Miller, Jacob, et al.
Published: (2025)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
by: Yang, Saelyne, et al.
Published: (2026)
by: Yang, Saelyne, et al.
Published: (2026)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
by: Kim, Jeonghyeon, et al.
Published: (2026)
by: Kim, Jeonghyeon, et al.
Published: (2026)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
TinyClick: Single-Turn Agent for Empowering GUI Automation
by: Pawlowski, Pawel, et al.
Published: (2024)
by: Pawlowski, Pawel, et al.
Published: (2024)
Qualitative Evaluation of LLM-Designed GUI
by: Sawicki, Bartosz, et al.
Published: (2026)
by: Sawicki, Bartosz, et al.
Published: (2026)
Large Language Model-Brained GUI Agents: A Survey
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
Less is More: Empowering GUI Agent with Context-Aware Simplification
by: Chen, Gongwei, et al.
Published: (2025)
by: Chen, Gongwei, et al.
Published: (2025)
Similar Items
-
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
by: Gu, Ming, et al.
Published: (2025) -
History-Aware Reasoning for GUI Agents
by: Wang, Ziwei, et al.
Published: (2025) -
PG-Agent: An Agent Powered by Page Graph
by: Chen, Weizhi, et al.
Published: (2025) -
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025) -
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)