Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Zongru, Cheng, Pengzhou, Wu, Zheng, Ju, Tianjie, Zhang, Zhuosheng, Liu, Gongshen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
von: Wu, Zheng, et al.
Veröffentlicht: (2025)
von: Wu, Zheng, et al.
Veröffentlicht: (2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
History-Aware Reasoning for GUI Agents
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
WinClick: GUI Grounding with Multimodal Large Language Models
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
von: Dong, Lingzhong, et al.
Veröffentlicht: (2025)
von: Dong, Lingzhong, et al.
Veröffentlicht: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
A Survey on (M)LLM-Based GUI Agents
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
You Only Look at Screens: Multimodal Chain-of-Action Agents
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
Learning from Implicit User Feedback, Emotions and Demographic Information in Task-Oriented and Document-Grounded Dialogues
von: Petrak, Dominic, et al.
Veröffentlicht: (2024)
von: Petrak, Dominic, et al.
Veröffentlicht: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
von: Wu, Zheng, et al.
Veröffentlicht: (2025)
von: Wu, Zheng, et al.
Veröffentlicht: (2025)
Bridging Reasoning and Action: Hybrid LLM-RL Framework for Efficient Cross-Domain Task-Oriented Dialogue
von: Zhao, Yangyang, et al.
Veröffentlicht: (2026)
von: Zhao, Yangyang, et al.
Veröffentlicht: (2026)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
On the Adaptive Psychological Persuasion of Large Language Models
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
von: Chen, Chaoran, et al.
Veröffentlicht: (2025)
SQLucid: Grounding Natural Language Database Queries with Interactive Explanations
von: Tian, Yuan, et al.
Veröffentlicht: (2024)
von: Tian, Yuan, et al.
Veröffentlicht: (2024)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)
von: Wang, Siting, et al.
Veröffentlicht: (2025)
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
SpiritSight Agent: Advanced GUI Agent with One Look
von: Huang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Huang, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025) -
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
von: Wu, Zongru, et al.
Veröffentlicht: (2025) -
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025) -
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024) -
GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
von: Wu, Zheng, et al.
Veröffentlicht: (2025)