See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zongru, Mao, Rui, Tian, Zhiyuan, Cheng, Pengzhou, Ju, Tianjie, Wu, Zheng, Dong, Lingzhong, Sheng, Haiyue, Zhang, Zhuosheng, Liu, Gongshen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
by: Dong, Lingzhong, et al.
Published: (2025)
by: Dong, Lingzhong, et al.
Published: (2025)
Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
Transferring Backdoors between Large Language Models by Knowledge Distillation
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
On the Adaptive Psychological Persuasion of Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
Faithful Mobile GUI Agents with Guided Advantage Estimator
by: Hu, Haowen, et al.
Published: (2026)
by: Hu, Haowen, et al.
Published: (2026)
OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents
by: Wu, Zheng, et al.
Published: (2026)
by: Wu, Zheng, et al.
Published: (2026)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
Disagreements in Reasoning: How a Model's Thinking Process Dictates Persuasion in Multi-Agent Systems
by: Zhao, Haodong, et al.
Published: (2025)
by: Zhao, Haodong, et al.
Published: (2025)
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
by: Deng, Zehao, et al.
Published: (2025)
by: Deng, Zehao, et al.
Published: (2025)
MKF-ADS: Multi-Knowledge Fusion Based Self-supervised Anomaly Detection System for Control Area Network
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection
by: Wu, Zheng, et al.
Published: (2026)
by: Wu, Zheng, et al.
Published: (2026)
Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
by: Ju, Tianjie, et al.
Published: (2024)
by: Ju, Tianjie, et al.
Published: (2024)
Investigating Multi-Hop Factual Shortcuts in Knowledge Editing of Large Language Models
by: Ju, Tianjie, et al.
Published: (2024)
by: Ju, Tianjie, et al.
Published: (2024)
Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought
by: Zhang, Yuyi, et al.
Published: (2025)
by: Zhang, Yuyi, et al.
Published: (2025)
Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
by: Zhao, Haodong, et al.
Published: (2025)
by: Zhao, Haodong, et al.
Published: (2025)
Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
by: Zhao, Haodong, et al.
Published: (2024)
by: Zhao, Haodong, et al.
Published: (2024)
How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study
by: Ju, Tianjie, et al.
Published: (2024)
by: Ju, Tianjie, et al.
Published: (2024)
Probing then Editing Response Personality of Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
by: Zhang, Yimeng, et al.
Published: (2025)
by: Zhang, Yimeng, et al.
Published: (2025)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
by: Cheng, Pengzhou, et al.
Published: (2023)
by: Cheng, Pengzhou, et al.
Published: (2023)
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
by: Guo, Yuan, et al.
Published: (2025)
by: Guo, Yuan, et al.
Published: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
by: Ma, Xinbei, et al.
Published: (2024)
by: Ma, Xinbei, et al.
Published: (2024)
On the Robustness of Editing Large Language Models
by: Ma, Xinbei, et al.
Published: (2024)
by: Ma, Xinbei, et al.
Published: (2024)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
See, Think, Learn: A Self-Taught Multimodal Reasoner
by: Sharma, Sourabh, et al.
Published: (2025)
by: Sharma, Sourabh, et al.
Published: (2025)
Similar Items
-
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
by: Wu, Zongru, et al.
Published: (2025) -
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025) -
GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
by: Wu, Zheng, et al.
Published: (2025) -
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025) -
Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
by: Dong, Lingzhong, et al.
Published: (2025)