OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Pengzhou, Wu, Zheng, Wu, Zongru, Zhang, Aston, Zhang, Zhuosheng, Liu, Gongshen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
You Only Look at Screens: Multimodal Chain-of-Action Agents
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
von: Sun, Qiushi, et al.
Veröffentlicht: (2025)
GUI Agents: A Survey
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Faithful Mobile GUI Agents with Guided Advantage Estimator
von: Hu, Haowen, et al.
Veröffentlicht: (2026)
von: Hu, Haowen, et al.
Veröffentlicht: (2026)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
von: Henry, Felix, et al.
Veröffentlicht: (2026)
von: Henry, Felix, et al.
Veröffentlicht: (2026)
API Agents vs. GUI Agents: Divergence and Convergence
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2026)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
von: Nong, Songqin, et al.
Veröffentlicht: (2025)
von: Nong, Songqin, et al.
Veröffentlicht: (2025)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
von: Tang, Liujian, et al.
Veröffentlicht: (2025)
von: Tang, Liujian, et al.
Veröffentlicht: (2025)
GUI Agents with Foundation Models: A Comprehensive Survey
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
Large Language Model-Brained GUI Agents: A Survey
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
UFO: A UI-Focused Agent for Windows OS Interaction
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
von: Sun, Qiushi, et al.
Veröffentlicht: (2024)
MKF-ADS: Multi-Knowledge Fusion Based Self-supervised Anomaly Detection System for Control Area Network
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
von: Chai, Yuxiang, et al.
Veröffentlicht: (2024)
von: Chai, Yuxiang, et al.
Veröffentlicht: (2024)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
von: Tang, Jingyu, et al.
Veröffentlicht: (2025)
von: Tang, Jingyu, et al.
Veröffentlicht: (2025)
TinyClick: Single-Turn Agent for Empowering GUI Automation
von: Pawlowski, Pawel, et al.
Veröffentlicht: (2024)
von: Pawlowski, Pawel, et al.
Veröffentlicht: (2024)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026)
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
AppAgent v2: Advanced Agent for Flexible Mobile Interactions
von: Li, Yanda, et al.
Veröffentlicht: (2024)
von: Li, Yanda, et al.
Veröffentlicht: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
von: Yin, Yiwen, et al.
Veröffentlicht: (2025)
von: Yin, Yiwen, et al.
Veröffentlicht: (2025)
Dynamic Planning for LLM-based Graphical User Interface Automation
von: Zhang, Shaoqing, et al.
Veröffentlicht: (2024)
von: Zhang, Shaoqing, et al.
Veröffentlicht: (2024)
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination
von: Liu, Jijia, et al.
Veröffentlicht: (2023)
von: Liu, Jijia, et al.
Veröffentlicht: (2023)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
von: Gong, Yichen, et al.
Veröffentlicht: (2026)
von: Gong, Yichen, et al.
Veröffentlicht: (2026)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
von: Lee, Jungjae, et al.
Veröffentlicht: (2025)
von: Lee, Jungjae, et al.
Veröffentlicht: (2025)
Aria-UI: Visual Grounding for GUI Instructions
von: Yang, Yuhao, et al.
Veröffentlicht: (2024)
von: Yang, Yuhao, et al.
Veröffentlicht: (2024)
A Survey on (M)LLM-Based GUI Agents
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
AdaptoML-UX: An Adaptive User-centered GUI-based AutoML Toolkit for Non-AI Experts and HCI Researchers
von: Gomaa, Amr, et al.
Veröffentlicht: (2024)
von: Gomaa, Amr, et al.
Veröffentlicht: (2024)
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
von: Xiao, Jingyu, et al.
Veröffentlicht: (2024)
von: Xiao, Jingyu, et al.
Veröffentlicht: (2024)
Artic: AI-oriented Real-time Communication for MLLM Video Assistant
von: Wu, Jiangkai, et al.
Veröffentlicht: (2026)
von: Wu, Jiangkai, et al.
Veröffentlicht: (2026)
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
von: Zhu, Ruiyan, et al.
Veröffentlicht: (2025)
von: Zhu, Ruiyan, et al.
Veröffentlicht: (2025)
PG-Agent: An Agent Powered by Page Graph
von: Chen, Weizhi, et al.
Veröffentlicht: (2025)
von: Chen, Weizhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
von: Wu, Zongru, et al.
Veröffentlicht: (2025) -
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
von: Wu, Zongru, et al.
Veröffentlicht: (2025) -
You Only Look at Screens: Multimodal Chain-of-Action Agents
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023) -
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024) -
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)