Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiwen, Wu, Jihao, Teng, Yihua, Liao, Minghui, Xu, Nuo, Xiao, Xiao, Wei, Zhongyu, Tang, Duyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025)
by: Lee, Jungjae, et al.
Published: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
A Survey on (M)LLM-Based GUI Agents
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
History-Aware Reasoning for GUI Agents
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Octo-planner: On-device Language Model for Planner-Action Agents
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024)
by: Chai, Yuxiang, et al.
Published: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
Large Language Model-Brained GUI Agents: A Survey
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
WinClick: GUI Grounding with Multimodal Large Language Models
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference
by: Pang, Zi Haur, et al.
Published: (2024)
by: Pang, Zi Haur, et al.
Published: (2024)
A Hopfieldian View-based Interpretation for Chain-of-Thought Reasoning
by: Hu, Lijie, et al.
Published: (2024)
by: Hu, Lijie, et al.
Published: (2024)
ElectionSim: Massive Population Election Simulation Powered by Large Language Model Driven Agents
by: Zhang, Xinnong, et al.
Published: (2024)
by: Zhang, Xinnong, et al.
Published: (2024)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025)
by: Qin, Yujia, et al.
Published: (2025)
Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
by: Sun, Qiushi, et al.
Published: (2024)
by: Sun, Qiushi, et al.
Published: (2024)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
Crisp: Cognitive Restructuring of Negative Thoughts through Multi-turn Supportive Dialogues
by: Zhou, Jinfeng, et al.
Published: (2025)
by: Zhou, Jinfeng, et al.
Published: (2025)
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
by: Sharma, Nikhil, et al.
Published: (2024)
by: Sharma, Nikhil, et al.
Published: (2024)
Taxonomy of User Needs and Actions
by: Shelby, Renee, et al.
Published: (2025)
by: Shelby, Renee, et al.
Published: (2025)
A Systematic Review on Prompt Engineering in Large Language Models for K-12 STEM Education
by: Chen, Eason, et al.
Published: (2024)
by: Chen, Eason, et al.
Published: (2024)
Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning
by: Xu, Yinggan, et al.
Published: (2025)
by: Xu, Yinggan, et al.
Published: (2025)
When ChatGPT is gone: Creativity reverts and homogeneity persists
by: Liu, Qinghan, et al.
Published: (2024)
by: Liu, Qinghan, et al.
Published: (2024)
SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain
by: Gao, Nan, et al.
Published: (2025)
by: Gao, Nan, et al.
Published: (2025)
One Agent Too Many: User Perspectives on Approaches to Multi-agent Conversational AI
by: Clarke, Christopher, et al.
Published: (2024)
by: Clarke, Christopher, et al.
Published: (2024)
A-MEM: Agentic Memory for LLM Agents
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
by: Jangra, Anubhav, et al.
Published: (2025)
by: Jangra, Anubhav, et al.
Published: (2025)
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
by: Liu, Guangyi, et al.
Published: (2025)
by: Liu, Guangyi, et al.
Published: (2025)
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
by: Ye, Rong, et al.
Published: (2025)
by: Ye, Rong, et al.
Published: (2025)
Similar Items
-
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024) -
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025) -
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025) -
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025) -
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)