You Only Look at Screens: Multimodal Chain-of-Action Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhuosheng, Zhang, Aston |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
por: Cheng, Pengzhou, et al.
Publicado: (2025)
por: Cheng, Pengzhou, et al.
Publicado: (2025)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
por: Wu, Zongru, et al.
Publicado: (2025)
por: Wu, Zongru, et al.
Publicado: (2025)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
por: Liu, Yuhang, et al.
Publicado: (2025)
por: Liu, Yuhang, et al.
Publicado: (2025)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
por: Sun, Qiang, et al.
Publicado: (2024)
por: Sun, Qiang, et al.
Publicado: (2024)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
por: Lee, Jungjae, et al.
Publicado: (2025)
por: Lee, Jungjae, et al.
Publicado: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
por: Wu, Zongru, et al.
Publicado: (2025)
por: Wu, Zongru, et al.
Publicado: (2025)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
por: Kim, Taesoo, et al.
Publicado: (2025)
por: Kim, Taesoo, et al.
Publicado: (2025)
Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
por: Li, Yuxuan, et al.
Publicado: (2025)
por: Li, Yuxuan, et al.
Publicado: (2025)
From Context to Action: Analysis of the Impact of State Representation and Context on the Generalization of Multi-Turn Web Navigation Agents
por: Tiwary, Nalin, et al.
Publicado: (2024)
por: Tiwary, Nalin, et al.
Publicado: (2024)
UFO: A UI-Focused Agent for Windows OS Interaction
por: Zhang, Chaoyun, et al.
Publicado: (2024)
por: Zhang, Chaoyun, et al.
Publicado: (2024)
Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation
por: Wang, Yu, et al.
Publicado: (2024)
por: Wang, Yu, et al.
Publicado: (2024)
Large Language Model-Brained GUI Agents: A Survey
por: Zhang, Chaoyun, et al.
Publicado: (2024)
por: Zhang, Chaoyun, et al.
Publicado: (2024)
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
por: Zhang, Hanzhong, et al.
Publicado: (2026)
por: Zhang, Hanzhong, et al.
Publicado: (2026)
Multimodal Transformer Models for Turn-taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction during Cooperative Gameplay
por: Bae, Young-Ho, et al.
Publicado: (2025)
por: Bae, Young-Ho, et al.
Publicado: (2025)
Towards End-to-End Open Conversational Machine Reading
por: Zhou, Sizhe, et al.
Publicado: (2022)
por: Zhou, Sizhe, et al.
Publicado: (2022)
Hey GPT, Can You be More Racist? Analysis from Crowdsourced Attempts to Elicit Biased Content from Generative AI
por: Guo, Hangzhi, et al.
Publicado: (2024)
por: Guo, Hangzhi, et al.
Publicado: (2024)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
por: Sung, Yoo Yeon, et al.
Publicado: (2025)
por: Sung, Yoo Yeon, et al.
Publicado: (2025)
Why Would You Suggest That? Human Trust in Language Model Responses
por: Sharma, Manasi, et al.
Publicado: (2024)
por: Sharma, Manasi, et al.
Publicado: (2024)
Empowering Private Tutoring by Chaining Large Language Models
por: Chen, Yulin, et al.
Publicado: (2023)
por: Chen, Yulin, et al.
Publicado: (2023)
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
por: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Publicado: (2025)
por: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Publicado: (2025)
Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
por: Grunde-McLaughlin, Madeleine, et al.
Publicado: (2023)
por: Grunde-McLaughlin, Madeleine, et al.
Publicado: (2023)
Cognition Chain for Explainable Psychological Stress Detection on Social Media
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer Use
por: Sapkota, Shardul, et al.
Publicado: (2026)
por: Sapkota, Shardul, et al.
Publicado: (2026)
Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?
por: Liu, Xiaoze, et al.
Publicado: (2026)
por: Liu, Xiaoze, et al.
Publicado: (2026)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
por: Shen, Hua, et al.
Publicado: (2025)
por: Shen, Hua, et al.
Publicado: (2025)
Analysing Explanation-Related Interactions in Collaborative Perception-Cognition-Communication-Action
por: Vilamala, Marc Roig, et al.
Publicado: (2024)
por: Vilamala, Marc Roig, et al.
Publicado: (2024)
Evaluating Large Language Models' Ability Using a Psychiatric Screening Tool Based on Metaphor and Sarcasm Scenarios
por: Yakura, Hiromu
Publicado: (2023)
por: Yakura, Hiromu
Publicado: (2023)
SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain
por: Gao, Nan, et al.
Publicado: (2025)
por: Gao, Nan, et al.
Publicado: (2025)
Chain of Empathy: Enhancing Empathetic Response of Large Language Models Based on Psychotherapy Models
por: Lee, Yoon Kyung, et al.
Publicado: (2023)
por: Lee, Yoon Kyung, et al.
Publicado: (2023)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
por: Qian, Cheng, et al.
Publicado: (2024)
por: Qian, Cheng, et al.
Publicado: (2024)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
por: Yang, Bufang, et al.
Publicado: (2025)
por: Yang, Bufang, et al.
Publicado: (2025)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
por: Yang, Bufang, et al.
Publicado: (2025)
por: Yang, Bufang, et al.
Publicado: (2025)
mrCAD: Multimodal Refinement of Computer-aided Designs
por: McCarthy, William P., et al.
Publicado: (2025)
por: McCarthy, William P., et al.
Publicado: (2025)
AgentCTG: Harnessing Multi-Agent Collaboration for Fine-Grained Precise Control in Text Generation
por: Zhou, Xinxu, et al.
Publicado: (2025)
por: Zhou, Xinxu, et al.
Publicado: (2025)
MathBuddy: A Multimodal System for Affective Math Tutoring
por: Kar, Debanjana, et al.
Publicado: (2025)
por: Kar, Debanjana, et al.
Publicado: (2025)
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
por: Stacchio, Lorenzo, et al.
Publicado: (2025)
por: Stacchio, Lorenzo, et al.
Publicado: (2025)
Visualization Literacy of Multimodal Large Language Models: A Comparative Study
por: Li, Zhimin, et al.
Publicado: (2024)
por: Li, Zhimin, et al.
Publicado: (2024)
Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality
por: Pei, Jiahuan, et al.
Publicado: (2024)
por: Pei, Jiahuan, et al.
Publicado: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
por: Zhu, Andrew, et al.
Publicado: (2025)
por: Zhu, Andrew, et al.
Publicado: (2025)
Automated Interpretability and Feature Discovery in Language Models with Agents
por: Marin-Llobet, Arnau, et al.
Publicado: (2026)
por: Marin-Llobet, Arnau, et al.
Publicado: (2026)
Ejemplares similares
-
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
por: Cheng, Pengzhou, et al.
Publicado: (2025) -
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
por: Wu, Zongru, et al.
Publicado: (2025) -
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
por: Liu, Yuhang, et al.
Publicado: (2025) -
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
por: Sun, Qiang, et al.
Publicado: (2024) -
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
por: Lee, Jungjae, et al.
Publicado: (2025)