ScreenAgent: A Vision Language Model-driven Computer Control Agent
Fuente:
arXiv
Salvato in:
| Autori principali: | Niu, Runliang, Li, Jindong, Wang, Shiqi, Fu, Yali, Hu, Xiyu, Leng, Xueyuan, Kong, He, Chang, Yi, Wang, Qi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
di: Zhang, Hanzhong, et al.
Pubblicazione: (2026)
di: Zhang, Hanzhong, et al.
Pubblicazione: (2026)
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
di: Yao, Bingsheng, et al.
Pubblicazione: (2026)
di: Yao, Bingsheng, et al.
Pubblicazione: (2026)
Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents
di: Liu, Jiateng, et al.
Pubblicazione: (2026)
di: Liu, Jiateng, et al.
Pubblicazione: (2026)
Personalizing Emotion-aware Conversational Agents? Exploring User Traits-driven Conversational Strategies for Enhanced Interaction
di: Zhang, Yuchong, et al.
Pubblicazione: (2025)
di: Zhang, Yuchong, et al.
Pubblicazione: (2025)
(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
di: Rudaz, Damien, et al.
Pubblicazione: (2026)
di: Rudaz, Damien, et al.
Pubblicazione: (2026)
Design and Evaluation of a Culturally Adapted Multimodal Virtual Agent for PTSD Screening
di: Ozel, Cengiz, et al.
Pubblicazione: (2026)
di: Ozel, Cengiz, et al.
Pubblicazione: (2026)
UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
di: Zhang, Ziyun, et al.
Pubblicazione: (2025)
di: Zhang, Ziyun, et al.
Pubblicazione: (2025)
Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study
di: Shao, Zekai, et al.
Pubblicazione: (2025)
di: Shao, Zekai, et al.
Pubblicazione: (2025)
Dynamic Human Trust Modeling of Autonomous Agents With Varying Capability and Strategy
di: Dekarske, Jason, et al.
Pubblicazione: (2024)
di: Dekarske, Jason, et al.
Pubblicazione: (2024)
AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
di: Leng, Zikang, et al.
Pubblicazione: (2025)
di: Leng, Zikang, et al.
Pubblicazione: (2025)
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
di: Wang, Jihong, et al.
Pubblicazione: (2026)
di: Wang, Jihong, et al.
Pubblicazione: (2026)
Comparing Human Oversight Strategies for Computer-Use Agents
di: Chen, Chaoran, et al.
Pubblicazione: (2026)
di: Chen, Chaoran, et al.
Pubblicazione: (2026)
MarkupLens: Balancing Computer Vision Assistance and Control in Professional Video Annotation for Video-Based Design Tasks
di: He, Tianhao, et al.
Pubblicazione: (2024)
di: He, Tianhao, et al.
Pubblicazione: (2024)
Learning from Brain Topography: A Hierarchical Local-Global Graph-Transformer Network for EEG Emotion Recognition
di: Zhou, Yijin, et al.
Pubblicazione: (2026)
di: Zhou, Yijin, et al.
Pubblicazione: (2026)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
AI-Gadget Kit: Integrating Swarm User Interfaces with LLM-driven Agents for Rich Tabletop Game Applications
di: Guo, Yijie, et al.
Pubblicazione: (2024)
di: Guo, Yijie, et al.
Pubblicazione: (2024)
SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
di: Guo, Longjie, et al.
Pubblicazione: (2025)
di: Guo, Longjie, et al.
Pubblicazione: (2025)
Large Language Model Use Impact Locus of Control
di: Fu, Jenny Xiyu, et al.
Pubblicazione: (2025)
di: Fu, Jenny Xiyu, et al.
Pubblicazione: (2025)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
di: Gao, Shenghan, et al.
Pubblicazione: (2026)
di: Gao, Shenghan, et al.
Pubblicazione: (2026)
Mapping the Design Space of User Experience for Computer Use Agents
di: Cheng, Ruijia, et al.
Pubblicazione: (2026)
di: Cheng, Ruijia, et al.
Pubblicazione: (2026)
Macaron-A2UI: A Model for Generative UI in Personal Agents
di: Kong, Fancy, et al.
Pubblicazione: (2026)
di: Kong, Fancy, et al.
Pubblicazione: (2026)
The Benefits of Prosociality towards AI Agents: Examining the Effects of Helping AI Agents on Human Well-Being
di: Zhu, Zicheng, et al.
Pubblicazione: (2025)
di: Zhu, Zicheng, et al.
Pubblicazione: (2025)
RevTogether: Supporting Science Story Revision with Multiple AI Agents
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration
di: Pan, Bo, et al.
Pubblicazione: (2024)
di: Pan, Bo, et al.
Pubblicazione: (2024)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
di: Sun, Lu, et al.
Pubblicazione: (2025)
di: Sun, Lu, et al.
Pubblicazione: (2025)
SceneScout: Towards AI Agent-driven Access to Street View Imagery for Blind Users
di: Jain, Gaurav, et al.
Pubblicazione: (2025)
di: Jain, Gaurav, et al.
Pubblicazione: (2025)
From Struggle to Success: Context-Aware Guidance for Screen Reader Users in Computer Use
di: Chen, Nan, et al.
Pubblicazione: (2026)
di: Chen, Nan, et al.
Pubblicazione: (2026)
CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
di: Sumyk, Marta, et al.
Pubblicazione: (2026)
di: Sumyk, Marta, et al.
Pubblicazione: (2026)
The Observability Gap: Why Output-Level Human Feedback Fails for LLM Coding Agents
di: Wang, Yinghao, et al.
Pubblicazione: (2026)
di: Wang, Yinghao, et al.
Pubblicazione: (2026)
Unremarkable to Remarkable AI Agent: Exploring Boundaries of Agent Intervention for Adults With and Without Cognitive Impairment
di: Chang, Mai Lee, et al.
Pubblicazione: (2025)
di: Chang, Mai Lee, et al.
Pubblicazione: (2025)
Families' Vision of Generative AI Agents for Household Safety Against Digital and Physical Threats
di: Wen, Zikai, et al.
Pubblicazione: (2025)
di: Wen, Zikai, et al.
Pubblicazione: (2025)
AgentA/B: Automated and Scalable Web A/BTesting with Interactive LLM Agents
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
di: Lu, Yuxuan, et al.
Pubblicazione: (2025)
Visual Embedding of Screen Sequences for User-Flow Search in Example-driven Communication
di: Jeong, Daeheon, et al.
Pubblicazione: (2025)
di: Jeong, Daeheon, et al.
Pubblicazione: (2025)
A Multi-Agent Framework for Democratizing XR Content Creation in K-12 Classrooms
di: Chang, Yuan, et al.
Pubblicazione: (2026)
di: Chang, Yuan, et al.
Pubblicazione: (2026)
A Design Space for Live Music Agents
di: Kim, Yewon, et al.
Pubblicazione: (2026)
di: Kim, Yewon, et al.
Pubblicazione: (2026)
Chartist: Task-driven Eye Movement Control for Chart Reading
di: Shi, Danqing, et al.
Pubblicazione: (2025)
di: Shi, Danqing, et al.
Pubblicazione: (2025)
TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews
di: Xu, Huimin, et al.
Pubblicazione: (2025)
di: Xu, Huimin, et al.
Pubblicazione: (2025)
You Only Look at Screens: Multimodal Chain-of-Action Agents
di: Zhang, Zhuosheng, et al.
Pubblicazione: (2023)
di: Zhang, Zhuosheng, et al.
Pubblicazione: (2023)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
di: Mohanbabu, Ananya Gubbi, et al.
Pubblicazione: (2026)
di: Mohanbabu, Ananya Gubbi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
di: Zhang, Hanzhong, et al.
Pubblicazione: (2026) -
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
di: Yao, Bingsheng, et al.
Pubblicazione: (2026) -
Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents
di: Liu, Jiateng, et al.
Pubblicazione: (2026) -
Personalizing Emotion-aware Conversational Agents? Exploring User Traits-driven Conversational Strategies for Enhanced Interaction
di: Zhang, Yuchong, et al.
Pubblicazione: (2025) -
(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
di: Rudaz, Damien, et al.
Pubblicazione: (2026)