ScreenAgent: A Vision Language Model-driven Computer Control Agent
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Niu, Runliang, Li, Jindong, Wang, Shiqi, Fu, Yali, Hu, Xiyu, Leng, Xueyuan, Kong, He, Chang, Yi, Wang, Qi |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
par: Zhang, Hanzhong, et autres
Publié: (2026)
par: Zhang, Hanzhong, et autres
Publié: (2026)
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
par: Yao, Bingsheng, et autres
Publié: (2026)
par: Yao, Bingsheng, et autres
Publié: (2026)
Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents
par: Liu, Jiateng, et autres
Publié: (2026)
par: Liu, Jiateng, et autres
Publié: (2026)
Personalizing Emotion-aware Conversational Agents? Exploring User Traits-driven Conversational Strategies for Enhanced Interaction
par: Zhang, Yuchong, et autres
Publié: (2025)
par: Zhang, Yuchong, et autres
Publié: (2025)
(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
par: Rudaz, Damien, et autres
Publié: (2026)
par: Rudaz, Damien, et autres
Publié: (2026)
Design and Evaluation of a Culturally Adapted Multimodal Virtual Agent for PTSD Screening
par: Ozel, Cengiz, et autres
Publié: (2026)
par: Ozel, Cengiz, et autres
Publié: (2026)
UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
par: Zhang, Ziyun, et autres
Publié: (2025)
par: Zhang, Ziyun, et autres
Publié: (2025)
Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study
par: Shao, Zekai, et autres
Publié: (2025)
par: Shao, Zekai, et autres
Publié: (2025)
Dynamic Human Trust Modeling of Autonomous Agents With Varying Capability and Strategy
par: Dekarske, Jason, et autres
Publié: (2024)
par: Dekarske, Jason, et autres
Publié: (2024)
AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
par: Leng, Zikang, et autres
Publié: (2025)
par: Leng, Zikang, et autres
Publié: (2025)
ColorBrowserAgent: Complex Long-Horizon Browser Agent with Adaptive Knowledge Evolution
par: Wang, Jihong, et autres
Publié: (2026)
par: Wang, Jihong, et autres
Publié: (2026)
Comparing Human Oversight Strategies for Computer-Use Agents
par: Chen, Chaoran, et autres
Publié: (2026)
par: Chen, Chaoran, et autres
Publié: (2026)
MarkupLens: Balancing Computer Vision Assistance and Control in Professional Video Annotation for Video-Based Design Tasks
par: He, Tianhao, et autres
Publié: (2024)
par: He, Tianhao, et autres
Publié: (2024)
Learning from Brain Topography: A Hierarchical Local-Global Graph-Transformer Network for EEG Emotion Recognition
par: Zhou, Yijin, et autres
Publié: (2026)
par: Zhou, Yijin, et autres
Publié: (2026)
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
par: Lu, Yuxuan, et autres
Publié: (2025)
par: Lu, Yuxuan, et autres
Publié: (2025)
AI-Gadget Kit: Integrating Swarm User Interfaces with LLM-driven Agents for Rich Tabletop Game Applications
par: Guo, Yijie, et autres
Publié: (2024)
par: Guo, Yijie, et autres
Publié: (2024)
SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
par: Guo, Longjie, et autres
Publié: (2025)
par: Guo, Longjie, et autres
Publié: (2025)
Large Language Model Use Impact Locus of Control
par: Fu, Jenny Xiyu, et autres
Publié: (2025)
par: Fu, Jenny Xiyu, et autres
Publié: (2025)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
par: Lu, Yuxuan, et autres
Publié: (2025)
par: Lu, Yuxuan, et autres
Publié: (2025)
SCSimulator: An Exploratory Visual Analytics Framework for Partner Selection in Supply Chains through LLM-driven Multi-Agent Simulation
par: Gao, Shenghan, et autres
Publié: (2026)
par: Gao, Shenghan, et autres
Publié: (2026)
Mapping the Design Space of User Experience for Computer Use Agents
par: Cheng, Ruijia, et autres
Publié: (2026)
par: Cheng, Ruijia, et autres
Publié: (2026)
Macaron-A2UI: A Model for Generative UI in Personal Agents
par: Kong, Fancy, et autres
Publié: (2026)
par: Kong, Fancy, et autres
Publié: (2026)
The Benefits of Prosociality towards AI Agents: Examining the Effects of Helping AI Agents on Human Well-Being
par: Zhu, Zicheng, et autres
Publié: (2025)
par: Zhu, Zicheng, et autres
Publié: (2025)
RevTogether: Supporting Science Story Revision with Multiple AI Agents
par: Zhang, Yu, et autres
Publié: (2025)
par: Zhang, Yu, et autres
Publié: (2025)
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration
par: Pan, Bo, et autres
Publié: (2024)
par: Pan, Bo, et autres
Publié: (2024)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
par: Sun, Lu, et autres
Publié: (2025)
par: Sun, Lu, et autres
Publié: (2025)
SceneScout: Towards AI Agent-driven Access to Street View Imagery for Blind Users
par: Jain, Gaurav, et autres
Publié: (2025)
par: Jain, Gaurav, et autres
Publié: (2025)
From Struggle to Success: Context-Aware Guidance for Screen Reader Users in Computer Use
par: Chen, Nan, et autres
Publié: (2026)
par: Chen, Nan, et autres
Publié: (2026)
CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
par: Sumyk, Marta, et autres
Publié: (2026)
par: Sumyk, Marta, et autres
Publié: (2026)
The Observability Gap: Why Output-Level Human Feedback Fails for LLM Coding Agents
par: Wang, Yinghao, et autres
Publié: (2026)
par: Wang, Yinghao, et autres
Publié: (2026)
Unremarkable to Remarkable AI Agent: Exploring Boundaries of Agent Intervention for Adults With and Without Cognitive Impairment
par: Chang, Mai Lee, et autres
Publié: (2025)
par: Chang, Mai Lee, et autres
Publié: (2025)
Families' Vision of Generative AI Agents for Household Safety Against Digital and Physical Threats
par: Wen, Zikai, et autres
Publié: (2025)
par: Wen, Zikai, et autres
Publié: (2025)
AgentA/B: Automated and Scalable Web A/BTesting with Interactive LLM Agents
par: Lu, Yuxuan, et autres
Publié: (2025)
par: Lu, Yuxuan, et autres
Publié: (2025)
Visual Embedding of Screen Sequences for User-Flow Search in Example-driven Communication
par: Jeong, Daeheon, et autres
Publié: (2025)
par: Jeong, Daeheon, et autres
Publié: (2025)
A Multi-Agent Framework for Democratizing XR Content Creation in K-12 Classrooms
par: Chang, Yuan, et autres
Publié: (2026)
par: Chang, Yuan, et autres
Publié: (2026)
A Design Space for Live Music Agents
par: Kim, Yewon, et autres
Publié: (2026)
par: Kim, Yewon, et autres
Publié: (2026)
Chartist: Task-driven Eye Movement Control for Chart Reading
par: Shi, Danqing, et autres
Publié: (2025)
par: Shi, Danqing, et autres
Publié: (2025)
TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews
par: Xu, Huimin, et autres
Publié: (2025)
par: Xu, Huimin, et autres
Publié: (2025)
You Only Look at Screens: Multimodal Chain-of-Action Agents
par: Zhang, Zhuosheng, et autres
Publié: (2023)
par: Zhang, Zhuosheng, et autres
Publié: (2023)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
par: Mohanbabu, Ananya Gubbi, et autres
Publié: (2026)
par: Mohanbabu, Ananya Gubbi, et autres
Publié: (2026)
Documents similaires
-
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
par: Zhang, Hanzhong, et autres
Publié: (2026) -
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
par: Yao, Bingsheng, et autres
Publié: (2026) -
Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents
par: Liu, Jiateng, et autres
Publié: (2026) -
Personalizing Emotion-aware Conversational Agents? Exploring User Traits-driven Conversational Strategies for Enhanced Interaction
par: Zhang, Yuchong, et autres
Publié: (2025) -
(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
par: Rudaz, Damien, et autres
Publié: (2026)