ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Difei, Ji, Lei, Bai, Zechen, Ouyang, Mingyu, Li, Peiran, Mao, Dongxing, Wu, Qinchen, Zhang, Weichen, Wang, Peiyi, Guo, Xiangwu, Wang, Hengxu, Zhou, Luowei, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GUI Action Narrator: Where and When Did That Action Take Place?
by: Wu, Qinchen, et al.
Published: (2024)
by: Wu, Qinchen, et al.
Published: (2024)
AUTO-Explorer: Automated Data Collection for GUI Agent
by: Guo, Xiangwu, et al.
Published: (2025)
by: Guo, Xiangwu, et al.
Published: (2025)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
by: Hu, Siyuan, et al.
Published: (2024)
by: Hu, Siyuan, et al.
Published: (2024)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
by: Zhao, Henry Hengyuan, et al.
Published: (2024)
by: Zhao, Henry Hengyuan, et al.
Published: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Visual Grounding Methods for Efficient Interaction with Desktop Graphical User Interfaces
by: Ettifouri, El Hassane, et al.
Published: (2024)
by: Ettifouri, El Hassane, et al.
Published: (2024)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
by: Bai, Zechen, et al.
Published: (2025)
by: Bai, Zechen, et al.
Published: (2025)
Impossible Videos
by: Bai, Zechen, et al.
Published: (2025)
by: Bai, Zechen, et al.
Published: (2025)
ShowUI-Aloha: Human-Taught GUI Agent
by: Zhang, Yichun, et al.
Published: (2026)
by: Zhang, Yichun, et al.
Published: (2026)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
AFFMAE: Scalable and Efficient Vision Pretraining for Desktop Graphics Cards
by: Smerkous, David, et al.
Published: (2026)
by: Smerkous, David, et al.
Published: (2026)
Dynamic Planning for LLM-based Graphical User Interface Automation
by: Zhang, Shaoqing, et al.
Published: (2024)
by: Zhang, Shaoqing, et al.
Published: (2024)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
When the Internet Knocks, Unlock the Desktop.
by: Nyerges, Mike
Published: (1999)
by: Nyerges, Mike
Published: (1999)
The Rise of the Graphical User Interface.
by: Edwards, Alastair D. N.
Published: (1996)
by: Edwards, Alastair D. N.
Published: (1996)
Overview of Graphical User Interfaces.
by: Hulser, Richard P.
Published: (1993)
by: Hulser, Richard P.
Published: (1993)
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
by: Hu, Haobo, et al.
Published: (2026)
by: Hu, Haobo, et al.
Published: (2026)
Training Users for Desktop Access.
by: King-Blandford, Marcia
Published: (1998)
by: King-Blandford, Marcia
Published: (1998)
VideoLLM-online: Online Video Large Language Model for Streaming Video
by: Chen, Joya, et al.
Published: (2024)
by: Chen, Joya, et al.
Published: (2024)
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
by: Liu, Xiaokang, et al.
Published: (2026)
by: Liu, Xiaokang, et al.
Published: (2026)
Factorized Learning for Temporally Grounded Video-Language Models
by: Zeng, Wenzheng, et al.
Published: (2025)
by: Zeng, Wenzheng, et al.
Published: (2025)
Low-code LLM: Graphical User Interface over Large Language Models
by: Cai, Yuzhe, et al.
Published: (2023)
by: Cai, Yuzhe, et al.
Published: (2023)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Factorized Visual Tokenization and Generation
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
by: Lin, Yiqi, et al.
Published: (2026)
by: Lin, Yiqi, et al.
Published: (2026)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)
by: Han, Zongbo, et al.
Published: (2024)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
GUIDE: Graphical User Interface Data for Execution
by: Chawla, Rajat, et al.
Published: (2024)
by: Chawla, Rajat, et al.
Published: (2024)
Graphical User Interfaces and Library Systems: End-User Reactions.
by: Zorn, Margaret, et al.
Published: (1995)
by: Zorn, Margaret, et al.
Published: (1995)
GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
by: Wang, Yuxuan, et al.
Published: (2022)
by: Wang, Yuxuan, et al.
Published: (2022)
Graphical User Interface Programming in Introductory Computer Science.
by: Skolnick, Michael M., et al.
Published: (1995)
by: Skolnick, Michael M., et al.
Published: (1995)
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
by: Hu, Juan, et al.
Published: (2024)
by: Hu, Juan, et al.
Published: (2024)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
by: Lu, Yuheng, et al.
Published: (2025)
by: Lu, Yuheng, et al.
Published: (2025)
Knowledge Fusion‐Based Neural Network Control for Uncertain Nonlinear Systems via Deterministic Learning
by: Qinchen Yang, et al.
Published: (2025)
by: Qinchen Yang, et al.
Published: (2025)
An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
Post Processing Graphical User Interface for Heat Flow Visualization
by: Olt, Lars, et al.
Published: (2025)
by: Olt, Lars, et al.
Published: (2025)
Similar Items
-
GUI Action Narrator: Where and When Did That Action Take Place?
by: Wu, Qinchen, et al.
Published: (2024) -
AUTO-Explorer: Automated Data Collection for GUI Agent
by: Guo, Xiangwu, et al.
Published: (2025) -
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
by: Hu, Siyuan, et al.
Published: (2024) -
LOVA3: Learning to Visual Question Answering, Asking and Assessment
by: Zhao, Henry Hengyuan, et al.
Published: (2024) -
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)