ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Kevin Qinghong, Li, Linjie, Gao, Difei, Yang, Zhengyuan, Wu, Shiwei, Bai, Zechen, Lei, Weixian, Wang, Lijuan, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Aria-UI: Visual Grounding for GUI Instructions
von: Yang, Yuhao, et al.
Veröffentlicht: (2024)
von: Yang, Yuhao, et al.
Veröffentlicht: (2024)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
GUI Action Narrator: Where and When Did That Action Take Place?
von: Wu, Qinchen, et al.
Veröffentlicht: (2024)
von: Wu, Qinchen, et al.
Veröffentlicht: (2024)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
ProactiveVA: Proactive Visual Analytics with LLM-Based UI Agent
von: Zhao, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yuheng, et al.
Veröffentlicht: (2025)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
von: Hu, Siyuan, et al.
Veröffentlicht: (2024)
von: Hu, Siyuan, et al.
Veröffentlicht: (2024)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
von: Hu, Haobo, et al.
Veröffentlicht: (2026)
von: Hu, Haobo, et al.
Veröffentlicht: (2026)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
SpiritSight Agent: Advanced GUI Agent with One Look
von: Huang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Huang, Zhiyuan, et al.
Veröffentlicht: (2025)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
Macaron-A2UI: A Model for Generative UI in Personal Agents
von: Kong, Fancy, et al.
Veröffentlicht: (2026)
von: Kong, Fancy, et al.
Veröffentlicht: (2026)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2024)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
von: Lee, Jungjae, et al.
Veröffentlicht: (2025)
von: Lee, Jungjae, et al.
Veröffentlicht: (2025)
Agent-Initiated Interaction in Phone UI Automation
von: Kahlon, Noam, et al.
Veröffentlicht: (2025)
von: Kahlon, Noam, et al.
Veröffentlicht: (2025)
What is the focus of XAI in UI design? Prioritizing UI design principles for enhancing XAI user experience
von: Lei, Dian, et al.
Veröffentlicht: (2024)
von: Lei, Dian, et al.
Veröffentlicht: (2024)
Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
von: Zhang, Shuning, et al.
Veröffentlicht: (2025)
von: Zhang, Shuning, et al.
Veröffentlicht: (2025)
UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
von: Zhang, Ziyun, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyun, et al.
Veröffentlicht: (2025)
MAIC-UI: Making Interactive Courseware with Generative UI
von: Tu, Shangqing, et al.
Veröffentlicht: (2026)
von: Tu, Shangqing, et al.
Veröffentlicht: (2026)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
von: Ni, Minheng, et al.
Veröffentlicht: (2026)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
von: Nayak, Shravan, et al.
Veröffentlicht: (2025)
von: Nayak, Shravan, et al.
Veröffentlicht: (2025)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
von: Yang, Saelyne, et al.
Veröffentlicht: (2026)
von: Yang, Saelyne, et al.
Veröffentlicht: (2026)
Code2World: A GUI World Model via Renderable Code Generation
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
DynaVis: Dynamically Synthesized UI Widgets for Visualization Editing
von: Vaithilingam, Priyan, et al.
Veröffentlicht: (2024)
von: Vaithilingam, Priyan, et al.
Veröffentlicht: (2024)
Morae: Proactively Pausing UI Agents for User Choices
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025) -
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026) -
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025) -
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026) -
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)