ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Siyuan, Lin, Kevin Qinghong, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Code2World: A GUI World Model via Renderable Code Generation
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
von: Zheng, Yuhao, et al.
Veröffentlicht: (2026)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
von: Qin, Yujia, et al.
Veröffentlicht: (2025)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
von: Yang, Saelyne, et al.
Veröffentlicht: (2026)
von: Yang, Saelyne, et al.
Veröffentlicht: (2026)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
RWKV-UI: UI Understanding with Enhanced Perception and Reasoning
von: Yang, Jiaxi, et al.
Veröffentlicht: (2025)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2025)
SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
von: Jing, Hongyi, et al.
Veröffentlicht: (2025)
UI-UG: A Unified MLLM for UI Understanding and Generation
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
von: Garg, Aryan, et al.
Veröffentlicht: (2025)
von: Garg, Aryan, et al.
Veröffentlicht: (2025)
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
von: You, Keen, et al.
Veröffentlicht: (2024)
von: You, Keen, et al.
Veröffentlicht: (2024)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
von: Li, Zisu, et al.
Veröffentlicht: (2025)
von: Li, Zisu, et al.
Veröffentlicht: (2025)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
V-Hands: Touchscreen-based Hand Tracking for Remote Whiteboard Interaction
von: Liu, Xinshuang, et al.
Veröffentlicht: (2024)
von: Liu, Xinshuang, et al.
Veröffentlicht: (2024)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
von: Luo, Run, et al.
Veröffentlicht: (2025)
von: Luo, Run, et al.
Veröffentlicht: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
UniHands: Unifying Various Wild-Collected Keypoints for Personalized Hand Reconstruction
von: Zhang, Menghe, et al.
Veröffentlicht: (2024)
von: Zhang, Menghe, et al.
Veröffentlicht: (2024)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
von: Garg, Mallika, et al.
Veröffentlicht: (2025)
von: Garg, Mallika, et al.
Veröffentlicht: (2025)
Morae: Proactively Pausing UI Agents for User Choices
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
HandDAGT: A Denoising Adaptive Graph Transformer for 3D Hand Pose Estimation
von: Cheng, Wencan, et al.
Veröffentlicht: (2024)
von: Cheng, Wencan, et al.
Veröffentlicht: (2024)
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
E-ANT: A Large-Scale Dataset for Efficient Automatic GUI NavigaTion
von: Wang, Ke, et al.
Veröffentlicht: (2024)
von: Wang, Ke, et al.
Veröffentlicht: (2024)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
von: Chang, Haochen, et al.
Veröffentlicht: (2025)
von: Chang, Haochen, et al.
Veröffentlicht: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
von: Han, Tianhao, et al.
Veröffentlicht: (2026)
von: Han, Tianhao, et al.
Veröffentlicht: (2026)
HaDR: Applying Domain Randomization for Generating Synthetic Multimodal Dataset for Hand Instance Segmentation in Cluttered Industrial Environments
von: Grushko, Stefan, et al.
Veröffentlicht: (2023)
von: Grushko, Stefan, et al.
Veröffentlicht: (2023)
SpiritSight Agent: Advanced GUI Agent with One Look
von: Huang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Huang, Zhiyuan, et al.
Veröffentlicht: (2025)
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition
von: Garg, Mallika, et al.
Veröffentlicht: (2024)
von: Garg, Mallika, et al.
Veröffentlicht: (2024)
HpEIS: Learning Hand Pose Embeddings for Multimedia Interactive Systems
von: Xu, Songpei, et al.
Veröffentlicht: (2024)
von: Xu, Songpei, et al.
Veröffentlicht: (2024)
AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding
von: Khan, Safwat Ali, et al.
Veröffentlicht: (2024)
von: Khan, Safwat Ali, et al.
Veröffentlicht: (2024)
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
von: Zhao, Yiming, et al.
Veröffentlicht: (2024)
von: Zhao, Yiming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024) -
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026) -
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026) -
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025) -
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)