GUI Action Narrator: Where and When Did That Action Take Place?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Qinchen, Gao, Difei, Lin, Kevin Qinghong, Wu, Zhuoyu, Guo, Xiangwu, Li, Peiran, Zhang, Weichen, Wang, Hengxu, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
AUTO-Explorer: Automated Data Collection for GUI Agent
von: Guo, Xiangwu, et al.
Veröffentlicht: (2025)
von: Guo, Xiangwu, et al.
Veröffentlicht: (2025)
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
von: Gao, Difei, et al.
Veröffentlicht: (2023)
von: Gao, Difei, et al.
Veröffentlicht: (2023)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
von: Hu, Siyuan, et al.
Veröffentlicht: (2024)
von: Hu, Siyuan, et al.
Veröffentlicht: (2024)
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
von: Hu, Haobo, et al.
Veröffentlicht: (2026)
von: Hu, Haobo, et al.
Veröffentlicht: (2026)
Paper2Video: Automatic Video Generation from Scientific Papers
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
VideoLLM-online: Online Video Large Language Model for Streaming Video
von: Chen, Joya, et al.
Veröffentlicht: (2024)
von: Chen, Joya, et al.
Veröffentlicht: (2024)
Bootstrapping SparseFormers from Vision Foundation Models
von: Gao, Ziteng, et al.
Veröffentlicht: (2023)
von: Gao, Ziteng, et al.
Veröffentlicht: (2023)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
Olaf-World: Orienting Latent Actions for Video World Modeling
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Large Model: A Survey
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
von: Wang, Yuxuan, et al.
Veröffentlicht: (2022)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2022)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation
von: Xiang, Sike, et al.
Veröffentlicht: (2026)
von: Xiang, Sike, et al.
Veröffentlicht: (2026)
Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
von: Hu, Juan, et al.
Veröffentlicht: (2024)
von: Hu, Juan, et al.
Veröffentlicht: (2024)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
Android in the Zoo: Chain-of-Action-Thought for GUI Agents
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2024)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
D-AR: Diffusion via Autoregressive Models
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
von: Zhang, David Junhao, et al.
Veröffentlicht: (2023)
von: Zhang, David Junhao, et al.
Veröffentlicht: (2023)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
ROICtrl: Boosting Instance Control for Visual Generation
von: Gu, Yuchao, et al.
Veröffentlicht: (2024)
von: Gu, Yuchao, et al.
Veröffentlicht: (2024)
Multi-human Interactive Talking Dataset
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
Automated Movie Generation via Multi-Agent CoT Planning
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024) -
AUTO-Explorer: Automated Data Collection for GUI Agent
von: Guo, Xiangwu, et al.
Veröffentlicht: (2025) -
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
von: Gao, Difei, et al.
Veröffentlicht: (2023) -
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024) -
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)