VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Dunjie, Xu, Yiheng, Wang, Junli, Wu, Haoyuan, Wang, Xinyuan, Wang, Zekun, Yang, Junlin, Su, Hongjin, Chen, Jixuan, Chen, Junda, Mao, Yuchen, Zhou, Jingren, Lin, Junyang, Hui, Binyuan, Yu, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
di: Xu, Yiheng, et al.
Pubblicazione: (2024)
di: Xu, Yiheng, et al.
Pubblicazione: (2024)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
di: Xie, Tianbao, et al.
Pubblicazione: (2025)
di: Xie, Tianbao, et al.
Pubblicazione: (2025)
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
di: Xu, Yiheng, et al.
Pubblicazione: (2024)
di: Xu, Yiheng, et al.
Pubblicazione: (2024)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
di: Wang, Bowen, et al.
Pubblicazione: (2026)
di: Wang, Bowen, et al.
Pubblicazione: (2026)
OpenCUA: Open Foundations for Computer-Use Agents
di: Wang, Xinyuan, et al.
Pubblicazione: (2025)
di: Wang, Xinyuan, et al.
Pubblicazione: (2025)
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
di: Bao, Peijun, et al.
Pubblicazione: (2024)
di: Bao, Peijun, et al.
Pubblicazione: (2024)
LumiVideo: An Intelligent Agentic System for Video Color Grading
di: Guo, Yuchen, et al.
Pubblicazione: (2026)
di: Guo, Yuchen, et al.
Pubblicazione: (2026)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
di: Wang, Shaowen, et al.
Pubblicazione: (2025)
di: Wang, Shaowen, et al.
Pubblicazione: (2025)
AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
di: Zhang, Haiyu, et al.
Pubblicazione: (2025)
di: Zhang, Haiyu, et al.
Pubblicazione: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
di: Peng, Bo, et al.
Pubblicazione: (2023)
di: Peng, Bo, et al.
Pubblicazione: (2023)
Literacy Trek
Pubblicazione: (2019)
Pubblicazione: (2019)
"The Writing Trek."
di: Bailey, Valeska
Pubblicazione: (2001)
di: Bailey, Valeska
Pubblicazione: (2001)
NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
di: Bin, Yanrui, et al.
Pubblicazione: (2025)
di: Bin, Yanrui, et al.
Pubblicazione: (2025)
Where No Librarian Has Gone Before...The 10 Best "Star Trek" Episodes on Video.
di: Romanko, Karen A.
Pubblicazione: (1993)
di: Romanko, Karen A.
Pubblicazione: (1993)
Radiance-Field Reinforced Pretraining: Scaling Localization Models with Unlabeled Wireless Signals
di: Wang, Guosheng, et al.
Pubblicazione: (2025)
di: Wang, Guosheng, et al.
Pubblicazione: (2025)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
di: Jia, Weinan, et al.
Pubblicazione: (2025)
di: Jia, Weinan, et al.
Pubblicazione: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
di: Fei, Jiajun, et al.
Pubblicazione: (2024)
di: Fei, Jiajun, et al.
Pubblicazione: (2024)
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
di: Chen, Yixiang, et al.
Pubblicazione: (2025)
di: Chen, Yixiang, et al.
Pubblicazione: (2025)
GaussVideoDreamer: 3D Scene Generation with Video Diffusion and Inconsistency-Aware Gaussian Splatting
di: Hao, Junlin, et al.
Pubblicazione: (2025)
di: Hao, Junlin, et al.
Pubblicazione: (2025)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
di: Ren, Zhongwei, et al.
Pubblicazione: (2025)
di: Ren, Zhongwei, et al.
Pubblicazione: (2025)
Baffin Island Trek
di: Dunn, John
Pubblicazione: (1996)
di: Dunn, John
Pubblicazione: (1996)
ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
di: Wu, Xiaoxue, et al.
Pubblicazione: (2025)
di: Wu, Xiaoxue, et al.
Pubblicazione: (2025)
Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval
di: Chen, Junyang, et al.
Pubblicazione: (2023)
di: Chen, Junyang, et al.
Pubblicazione: (2023)
PARE: Pruning and Adaptive Routing for Efficient Video Generation
di: Wang, Yutong, et al.
Pubblicazione: (2026)
di: Wang, Yutong, et al.
Pubblicazione: (2026)
VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
di: Chen, Xinyu, et al.
Pubblicazione: (2025)
di: Chen, Xinyu, et al.
Pubblicazione: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
di: Liu, Yunze, et al.
Pubblicazione: (2025)
di: Liu, Yunze, et al.
Pubblicazione: (2025)
Harnessing LLMs for Automated Video Content Analysis: An Exploratory Workflow of Short Videos on Depression
di: Liu, Jiaying Lizzy, et al.
Pubblicazione: (2024)
di: Liu, Jiaying Lizzy, et al.
Pubblicazione: (2024)
Harvest Video Foundation Models via Efficient Post-Pretraining
di: Li, Yizhuo, et al.
Pubblicazione: (2023)
di: Li, Yizhuo, et al.
Pubblicazione: (2023)
4Diffusion: Multi-view Video Diffusion Model for 4D Generation
di: Zhang, Haiyu, et al.
Pubblicazione: (2024)
di: Zhang, Haiyu, et al.
Pubblicazione: (2024)
Latent Action Pretraining from Videos
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)
di: Ye, Seonghyeon, et al.
Pubblicazione: (2024)
Learning from Online Videos at Inference Time for Computer-Use Agents
di: Liu, Yujian, et al.
Pubblicazione: (2025)
di: Liu, Yujian, et al.
Pubblicazione: (2025)
$R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
di: Liu, Ye, et al.
Pubblicazione: (2024)
di: Liu, Ye, et al.
Pubblicazione: (2024)
Automatic Retrieval of Specific Cows from Unlabeled Videos
di: Lyu, Jiawen, et al.
Pubblicazione: (2025)
di: Lyu, Jiawen, et al.
Pubblicazione: (2025)
Knowledge Circuits in Pretrained Transformers
di: Yao, Yunzhi, et al.
Pubblicazione: (2024)
di: Yao, Yunzhi, et al.
Pubblicazione: (2024)
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
di: Huang, Binyuan, et al.
Pubblicazione: (2026)
di: Huang, Binyuan, et al.
Pubblicazione: (2026)
LEO: Generative Latent Image Animator for Human Video Synthesis
di: Wang, Yaohui, et al.
Pubblicazione: (2023)
di: Wang, Yaohui, et al.
Pubblicazione: (2023)
CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
di: Wu, Xiaoxue, et al.
Pubblicazione: (2025)
di: Wu, Xiaoxue, et al.
Pubblicazione: (2025)
VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
di: Wang, Yutong, et al.
Pubblicazione: (2025)
di: Wang, Yutong, et al.
Pubblicazione: (2025)
Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning
di: Zhou, Zhaomeng, et al.
Pubblicazione: (2026)
di: Zhou, Zhaomeng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
di: Xu, Yiheng, et al.
Pubblicazione: (2024) -
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
di: Xie, Tianbao, et al.
Pubblicazione: (2025) -
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
di: Xu, Yiheng, et al.
Pubblicazione: (2024) -
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
di: Wang, Bowen, et al.
Pubblicazione: (2026) -
OpenCUA: Open Foundations for Computer-Use Agents
di: Wang, Xinyuan, et al.
Pubblicazione: (2025)