VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Kevin Qinghong, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Paper2Video: Automatic Video Generation from Scientific Papers
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
von: Hu, Siyuan, et al.
Veröffentlicht: (2025)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
Bootstrapping SparseFormers from Vision Foundation Models
von: Gao, Ziteng, et al.
Veröffentlicht: (2023)
von: Gao, Ziteng, et al.
Veröffentlicht: (2023)
VideoLLM-online: Online Video Large Language Model for Streaming Video
von: Chen, Joya, et al.
Veröffentlicht: (2024)
von: Chen, Joya, et al.
Veröffentlicht: (2024)
GUI Action Narrator: Where and When Did That Action Take Place?
von: Wu, Qinchen, et al.
Veröffentlicht: (2024)
von: Wu, Qinchen, et al.
Veröffentlicht: (2024)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
von: Zhao, Henghao, et al.
Veröffentlicht: (2023)
von: Zhao, Henghao, et al.
Veröffentlicht: (2023)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
TPDiff: Temporal Pyramid Video Diffusion Model
von: Ran, Lingmin, et al.
Veröffentlicht: (2025)
von: Ran, Lingmin, et al.
Veröffentlicht: (2025)
Learning Long-form Video Prior via Generative Pre-Training
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
ROICtrl: Boosting Instance Control for Visual Generation
von: Gu, Yuchao, et al.
Veröffentlicht: (2024)
von: Gu, Yuchao, et al.
Veröffentlicht: (2024)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
von: Gu, Yuchao, et al.
Veröffentlicht: (2025)
von: Gu, Yuchao, et al.
Veröffentlicht: (2025)
Impossible Videos
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
Mitty: Diffusion-based Human-to-Robot Video Generation
von: Song, Yiren, et al.
Veröffentlicht: (2025)
von: Song, Yiren, et al.
Veröffentlicht: (2025)
Computer-Use Agents as Judges for Generative User Interface
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Large Model: A Survey
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
von: Wu, Weijia, et al.
Veröffentlicht: (2025)
P-Flow: Prompting Visual Effects Generation
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
Unsupervised Open-Vocabulary Object Localization in Videos
von: Fan, Ke, et al.
Veröffentlicht: (2023)
von: Fan, Ke, et al.
Veröffentlicht: (2023)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
D-AR: Diffusion via Autoregressive Models
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
von: Gao, Ziteng, et al.
Veröffentlicht: (2025)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
von: Zhao, Rui, et al.
Veröffentlicht: (2025)
von: Zhao, Rui, et al.
Veröffentlicht: (2025)
Ego-centric Predictive Model Conditioned on Hand Trajectories
von: Zhang, Binjie, et al.
Veröffentlicht: (2025)
von: Zhang, Binjie, et al.
Veröffentlicht: (2025)
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
von: Xie, Jinheng, et al.
Veröffentlicht: (2024)
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2023)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2023)
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
von: Mei, Haiyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Paper2Video: Automatic Video Generation from Scientific Papers
von: Zhu, Zeyu, et al.
Veröffentlicht: (2025) -
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025) -
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
von: Hu, Siyuan, et al.
Veröffentlicht: (2025) -
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025) -
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)