VRAG: Learning World Models for Interactive Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Taiye, Hu, Xun, Ding, Zihan, Jin, Chi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
von: Dai, Josef, et al.
Veröffentlicht: (2024)
von: Dai, Josef, et al.
Veröffentlicht: (2024)
Rethinking Video Generation Model for the Embodied World
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
WorldModelBench: Judging Video Generation Models As World Models
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
Yume: An Interactive World Generation Model
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
GameGen-X: Interactive Open-world Game Video Generation
von: Che, Haoxuan, et al.
Veröffentlicht: (2024)
von: Che, Haoxuan, et al.
Veröffentlicht: (2024)
A Mechanistic View on Video Generation as World Models: State and Dynamics
von: Wang, Luozhou, et al.
Veröffentlicht: (2026)
von: Wang, Luozhou, et al.
Veröffentlicht: (2026)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2025)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2025)
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
von: Fang, Jianjie, et al.
Veröffentlicht: (2026)
von: Fang, Jianjie, et al.
Veröffentlicht: (2026)
Pre-Trained Video Generative Models as World Simulators
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
von: Chen, Kaijin, et al.
Veröffentlicht: (2026)
von: Chen, Kaijin, et al.
Veröffentlicht: (2026)
Yan: Foundational Interactive Video Generation
von: Ye, Deheng, et al.
Veröffentlicht: (2025)
von: Ye, Deheng, et al.
Veröffentlicht: (2025)
Latent Video Prediction Learns Better World Models
von: Alrasheed, Ali J, et al.
Veröffentlicht: (2026)
von: Alrasheed, Ali J, et al.
Veröffentlicht: (2026)
RISE-Video: Can Video Generators Decode Implicit World Rules?
von: Liu, Mingxin, et al.
Veröffentlicht: (2026)
von: Liu, Mingxin, et al.
Veröffentlicht: (2026)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
von: Zhang, Mengchen, et al.
Veröffentlicht: (2025)
von: Zhang, Mengchen, et al.
Veröffentlicht: (2025)
ContactGaussian-WM: Learning Physics-Grounded World Model from Videos
von: Wang, Meizhong, et al.
Veröffentlicht: (2026)
von: Wang, Meizhong, et al.
Veröffentlicht: (2026)
Category Query Learning for Human-Object Interaction Classification
von: Xie, Chi, et al.
Veröffentlicht: (2023)
von: Xie, Chi, et al.
Veröffentlicht: (2023)
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
von: Huang, Zitong, et al.
Veröffentlicht: (2026)
von: Huang, Zitong, et al.
Veröffentlicht: (2026)
FlowAct-R1: Towards Interactive Humanoid Video Generation
von: Wang, Lizhen, et al.
Veröffentlicht: (2026)
von: Wang, Lizhen, et al.
Veröffentlicht: (2026)
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
von: PAN Team, et al.
Veröffentlicht: (2025)
von: PAN Team, et al.
Veröffentlicht: (2025)
Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
Conditional Video Generation for High-Efficiency Video Compression
von: Yi, Fangqiu, et al.
Veröffentlicht: (2025)
von: Yi, Fangqiu, et al.
Veröffentlicht: (2025)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025)
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026)
Interpreting Physics in Video World Models
von: Joseph, Sonia, et al.
Veröffentlicht: (2026)
von: Joseph, Sonia, et al.
Veröffentlicht: (2026)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
von: Xie, Yifan, et al.
Veröffentlicht: (2025)
von: Xie, Yifan, et al.
Veröffentlicht: (2025)
PanoWorld: Geometry-Consistent Panoramic Video World Modeling
von: Jiang, Le, et al.
Veröffentlicht: (2026)
von: Jiang, Le, et al.
Veröffentlicht: (2026)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
von: Romero, David, et al.
Veröffentlicht: (2025)
von: Romero, David, et al.
Veröffentlicht: (2025)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
Pandora: Towards General World Model with Natural Language Actions and Video States
von: Xiang, Jiannan, et al.
Veröffentlicht: (2024)
von: Xiang, Jiannan, et al.
Veröffentlicht: (2024)
How Far is Video Generation from World Model: A Physical Law Perspective
von: Kang, Bingyi, et al.
Veröffentlicht: (2024)
von: Kang, Bingyi, et al.
Veröffentlicht: (2024)
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
von: Hu, Teng, et al.
Veröffentlicht: (2025)
von: Hu, Teng, et al.
Veröffentlicht: (2025)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
von: Chen, Zhifei, et al.
Veröffentlicht: (2025)
von: Chen, Zhifei, et al.
Veröffentlicht: (2025)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
von: Yang, Deshun, et al.
Veröffentlicht: (2024)
von: Yang, Deshun, et al.
Veröffentlicht: (2024)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
von: Gong, Zhantao, et al.
Veröffentlicht: (2025)
von: Gong, Zhantao, et al.
Veröffentlicht: (2025)
Matrix-Game: Interactive World Foundation Model
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
von: Gillman, Nate, et al.
Veröffentlicht: (2025)
von: Gillman, Nate, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
Video-Bench: Human-Aligned Video Generation Benchmark
von: Han, Hui, et al.
Veröffentlicht: (2025)
von: Han, Hui, et al.
Veröffentlicht: (2025)
VideoCAD: A Dataset and Model for Learning Long-Horizon 3D CAD UI Interactions from Video
von: Man, Brandon, et al.
Veröffentlicht: (2025)
von: Man, Brandon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025) -
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
von: Dai, Josef, et al.
Veröffentlicht: (2024) -
Rethinking Video Generation Model for the Embodied World
von: Deng, Yufan, et al.
Veröffentlicht: (2026) -
WorldModelBench: Judging Video Generation Models As World Models
von: Li, Dacheng, et al.
Veröffentlicht: (2025) -
Yume: An Interactive World Generation Model
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)