Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hassan, Mariam, Van Delft, Bastien, Li, Wuyang, Alahi, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
by: Hassan, Mariam, et al.
Published: (2026)
by: Hassan, Mariam, et al.
Published: (2026)
LayerSync: Self-aligning Intermediate Layers
by: Haghighi, Yasaman, et al.
Published: (2025)
by: Haghighi, Yasaman, et al.
Published: (2025)
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by: Li, Wuyang, et al.
Published: (2025)
by: Li, Wuyang, et al.
Published: (2025)
CODE: Confident Ordinary Differential Editing
by: van Delft, Bastien, et al.
Published: (2024)
by: van Delft, Bastien, et al.
Published: (2024)
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Drift-Resistant Navigation World Model with Anchored Epipolar Guidance
by: Luan, Po-Chien, et al.
Published: (2026)
by: Luan, Po-Chien, et al.
Published: (2026)
VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection
by: Li, Wuyang, et al.
Published: (2025)
by: Li, Wuyang, et al.
Published: (2025)
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
by: Luan, Po-Chien, et al.
Published: (2026)
by: Luan, Po-Chien, et al.
Published: (2026)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026)
by: Li, Wuyang, et al.
Published: (2026)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
by: Dou, Huanzhang, et al.
Published: (2024)
by: Dou, Huanzhang, et al.
Published: (2024)
Decoupled Video Generation with Chain of Training-free Diffusion Model Experts
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
Anchored Diffusion for Video Face Reenactment
by: Kligvasser, Idan, et al.
Published: (2024)
by: Kligvasser, Idan, et al.
Published: (2024)
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
by: Wang, Yuji, et al.
Published: (2025)
by: Wang, Yuji, et al.
Published: (2025)
MAD: Motion Appearance Decoupling for efficient Driving World Models
by: Rahimi, Ahmad, et al.
Published: (2026)
by: Rahimi, Ahmad, et al.
Published: (2026)
Grid Diffusion Models for Text-to-Video Generation
by: Lee, Taegyeong, et al.
Published: (2024)
by: Lee, Taegyeong, et al.
Published: (2024)
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
by: Haghighi, Yasaman, et al.
Published: (2026)
by: Haghighi, Yasaman, et al.
Published: (2026)
Decoupled Diffusion Sparks Adaptive Scene Generation
by: Zhou, Yunsong, et al.
Published: (2025)
by: Zhou, Yunsong, et al.
Published: (2025)
PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education
by: M, Megha Mariam K., et al.
Published: (2026)
by: M, Megha Mariam K., et al.
Published: (2026)
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
by: Liu, Yifei, et al.
Published: (2026)
by: Liu, Yifei, et al.
Published: (2026)
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
by: Chen, Mu, et al.
Published: (2025)
by: Chen, Mu, et al.
Published: (2025)
Endora: Video Generation Models as Endoscopy Simulators
by: Li, Chenxin, et al.
Published: (2024)
by: Li, Chenxin, et al.
Published: (2024)
VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
by: Zhang, Yabo, et al.
Published: (2024)
by: Zhang, Yabo, et al.
Published: (2024)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
by: Shen, Cuifeng, et al.
Published: (2025)
by: Shen, Cuifeng, et al.
Published: (2025)
HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
by: Zhou, Haiyang, et al.
Published: (2025)
by: Zhou, Haiyang, et al.
Published: (2025)
Video Perception Models for 3D Scene Synthesis
by: Huang, Rui, et al.
Published: (2025)
by: Huang, Rui, et al.
Published: (2025)
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models
by: Cai, Yufei, et al.
Published: (2025)
by: Cai, Yufei, et al.
Published: (2025)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
Inflation with Diffusion: Efficient Temporal Adaptation for Text-to-Video Super-Resolution
by: Yuan, Xin, et al.
Published: (2024)
by: Yuan, Xin, et al.
Published: (2024)
Edit Temporal-Consistent Videos with Image Diffusion Model
by: Wang, Yuanzhi, et al.
Published: (2023)
by: Wang, Yuanzhi, et al.
Published: (2023)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
by: Bendel, Matthew, et al.
Published: (2026)
by: Bendel, Matthew, et al.
Published: (2026)
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
by: He, Haibin, et al.
Published: (2026)
by: He, Haibin, et al.
Published: (2026)
Layout Agnostic Scene Text Image Synthesis with Diffusion Models
by: Zhangli, Qilong, et al.
Published: (2024)
by: Zhangli, Qilong, et al.
Published: (2024)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
Similar Items
-
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
by: Hassan, Mariam, et al.
Published: (2026) -
LayerSync: Self-aligning Intermediate Layers
by: Haghighi, Yasaman, et al.
Published: (2025) -
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by: Li, Wuyang, et al.
Published: (2025) -
CODE: Confident Ordinary Differential Editing
by: van Delft, Bastien, et al.
Published: (2024) -
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025)