A Survey: Spatiotemporal Consistency in Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Zhiyu, Chen, Kehai, Bai, Xuefeng, Jiang, Ruili, Li, Juntao, Li, Hongdong, Liu, Jin, Xiang, Yang, Yu, Jun, Zhang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026)
Mitigating Multimodal Hallucination via Phase-wise Self-reward
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors
von: Wang, Rong, et al.
Veröffentlicht: (2026)
von: Wang, Rong, et al.
Veröffentlicht: (2026)
BachVid: Training-Free Video Generation with Consistent Background and Character
von: Yan, Han, et al.
Veröffentlicht: (2025)
von: Yan, Han, et al.
Veröffentlicht: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
von: Zheng, Xiangqing, et al.
Veröffentlicht: (2025)
von: Zheng, Xiangqing, et al.
Veröffentlicht: (2025)
RAIN: Real-time Animation of Infinite Video Stream
von: Shu, Zhilei, et al.
Veröffentlicht: (2024)
von: Shu, Zhilei, et al.
Veröffentlicht: (2024)
Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAE
von: Yang, Yiying, et al.
Veröffentlicht: (2024)
von: Yang, Yiying, et al.
Veröffentlicht: (2024)
DRFusion: Drift-Resilient Temporally Consistent Infrared-Visible Video Fusion
von: Li, Xingyuan, et al.
Veröffentlicht: (2026)
von: Li, Xingyuan, et al.
Veröffentlicht: (2026)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
von: Yang, Ying, et al.
Veröffentlicht: (2026)
von: Yang, Ying, et al.
Veröffentlicht: (2026)
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
von: Zhang, Guiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Guiyu, et al.
Veröffentlicht: (2025)
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
von: Song, Yuchen, et al.
Veröffentlicht: (2025)
von: Song, Yuchen, et al.
Veröffentlicht: (2025)
Omni-Video: Democratizing Unified Video Understanding and Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
von: Ouyang, Hao, et al.
Veröffentlicht: (2023)
von: Ouyang, Hao, et al.
Veröffentlicht: (2023)
A Survey on Long Video Generation: Challenges, Methods, and Prospects
von: Li, Chengxuan, et al.
Veröffentlicht: (2024)
von: Li, Chengxuan, et al.
Veröffentlicht: (2024)
Reducing Annotation Burden: Exploiting Image Knowledge for Few-Shot Medical Video Object Segmentation via Spatiotemporal Consistency Relearning
von: Zheng, Zixuan, et al.
Veröffentlicht: (2025)
von: Zheng, Zixuan, et al.
Veröffentlicht: (2025)
PADS: Plug-and-Play 3D Human Pose Analysis via Diffusion Generative Modeling
von: Ji, Haorui, et al.
Veröffentlicht: (2024)
von: Ji, Haorui, et al.
Veröffentlicht: (2024)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
Consistency-Preserving Diverse Video Generation
von: Liu, Xinshuang, et al.
Veröffentlicht: (2026)
von: Liu, Xinshuang, et al.
Veröffentlicht: (2026)
A Survey of AI-Generated Video Evaluation
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
Unsupervised Stereo via Multi-Baseline Geometry-Consistent Self-Training
von: Xu, Peng, et al.
Veröffentlicht: (2025)
von: Xu, Peng, et al.
Veröffentlicht: (2025)
3AM: 3egment Anything with Geometric Consistency in Videos
von: Sun, Yang-Che, et al.
Veröffentlicht: (2026)
von: Sun, Yang-Che, et al.
Veröffentlicht: (2026)
A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
von: Elmoghany, Mohamed, et al.
Veröffentlicht: (2025)
von: Elmoghany, Mohamed, et al.
Veröffentlicht: (2025)
Latent Spatiotemporal Adaptation for Generalized Face Forgery Video Detection
von: Zhang, Daichi, et al.
Veröffentlicht: (2023)
von: Zhang, Daichi, et al.
Veröffentlicht: (2023)
TIE: Time Interval Encoding for Video Generation over Events
von: Shu, Zhilei, et al.
Veröffentlicht: (2026)
von: Shu, Zhilei, et al.
Veröffentlicht: (2026)
Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)
Gloria: Consistent Character Video Generation via Content Anchors
von: Yang, Yuhang, et al.
Veröffentlicht: (2026)
von: Yang, Yuhang, et al.
Veröffentlicht: (2026)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
von: Jiang, Xixi, et al.
Veröffentlicht: (2025)
von: Jiang, Xixi, et al.
Veröffentlicht: (2025)
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
JADE: Joint-aware Latent Diffusion for 3D Human Generative Modeling
von: Ji, Haorui, et al.
Veröffentlicht: (2024)
von: Ji, Haorui, et al.
Veröffentlicht: (2024)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
von: Wang, Wenjing, et al.
Veröffentlicht: (2023)
von: Wang, Wenjing, et al.
Veröffentlicht: (2023)
Human Motion Video Generation: A Survey
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation
von: Chen, Hanlin, et al.
Veröffentlicht: (2026)
von: Chen, Hanlin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026) -
Beyond Rigid: Benchmarking Non-Rigid Video Editing
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026) -
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024) -
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
von: Yin, Zhiyu, et al.
Veröffentlicht: (2026) -
Mitigating Multimodal Hallucination via Phase-wise Self-reward
von: Zhang, Yu, et al.
Veröffentlicht: (2026)