Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Maduabuchi, Chika, Chen, Hao, Han, Yujin, Wang, Jindong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Event-Driven Video Generation
by: Maduabuchi, Chika
Published: (2026)
by: Maduabuchi, Chika
Published: (2026)
Temporal Pair Consistency for Variance-Reduced Flow Matching
by: Maduabuchi, Chika, et al.
Published: (2026)
by: Maduabuchi, Chika, et al.
Published: (2026)
VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
by: Maduabuchi, Chika, et al.
Published: (2024)
by: Maduabuchi, Chika, et al.
Published: (2024)
Entropy-Controlled Flow Matching
by: Maduabuchi, Chika
Published: (2026)
by: Maduabuchi, Chika
Published: (2026)
MSEG-VCUQ: Multimodal SEGmentation with Enhanced Vision Foundation Models, Convolutional Neural Networks, and Uncertainty Quantification for High-Speed Video Phase Detection Data
by: Maduabuchi, Chika, et al.
Published: (2024)
by: Maduabuchi, Chika, et al.
Published: (2024)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Multi-event Video-Text Retrieval
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
by: Bao, Fan, et al.
Published: (2024)
by: Bao, Fan, et al.
Published: (2024)
Unlearning Concepts from Text-to-Video Diffusion Models
by: Liu, Shiqi, et al.
Published: (2024)
by: Liu, Shiqi, et al.
Published: (2024)
LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
by: Varshney, Payal, et al.
Published: (2025)
by: Varshney, Payal, et al.
Published: (2025)
MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation
by: Yu, Sihyun, et al.
Published: (2025)
by: Yu, Sihyun, et al.
Published: (2025)
Expanding the Role of Diffusion Models for Robust Classifier Training
by: Huang, Pin-Han, et al.
Published: (2026)
by: Huang, Pin-Han, et al.
Published: (2026)
Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Leveraging Programmatically Generated Synthetic Data for Differentially Private Diffusion Training
by: Choi, Yujin, et al.
Published: (2024)
by: Choi, Yujin, et al.
Published: (2024)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
Adaptive Hybrid Caching for Efficient Text-to-Video Diffusion Model Acceleration
by: Wei, Yuanxin, et al.
Published: (2025)
by: Wei, Yuanxin, et al.
Published: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Test-Time Training on Video Streams
by: Wang, Renhao, et al.
Published: (2023)
by: Wang, Renhao, et al.
Published: (2023)
LVSA: Training-Free Sparse Attention for Long Video Diffusion
by: Glorian, Gael, et al.
Published: (2026)
by: Glorian, Gael, et al.
Published: (2026)
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance
by: Bompai, Stelio, et al.
Published: (2026)
by: Bompai, Stelio, et al.
Published: (2026)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
by: Cao, Cong, et al.
Published: (2024)
by: Cao, Cong, et al.
Published: (2024)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
SurGen: Text-Guided Diffusion Model for Surgical Video Generation
by: Cho, Joseph, et al.
Published: (2024)
by: Cho, Joseph, et al.
Published: (2024)
GSURE-Based Diffusion Model Training with Corrupted Data
by: Kawar, Bahjat, et al.
Published: (2023)
by: Kawar, Bahjat, et al.
Published: (2023)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
by: Lin, Han, et al.
Published: (2024)
by: Lin, Han, et al.
Published: (2024)
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
by: Zhang, Shuhai, et al.
Published: (2025)
by: Zhang, Shuhai, et al.
Published: (2025)
Video Diffusion Models: A Survey
by: Melnik, Andrew, et al.
Published: (2024)
by: Melnik, Andrew, et al.
Published: (2024)
Diffusion Adversarial Post-Training for One-Step Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
by: Gupta, Samarth, et al.
Published: (2025)
by: Gupta, Samarth, et al.
Published: (2025)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
by: Romero, David, et al.
Published: (2025)
by: Romero, David, et al.
Published: (2025)
A Survey on Video Diffusion Models
by: Xing, Zhen, et al.
Published: (2023)
by: Xing, Zhen, et al.
Published: (2023)
Similar Items
-
Event-Driven Video Generation
by: Maduabuchi, Chika
Published: (2026) -
Temporal Pair Consistency for Variance-Reduced Flow Matching
by: Maduabuchi, Chika, et al.
Published: (2026) -
VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
by: Maduabuchi, Chika, et al.
Published: (2024) -
Entropy-Controlled Flow Matching
by: Maduabuchi, Chika
Published: (2026) -
MSEG-VCUQ: Multimodal SEGmentation with Enhanced Vision Foundation Models, Convolutional Neural Networks, and Uncertainty Quantification for High-Speed Video Phase Detection Data
by: Maduabuchi, Chika, et al.
Published: (2024)