Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Dohun, Jeong, Hyeonho, Kim, Jiwook, Ceylan, Duygu, Ye, Jong Chul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
di: Lee, Dohun, et al.
Pubblicazione: (2026)
di: Lee, Dohun, et al.
Pubblicazione: (2026)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
di: Jeong, Hyeonho, et al.
Pubblicazione: (2024)
di: Jeong, Hyeonho, et al.
Pubblicazione: (2024)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
di: Jeong, Hyeonho, et al.
Pubblicazione: (2023)
di: Jeong, Hyeonho, et al.
Pubblicazione: (2023)
VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
di: Lee, Dohun, et al.
Pubblicazione: (2024)
di: Lee, Dohun, et al.
Pubblicazione: (2024)
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
di: Jeong, Hyeonho, et al.
Pubblicazione: (2025)
di: Jeong, Hyeonho, et al.
Pubblicazione: (2025)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models
di: Park, Geon Yeong, et al.
Pubblicazione: (2024)
di: Park, Geon Yeong, et al.
Pubblicazione: (2024)
TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video Generation
di: Kwon, Gihyun, et al.
Pubblicazione: (2024)
di: Kwon, Gihyun, et al.
Pubblicazione: (2024)
DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing
di: Jeong, Hyeonho, et al.
Pubblicazione: (2024)
di: Jeong, Hyeonho, et al.
Pubblicazione: (2024)
ACDC: Autoregressive Coherent Multimodal Generation using Diffusion Correction
di: Chung, Hyungjin, et al.
Pubblicazione: (2024)
di: Chung, Hyungjin, et al.
Pubblicazione: (2024)
Optical-Flow Guided Prompt Optimization for Coherent Video Generation
di: Nam, Hyelin, et al.
Pubblicazione: (2024)
di: Nam, Hyelin, et al.
Pubblicazione: (2024)
JOG3R: Towards 3D-Consistent Video Generators
di: Huang, Chun-Hao Paul, et al.
Pubblicazione: (2025)
di: Huang, Chun-Hao Paul, et al.
Pubblicazione: (2025)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
PromptLoop: Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
di: Lee, Suhyeon, et al.
Pubblicazione: (2025)
di: Lee, Suhyeon, et al.
Pubblicazione: (2025)
Representation Alignment for Just Image Transformers is not Easier than You Think
di: Shin, Jaeyo, et al.
Pubblicazione: (2026)
di: Shin, Jaeyo, et al.
Pubblicazione: (2026)
Zero4D: Training-Free 4D Video Generation From Single Video Using Off-the-Shelf Video Diffusion
di: Park, Jangho, et al.
Pubblicazione: (2025)
di: Park, Jangho, et al.
Pubblicazione: (2025)
InvFusion: Bridging Supervised and Zero-shot Diffusion for Inverse Problems
di: Elata, Noam, et al.
Pubblicazione: (2025)
di: Elata, Noam, et al.
Pubblicazione: (2025)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
di: Lee, Seonho, et al.
Pubblicazione: (2024)
di: Lee, Seonho, et al.
Pubblicazione: (2024)
Boosting Camera Motion Control for Video Diffusion Transformers
di: Cheong, Soon Yau, et al.
Pubblicazione: (2024)
di: Cheong, Soon Yau, et al.
Pubblicazione: (2024)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
di: Lee, Jaa-Yeon, et al.
Pubblicazione: (2026)
di: Lee, Jaa-Yeon, et al.
Pubblicazione: (2026)
MindFormer: Semantic Alignment of Multi-Subject fMRI for Brain Decoding
di: Han, Inhwa, et al.
Pubblicazione: (2024)
di: Han, Inhwa, et al.
Pubblicazione: (2024)
Solving Video Inverse Problems Using Image Diffusion Models
di: Kwon, Taesung, et al.
Pubblicazione: (2024)
di: Kwon, Taesung, et al.
Pubblicazione: (2024)
GeoFusionLRM: Geometry-Aware Self-Correction for Consistent 3D Reconstruction
di: Yildirim, Ahmet Burak, et al.
Pubblicazione: (2026)
di: Yildirim, Ahmet Burak, et al.
Pubblicazione: (2026)
Class-Continuous Conditional Generative Neural Radiance Field
di: Kim, Jiwook, et al.
Pubblicazione: (2023)
di: Kim, Jiwook, et al.
Pubblicazione: (2023)
Self-Guided Generation of Minority Samples Using Diffusion Models
di: Um, Soobin, et al.
Pubblicazione: (2024)
di: Um, Soobin, et al.
Pubblicazione: (2024)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
di: Kara, Ozgur, et al.
Pubblicazione: (2025)
di: Kara, Ozgur, et al.
Pubblicazione: (2025)
VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion Models
di: Kwon, Taesung, et al.
Pubblicazione: (2024)
di: Kwon, Taesung, et al.
Pubblicazione: (2024)
ContextMRI: Enhancing Compressed Sensing MRI through Metadata Conditioning
di: Chung, Hyungjin, et al.
Pubblicazione: (2025)
di: Chung, Hyungjin, et al.
Pubblicazione: (2025)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
di: Vandersanden, Jente, et al.
Pubblicazione: (2026)
di: Vandersanden, Jente, et al.
Pubblicazione: (2026)
Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
di: Jeon, Wooseok, et al.
Pubblicazione: (2026)
di: Jeon, Wooseok, et al.
Pubblicazione: (2026)
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
di: Kim, Bryan Sangwoo, et al.
Pubblicazione: (2025)
di: Kim, Bryan Sangwoo, et al.
Pubblicazione: (2025)
RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
di: Lee, Junhee, et al.
Pubblicazione: (2025)
di: Lee, Junhee, et al.
Pubblicazione: (2025)
ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler
di: Yang, Serin, et al.
Pubblicazione: (2024)
di: Yang, Serin, et al.
Pubblicazione: (2024)
MD-ProjTex: Texturing 3D Shapes with Multi-Diffusion Projection
di: Yildirim, Ahmet Burak, et al.
Pubblicazione: (2025)
di: Yildirim, Ahmet Burak, et al.
Pubblicazione: (2025)
Unified Editing of Panorama, 3D Scenes, and Videos Through Disentangled Self-Attention Injection
di: Kwon, Gihyun, et al.
Pubblicazione: (2024)
di: Kwon, Gihyun, et al.
Pubblicazione: (2024)
Contrastive CFG: Improving CFG in Diffusion Models by Contrasting Positive and Negative Concepts
di: Chang, Jinho, et al.
Pubblicazione: (2024)
di: Chang, Jinho, et al.
Pubblicazione: (2024)
Latent Schrodinger Bridge: Prompting Latent Diffusion for Fast Unpaired Image-to-Image Translation
di: Kim, Jeongsol, et al.
Pubblicazione: (2024)
di: Kim, Jeongsol, et al.
Pubblicazione: (2024)
Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
di: Chang, Jinho, et al.
Pubblicazione: (2025)
di: Chang, Jinho, et al.
Pubblicazione: (2025)
Gradient-Free Noise Optimization for Reward Alignment in Generative Models
di: Kim, Jeongsol, et al.
Pubblicazione: (2026)
di: Kim, Jeongsol, et al.
Pubblicazione: (2026)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
di: Lee, Jaeseong, et al.
Pubblicazione: (2024)
di: Lee, Jaeseong, et al.
Pubblicazione: (2024)
Don't Play Favorites: Minority Guidance for Diffusion Models
di: Um, Soobin, et al.
Pubblicazione: (2023)
di: Um, Soobin, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
di: Lee, Dohun, et al.
Pubblicazione: (2026) -
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
di: Jeong, Hyeonho, et al.
Pubblicazione: (2024) -
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
di: Jeong, Hyeonho, et al.
Pubblicazione: (2023) -
VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
di: Lee, Dohun, et al.
Pubblicazione: (2024) -
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
di: Jeong, Hyeonho, et al.
Pubblicazione: (2025)