Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zheyuan, Wang, Junyan, Duan, Zicheng, Rodriguez-Opazo, Cristian, Hengel, Anton van den |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026)
by: Xia, Jiatong, et al.
Published: (2026)
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
by: Zhang, Frederic Z., et al.
Published: (2024)
by: Zhang, Frederic Z., et al.
Published: (2024)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024)
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024)
RandLoRA: Full-rank parameter-efficient fine-tuning of large models
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
by: Hwang, Sungwon, et al.
Published: (2025)
by: Hwang, Sungwon, et al.
Published: (2025)
Source-Free Unsupervised Domain Adaptation with Hypothesis Consolidation of Prediction Rationale
by: Shu, Yangyang, et al.
Published: (2024)
by: Shu, Yangyang, et al.
Published: (2024)
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
by: Zhu, Tianyi, et al.
Published: (2024)
by: Zhu, Tianyi, et al.
Published: (2024)
Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
by: Lu, Haodong, et al.
Published: (2025)
by: Lu, Haodong, et al.
Published: (2025)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
by: Tian, Mingkai, et al.
Published: (2025)
by: Tian, Mingkai, et al.
Published: (2025)
Hierarchical Process Reward Models are Symbolic Vision Learners
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
by: Yin, Wei, et al.
Published: (2022)
by: Yin, Wei, et al.
Published: (2022)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
Premonition: Using Generative Models to Preempt Future Data Changes in Continual Learning
by: McDonnell, Mark D., et al.
Published: (2024)
by: McDonnell, Mark D., et al.
Published: (2024)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Fine Tuning Text-to-Image Diffusion Models for Correcting Anomalous Images
by: Yoo, Hyunwoo
Published: (2024)
by: Yoo, Hyunwoo
Published: (2024)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
by: Rahman, Aimon, et al.
Published: (2024)
by: Rahman, Aimon, et al.
Published: (2024)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
by: Phunyaphibarn, Prin, et al.
Published: (2025)
by: Phunyaphibarn, Prin, et al.
Published: (2025)
ViewFusion: Towards Multi-View Consistency via Interpolated Denoising
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
Raccoon: Multi-stage Diffusion Training with Coarse-to-Fine Curating Videos
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
RanPAC: Random Projections and Pre-trained Models for Continual Learning
by: McDonnell, Mark D., et al.
Published: (2023)
by: McDonnell, Mark D., et al.
Published: (2023)
Membership Inference Attacks for Face Images Against Fine-Tuned Latent Diffusion Models
by: Holme, Lauritz Christian, et al.
Published: (2025)
by: Holme, Lauritz Christian, et al.
Published: (2025)
SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
by: Ye, Xi, et al.
Published: (2026)
by: Ye, Xi, et al.
Published: (2026)
Fine-Tuning Video-Text Contrastive Model for Primate Behavior Retrieval from Unlabeled Raw Videos
by: Santo, Giulio Cesare Mastrocinque, et al.
Published: (2025)
by: Santo, Giulio Cesare Mastrocinque, et al.
Published: (2025)
ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
Fine-Tuning Text-To-Image Diffusion Models for Class-Wise Spurious Feature Generation
by: MaungMaung, AprilPyone, et al.
Published: (2024)
by: MaungMaung, AprilPyone, et al.
Published: (2024)
Detecting Dataset Abuse in Fine-Tuning Stable Diffusion Models for Text-to-Image Synthesis
by: Wang, Songrui, et al.
Published: (2024)
by: Wang, Songrui, et al.
Published: (2024)
ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation
by: Yang, Shaoshu, et al.
Published: (2024)
by: Yang, Shaoshu, et al.
Published: (2024)
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025)
by: Kim, Kinam, et al.
Published: (2025)
An empirical study of the effect of video encoders on Temporal Video Grounding
by: De la Jara, Ignacio M., et al.
Published: (2025)
by: De la Jara, Ignacio M., et al.
Published: (2025)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models
by: Wang, Zilan, et al.
Published: (2024)
by: Wang, Zilan, et al.
Published: (2024)
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models
by: Cai, Yufei, et al.
Published: (2025)
by: Cai, Yufei, et al.
Published: (2025)
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
by: Yin, Bo, et al.
Published: (2025)
by: Yin, Bo, et al.
Published: (2025)
DAPE: Dual-Stage Parameter-Efficient Fine-Tuning for Consistent Video Editing with Diffusion Models
by: Xia, Junhao, et al.
Published: (2025)
by: Xia, Junhao, et al.
Published: (2025)
Similar Items
-
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025) -
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
by: Xia, Jiatong, et al.
Published: (2026) -
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
by: Zhang, Frederic Z., et al.
Published: (2024) -
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
by: Rodriguez-Opazo, Cristian, et al.
Published: (2024) -
RandLoRA: Full-rank parameter-efficient fine-tuning of large models
by: Albert, Paul, et al.
Published: (2025)