ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Weiqi, Zhang, Zehao, Lin, Liang, Wang, Guangrun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
GS: Generative Segmentation via Label Diffusion
by: Chen, Yuhao, et al.
Published: (2025)
by: Chen, Yuhao, et al.
Published: (2025)
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
by: Liu, Jinxi, et al.
Published: (2025)
by: Liu, Jinxi, et al.
Published: (2025)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
DNA Family: Boosting Weight-Sharing NAS with Block-Wise Supervisions
by: Wang, Guangrun, et al.
Published: (2024)
by: Wang, Guangrun, et al.
Published: (2024)
In-Situ Tweedie Discrete Diffusion Models
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment
by: Xu, Yuanfeng, et al.
Published: (2024)
by: Xu, Yuanfeng, et al.
Published: (2024)
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
by: He, Zijian, et al.
Published: (2025)
by: He, Zijian, et al.
Published: (2025)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
SDXS: Real-Time One-Step Latent Diffusion Models with Image Conditions
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Controllable Weather Synthesis and Removal with Video Diffusion Models
by: Lin, Chih-Hao, et al.
Published: (2025)
by: Lin, Chih-Hao, et al.
Published: (2025)
Re-Attentional Controllable Video Diffusion Editing
by: Wang, Yuanzhi, et al.
Published: (2024)
by: Wang, Yuanzhi, et al.
Published: (2024)
Geometry aware 3D generation from in-the-wild images in ImageNet
by: Shen, Qijia, et al.
Published: (2024)
by: Shen, Qijia, et al.
Published: (2024)
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution
by: Chen, Bin, et al.
Published: (2026)
by: Chen, Bin, et al.
Published: (2026)
Blended Latent Diffusion under Attention Control for Real-World Video Editing
by: Liu, Deyin, et al.
Published: (2024)
by: Liu, Deyin, et al.
Published: (2024)
Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention Learning
by: Li, Maomao, et al.
Published: (2025)
by: Li, Maomao, et al.
Published: (2025)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
by: Ma, Junxian, et al.
Published: (2025)
by: Ma, Junxian, et al.
Published: (2025)
Understanding Attention Mechanism in Video Diffusion Models
by: Liu, Bingyan, et al.
Published: (2025)
by: Liu, Bingyan, et al.
Published: (2025)
DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
by: Long, Yongji, et al.
Published: (2026)
by: Long, Yongji, et al.
Published: (2026)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
by: Peng, Liang, et al.
Published: (2025)
by: Peng, Liang, et al.
Published: (2025)
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026)
by: Chen, Tianshui, et al.
Published: (2026)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
by: Li, Chunyu, et al.
Published: (2024)
by: Li, Chunyu, et al.
Published: (2024)
Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion
by: Dominici, Edoardo A., et al.
Published: (2026)
by: Dominici, Edoardo A., et al.
Published: (2026)
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
by: Li, Weiqi, et al.
Published: (2024)
by: Li, Weiqi, et al.
Published: (2024)
CCEdit: Creative and Controllable Video Editing via Diffusion Models
by: Feng, Ruoyu, et al.
Published: (2023)
by: Feng, Ruoyu, et al.
Published: (2023)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
by: Zhang, Ruixin, et al.
Published: (2025)
by: Zhang, Ruixin, et al.
Published: (2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Bidirectional Sparse Attention for Faster Video Diffusion Training
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
Retinex-Diffusion: On Controlling Illumination Conditions in Diffusion Models via Retinex Theory
by: Xing, Xiaoyan, et al.
Published: (2024)
by: Xing, Xiaoyan, et al.
Published: (2024)
Noise Controlled CT Super-Resolution with Conditional Diffusion Model
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023)
by: Li, Pengxiang, et al.
Published: (2023)
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
Making Large Language Models Better Planners with Reasoning-Decision Alignment
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
Similar Items
-
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
by: Li, Weiqi, et al.
Published: (2025) -
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024) -
GS: Generative Segmentation via Label Diffusion
by: Chen, Yuhao, et al.
Published: (2025) -
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
by: Liu, Jinxi, et al.
Published: (2025) -
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)