ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Ruonan, Tan, Zhenxiong, Chen, Zigeng, Liu, Songhua, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
SpotEdit: Selective Region Editing in Diffusion Transformers
by: Qin, Zhibin, et al.
Published: (2025)
by: Qin, Zhibin, et al.
Published: (2025)
Ultra-Resolution Adaptation with Ease
by: Yu, Ruonan, et al.
Published: (2025)
by: Yu, Ruonan, et al.
Published: (2025)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Heavy Labels Out! Dataset Distillation with Label Space Lightening
by: Yu, Ruonan, et al.
Published: (2024)
by: Yu, Ruonan, et al.
Published: (2024)
OminiControl2: Efficient Conditioning for Diffusion Transformers
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
Image Editing As Programs with Diffusion Models
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
by: Liu, Yuhe, et al.
Published: (2026)
by: Liu, Yuhe, et al.
Published: (2026)
OminiControl: Minimal and Universal Control for Diffusion Transformer
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
LinFusion: 1 GPU, 1 Minute, 16K Image
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
MindBridge: A Cross-Subject Brain Decoding Framework
by: Wang, Shizun, et al.
Published: (2024)
by: Wang, Shizun, et al.
Published: (2024)
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
by: Wu, Yunfeng, et al.
Published: (2025)
by: Wu, Yunfeng, et al.
Published: (2025)
Teddy: Efficient Large-Scale Dataset Distillation via Taylor-Approximated Matching
by: Yu, Ruonan, et al.
Published: (2024)
by: Yu, Ruonan, et al.
Published: (2024)
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
by: Zhou, Letian, et al.
Published: (2025)
by: Zhou, Letian, et al.
Published: (2025)
Distilled Datamodel with Reverse Gradient Matching
by: Ye, Jingwen, et al.
Published: (2024)
by: Ye, Jingwen, et al.
Published: (2024)
Minute-Long Videos with Dual Parallelisms
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
by: Wu, Yunfeng, et al.
Published: (2026)
by: Wu, Yunfeng, et al.
Published: (2026)
SlimSAM: 0.1% Data Makes Segment Anything Slim
by: Chen, Zigeng, et al.
Published: (2023)
by: Chen, Zigeng, et al.
Published: (2023)
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
by: Zuo, Yi, et al.
Published: (2024)
by: Zuo, Yi, et al.
Published: (2024)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
FreeTuner: Any Subject in Any Style with Training-free Diffusion
by: Xu, Youcan, et al.
Published: (2024)
by: Xu, Youcan, et al.
Published: (2024)
V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes
by: Zhang, Yanming, et al.
Published: (2025)
by: Zhang, Yanming, et al.
Published: (2025)
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
One-shot Federated Learning via Synthetic Distiller-Distillate Communication
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
Understanding Dataset Distillation via Spectral Filtering
by: Bo, Deyu, et al.
Published: (2025)
by: Bo, Deyu, et al.
Published: (2025)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
Edit Temporal-Consistent Videos with Image Diffusion Model
by: Wang, Yuanzhi, et al.
Published: (2023)
by: Wang, Yuanzhi, et al.
Published: (2023)
Edit-Your-Interest: Efficient Video Editing via Feature Most-Similar Propagation
by: Zuo, Yi, et al.
Published: (2025)
by: Zuo, Yi, et al.
Published: (2025)
Control and Realism: Best of Both Worlds in Layout-to-Image without Training
by: Li, Bonan, et al.
Published: (2025)
by: Li, Bonan, et al.
Published: (2025)
Flash Sculptor: Modular 3D Worlds from Objects
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
Compositional Video Generation as Flow Equalization
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
Sparse-Dense Side-Tuner for efficient Video Temporal Grounding
by: Pujol-Perich, David, et al.
Published: (2025)
by: Pujol-Perich, David, et al.
Published: (2025)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
by: li, Bonan, et al.
Published: (2025)
by: li, Bonan, et al.
Published: (2025)
MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation
by: Zheng, Longtao, et al.
Published: (2024)
by: Zheng, Longtao, et al.
Published: (2024)
JointTuner: Appearance-Motion Adaptive Joint Training for Customized Video Generation
by: Chen, Fangda, et al.
Published: (2025)
by: Chen, Fangda, et al.
Published: (2025)
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
by: Tang, Siao, et al.
Published: (2026)
by: Tang, Siao, et al.
Published: (2026)
Similar Items
-
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
by: Liu, Songhua, et al.
Published: (2024) -
SpotEdit: Selective Region Editing in Diffusion Transformers
by: Qin, Zhibin, et al.
Published: (2025) -
Ultra-Resolution Adaptation with Ease
by: Yu, Ruonan, et al.
Published: (2025) -
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024) -
Heavy Labels Out! Dataset Distillation with Label Space Lightening
by: Yu, Ruonan, et al.
Published: (2024)