Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Yongjia, Chen, Junlin, Di, Donglin, Xie, Qi, Fan, Lei, Chen, Wei, Gou, Xiaofei, Zhao, Na, Yang, Xun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
by: Xie, Qi, et al.
Published: (2025)
by: Xie, Qi, et al.
Published: (2025)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
by: Di, Donglin, et al.
Published: (2024)
by: Di, Donglin, et al.
Published: (2024)
Real Face Video Animation Platform
by: Chen, Xiaokai, et al.
Published: (2024)
by: Chen, Xiaokai, et al.
Published: (2024)
GRPose: Learning Graph Relations for Human Image Generation with Pose Priors
by: Yin, Xiangchen, et al.
Published: (2024)
by: Yin, Xiangchen, et al.
Published: (2024)
CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning
by: Feng, He, et al.
Published: (2026)
by: Feng, He, et al.
Published: (2026)
QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation
by: Yang, Jiahui, et al.
Published: (2025)
by: Yang, Jiahui, et al.
Published: (2025)
One-Shot Pose-Driving Face Animation Platform
by: Feng, He, et al.
Published: (2024)
by: Feng, He, et al.
Published: (2024)
TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian Splatting Manipulation
by: Luo, Chaofan, et al.
Published: (2024)
by: Luo, Chaofan, et al.
Published: (2024)
UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation
by: Sun, Wenzhang, et al.
Published: (2025)
by: Sun, Wenzhang, et al.
Published: (2025)
Adams Bashforth Moulton Solver for Inversion and Editing in Rectified Flow
by: Ma, Yongjia, et al.
Published: (2025)
by: Ma, Yongjia, et al.
Published: (2025)
DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation
by: Feng, He, et al.
Published: (2025)
by: Feng, He, et al.
Published: (2025)
EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
by: Ling, Haotian, et al.
Published: (2025)
by: Ling, Haotian, et al.
Published: (2025)
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
Global-Local Aware Scene Text Editing
by: Yang, Fuxiang, et al.
Published: (2025)
by: Yang, Fuxiang, et al.
Published: (2025)
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion
by: Chen, Jingyuan, et al.
Published: (2025)
by: Chen, Jingyuan, et al.
Published: (2025)
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
by: Qiu, Haonan, et al.
Published: (2024)
by: Qiu, Haonan, et al.
Published: (2024)
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
by: Yu, Haodong, et al.
Published: (2026)
by: Yu, Haodong, et al.
Published: (2026)
Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion
by: Liu, Huaize, et al.
Published: (2025)
by: Liu, Huaize, et al.
Published: (2025)
Discriminator-Free Direct Preference Optimization for Video Diffusion
by: Cheng, Haoran, et al.
Published: (2025)
by: Cheng, Haoran, et al.
Published: (2025)
Long Context Tuning for Video Generation
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
CoNo: Consistency Noise Injection for Tuning-free Long Video Diffusion
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
PhyRPR: Training-Free Physics-Constrained Video Generation
by: Zhao, Yibo, et al.
Published: (2026)
by: Zhao, Yibo, et al.
Published: (2026)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
by: Qiu, Di, et al.
Published: (2024)
by: Qiu, Di, et al.
Published: (2024)
Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
by: Di, Donglin, et al.
Published: (2024)
by: Di, Donglin, et al.
Published: (2024)
A Self-supervised Motion Representation for Portrait Video Generation
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
Comparative Study of Neighbor-based Methods for Local Outlier Detection
by: Qi, Zhuang, et al.
Published: (2024)
by: Qi, Zhuang, et al.
Published: (2024)
Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection
by: Peng, Xingyu, et al.
Published: (2024)
by: Peng, Xingyu, et al.
Published: (2024)
FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction
by: Chen, Fangda, et al.
Published: (2026)
by: Chen, Fangda, et al.
Published: (2026)
GLRD: Global-Local Collaborative Reason and Debate with PSL for 3D Open-Vocabulary Detection
by: Peng, Xingyu, et al.
Published: (2025)
by: Peng, Xingyu, et al.
Published: (2025)
ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
by: Li, Zongyi, et al.
Published: (2024)
by: Li, Zongyi, et al.
Published: (2024)
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
by: Chen, Nan, et al.
Published: (2025)
by: Chen, Nan, et al.
Published: (2025)
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
by: Xue, Haotian, et al.
Published: (2025)
by: Xue, Haotian, et al.
Published: (2025)
LongLive: Real-time Interactive Long Video Generation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
by: Fan, Shixuan, et al.
Published: (2024)
by: Fan, Shixuan, et al.
Published: (2024)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Similar Items
-
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
by: Xie, Qi, et al.
Published: (2025) -
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
by: Di, Donglin, et al.
Published: (2024) -
Real Face Video Animation Platform
by: Chen, Xiaokai, et al.
Published: (2024) -
GRPose: Learning Graph Relations for Human Image Generation with Pose Priors
by: Yin, Xiangchen, et al.
Published: (2024) -
CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning
by: Feng, He, et al.
Published: (2026)