FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Jiasong, Ma, Ao, Wang, Jing, Cao, Ke, Zhang, Zhanjie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
by: Ma, Ao, et al.
Published: (2025)
by: Ma, Ao, et al.
Published: (2025)
FedCVU: Federated Learning for Cross-View Video Understanding
by: Zhang, Shenghan, et al.
Published: (2026)
by: Zhang, Shenghan, et al.
Published: (2026)
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Consistent Video Colorization via Palette Guidance
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
VideoMemory: Toward Consistent Video Generation via Memory Integration
by: Zhou, Jinsong, et al.
Published: (2026)
by: Zhou, Jinsong, et al.
Published: (2026)
Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark
by: Cao, Ke, et al.
Published: (2026)
by: Cao, Ke, et al.
Published: (2026)
RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
by: Cao, Ke, et al.
Published: (2025)
by: Cao, Ke, et al.
Published: (2025)
RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature Guidance
by: Fan, Jiaojiao, et al.
Published: (2024)
by: Fan, Jiaojiao, et al.
Published: (2024)
MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
by: Ling, Run, et al.
Published: (2025)
by: Ling, Run, et al.
Published: (2025)
U-StyDiT: Ultra-high Quality Artistic Style Transfer Using Diffusion Transformers
by: Zhang, Zhanjie, et al.
Published: (2025)
by: Zhang, Zhanjie, et al.
Published: (2025)
Denoising Reuse: Exploiting Inter-frame Motion Consistency for Efficient Video Latent Generation
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
PiercingEye: Dual-Space Video Violence Detection with Hyperbolic Vision-Language Guidance
by: Leng, Jiaxu, et al.
Published: (2025)
by: Leng, Jiaxu, et al.
Published: (2025)
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
by: Feng, Sicong, et al.
Published: (2025)
by: Feng, Sicong, et al.
Published: (2025)
BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
by: Hu, Zhanjie, et al.
Published: (2026)
by: Hu, Zhanjie, et al.
Published: (2026)
Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task
by: Wang, Jing, et al.
Published: (2024)
by: Wang, Jing, et al.
Published: (2024)
Gloria: Consistent Character Video Generation via Content Anchors
by: Yang, Yuhang, et al.
Published: (2026)
by: Yang, Yuhang, et al.
Published: (2026)
NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
by: Yuan, Yu, et al.
Published: (2025)
by: Yuan, Yu, et al.
Published: (2025)
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Object-Aware Video Matting with Cross-Frame Guidance
by: Zhang, Huayu, et al.
Published: (2025)
by: Zhang, Huayu, et al.
Published: (2025)
Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
by: Zeng, Jianshu, et al.
Published: (2025)
by: Zeng, Jianshu, et al.
Published: (2025)
VideoAuteur: Towards Long Narrative Video Generation
by: Xiao, Junfei, et al.
Published: (2025)
by: Xiao, Junfei, et al.
Published: (2025)
Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors
by: Wang, Rong, et al.
Published: (2026)
by: Wang, Rong, et al.
Published: (2026)
Compositional Video Generation via Inference-Time Guidance
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
by: Fu, Junhu, et al.
Published: (2026)
by: Fu, Junhu, et al.
Published: (2026)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
by: Zhu, Tianrui, et al.
Published: (2025)
by: Zhu, Tianrui, et al.
Published: (2025)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
by: Chu, Ruihang, et al.
Published: (2025)
by: Chu, Ruihang, et al.
Published: (2025)
Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance
by: Zhang, Qing, et al.
Published: (2024)
by: Zhang, Qing, et al.
Published: (2024)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
by: Yang, Ying, et al.
Published: (2026)
by: Yang, Ying, et al.
Published: (2026)
NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration
by: Dong, Haotian, et al.
Published: (2025)
by: Dong, Haotian, et al.
Published: (2025)
InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
by: Zhang, Ziqing, et al.
Published: (2025)
by: Zhang, Ziqing, et al.
Published: (2025)
Video2LoRA: Unified Semantic-Controlled Video Generation via Per-Reference-Video LoRA
by: Wu, Zexi, et al.
Published: (2026)
by: Wu, Zexi, et al.
Published: (2026)
Match-Stereo-Videos: Bidirectional Alignment for Consistent Dynamic Stereo Matching
by: Jing, Junpeng, et al.
Published: (2024)
by: Jing, Junpeng, et al.
Published: (2024)
Video Generation with Consistency Tuning
by: Wang, Chaoyi, et al.
Published: (2024)
by: Wang, Chaoyi, et al.
Published: (2024)
Detecting AI-Generated Video via Frame Consistency
by: Ma, Long, et al.
Published: (2024)
by: Ma, Long, et al.
Published: (2024)
Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
by: Aoshima, Takehiro, et al.
Published: (2025)
by: Aoshima, Takehiro, et al.
Published: (2025)
AnimateAnything: Consistent and Controllable Animation for Video Generation
by: Lei, Guojun, et al.
Published: (2024)
by: Lei, Guojun, et al.
Published: (2024)
MoCha:End-to-End Video Character Replacement without Structural Guidance
by: Xu, Zhengbo, et al.
Published: (2026)
by: Xu, Zhengbo, et al.
Published: (2026)
Similar Items
-
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
by: Ma, Ao, et al.
Published: (2025) -
FedCVU: Federated Learning for Cross-View Video Understanding
by: Zhang, Shenghan, et al.
Published: (2026) -
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
by: Wang, Jing, et al.
Published: (2025) -
Consistent Video Colorization via Palette Guidance
by: Wang, Han, et al.
Published: (2025) -
VideoMemory: Toward Consistent Video Generation via Memory Integration
by: Zhou, Jinsong, et al.
Published: (2026)