CascadeV: An Implementation of Wurstchen Architecture for Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Wenfeng, Wei, Jiangchuan, Liu, Boyuan, Zhang, Yichen, Yan, Shiyue, Guo, Mingyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025)
by: Wei, Jiangchuan, et al.
Published: (2025)
ContentV: Efficient Training of Video Generation Models with Limited Compute
by: Lin, Wenfeng, et al.
Published: (2025)
by: Lin, Wenfeng, et al.
Published: (2025)
Towards Self-Improvement of Diffusion Models via Group Preference Optimization
by: Chen, Renjie, et al.
Published: (2025)
by: Chen, Renjie, et al.
Published: (2025)
CascadeV-Det: Cascade Point Voting for 3D Object Detection
by: Liang, Yingping, et al.
Published: (2024)
by: Liang, Yingping, et al.
Published: (2024)
LVFace: Progressive Cluster Optimization for Large Vision Models in Face Recognition
by: You, Jinghan, et al.
Published: (2025)
by: You, Jinghan, et al.
Published: (2025)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
MultiModal Action Conditioned Video Generation
by: Li, Yichen, et al.
Published: (2025)
by: Li, Yichen, et al.
Published: (2025)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
UniVBench: Towards Unified Evaluation for Video Foundation Models
by: Wei, Jianhui, et al.
Published: (2026)
by: Wei, Jianhui, et al.
Published: (2026)
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
by: Jin, Hongbo, et al.
Published: (2025)
by: Jin, Hongbo, et al.
Published: (2025)
COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation
by: Zhang, Xinlong, et al.
Published: (2026)
by: Zhang, Xinlong, et al.
Published: (2026)
Generating Human Motion Videos using a Cascaded Text-to-Video Framework
by: Nam, Hyelin, et al.
Published: (2025)
by: Nam, Hyelin, et al.
Published: (2025)
Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
by: Pan, Yueming, et al.
Published: (2025)
by: Pan, Yueming, et al.
Published: (2025)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
by: Yang, Kaixing, et al.
Published: (2025)
by: Yang, Kaixing, et al.
Published: (2025)
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
by: Chen, Yunuo, et al.
Published: (2026)
by: Chen, Yunuo, et al.
Published: (2026)
GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections
by: Zhang, Shiyue, et al.
Published: (2024)
by: Zhang, Shiyue, et al.
Published: (2024)
Generative Video Matting
by: Ge, Yongtao, et al.
Published: (2025)
by: Ge, Yongtao, et al.
Published: (2025)
UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving
by: Chen, Rui, et al.
Published: (2024)
by: Chen, Rui, et al.
Published: (2024)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
by: SII-GAIR, et al.
Published: (2026)
by: SII-GAIR, et al.
Published: (2026)
LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
How Far Are Video Models from True Multimodal Reasoning?
by: Zhang, Xiaotian, et al.
Published: (2026)
by: Zhang, Xiaotian, et al.
Published: (2026)
StableV2V: Stablizing Shape Consistency in Video-to-Video Editing
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
FlexiFilm: Long Video Generation with Flexible Conditions
by: Ouyang, Yichen, et al.
Published: (2024)
by: Ouyang, Yichen, et al.
Published: (2024)
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
by: Wang, Weimin, et al.
Published: (2024)
by: Wang, Weimin, et al.
Published: (2024)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
Multi-Garment Customized Model Generation
by: Liu, Yichen, et al.
Published: (2024)
by: Liu, Yichen, et al.
Published: (2024)
Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
by: He, Muyang, et al.
Published: (2026)
by: He, Muyang, et al.
Published: (2026)
Corner2Net: Detecting Objects as Cascade Corners
by: Liu, Chenglong, et al.
Published: (2024)
by: Liu, Chenglong, et al.
Published: (2024)
RoadBEV: Road Surface Reconstruction in Bird's Eye View
by: Zhao, Tong, et al.
Published: (2024)
by: Zhao, Tong, et al.
Published: (2024)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
by: Ni, Jingcheng, et al.
Published: (2025)
by: Ni, Jingcheng, et al.
Published: (2025)
InstanceV: Instance-Level Video Generation
by: Chen, Yuheng, et al.
Published: (2025)
by: Chen, Yuheng, et al.
Published: (2025)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
by: Wu, Weijia, et al.
Published: (2024)
by: Wu, Weijia, et al.
Published: (2024)
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
by: He, Zhixian, et al.
Published: (2024)
by: He, Zhixian, et al.
Published: (2024)
Cascading Refinement Video Denoising with Uncertainty Adaptivity
by: Yu, Xinyuan
Published: (2024)
by: Yu, Xinyuan
Published: (2024)
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Similar Items
-
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
by: Wei, Jiangchuan, et al.
Published: (2025) -
ContentV: Efficient Training of Video Generation Models with Limited Compute
by: Lin, Wenfeng, et al.
Published: (2025) -
Towards Self-Improvement of Diffusion Models via Group Preference Optimization
by: Chen, Renjie, et al.
Published: (2025) -
CascadeV-Det: Cascade Point Voting for 3D Object Detection
by: Liang, Yingping, et al.
Published: (2024) -
LVFace: Progressive Cluster Optimization for Large Vision Models in Face Recognition
by: You, Jinghan, et al.
Published: (2025)