VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chi, Liang, Yuanzhi, Qiu, Xi, Yi, Fangqiu, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
by: Ni, Ziqi, et al.
Published: (2025)
by: Ni, Ziqi, et al.
Published: (2025)
InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the Wild
by: Ma, Yiyi, et al.
Published: (2025)
by: Ma, Yiyi, et al.
Published: (2025)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
AnimateAnything: Consistent and Controllable Animation for Video Generation
by: Lei, Guojun, et al.
Published: (2024)
by: Lei, Guojun, et al.
Published: (2024)
Spatial-Temporal State Propagation Autoregressive Model for 4D Object Generation
by: Yang, Liying, et al.
Published: (2026)
by: Yang, Liying, et al.
Published: (2026)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024)
by: Lu, Yu, et al.
Published: (2024)
Edit Temporal-Consistent Videos with Image Diffusion Model
by: Wang, Yuanzhi, et al.
Published: (2023)
by: Wang, Yuanzhi, et al.
Published: (2023)
MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control
by: Yao, Yining, et al.
Published: (2024)
by: Yao, Yining, et al.
Published: (2024)
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
by: Liu, Biao, et al.
Published: (2025)
by: Liu, Biao, et al.
Published: (2025)
Reward-Aware Trajectory Shaping for Few-step Visual Generation
by: Li, Rui, et al.
Published: (2026)
by: Li, Rui, et al.
Published: (2026)
World Consistency Score: A Unified Metric for Video Generation Quality
by: Rakheja, Akshat, et al.
Published: (2025)
by: Rakheja, Akshat, et al.
Published: (2025)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Diffusion Models
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
TeleStyle: Content-Preserving Style Transfer in Images and Videos
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
by: Guo, Yanjun, et al.
Published: (2026)
by: Guo, Yanjun, et al.
Published: (2026)
Unified Camera Positional Encoding for Controlled Video Generation
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
Seedance 1.0: Exploring the Boundaries of Video Generation Models
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
by: Wang, Yukun, et al.
Published: (2026)
by: Wang, Yukun, et al.
Published: (2026)
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
by: Qin, Qi, et al.
Published: (2025)
by: Qin, Qi, et al.
Published: (2025)
Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
by: Liang, Yuanzhi, et al.
Published: (2025)
by: Liang, Yuanzhi, et al.
Published: (2025)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
by: Shi, Xiaoyu, et al.
Published: (2024)
by: Shi, Xiaoyu, et al.
Published: (2024)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
by: Ma, Yukuo, et al.
Published: (2025)
by: Ma, Yukuo, et al.
Published: (2025)
Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation
by: Li, Rui, et al.
Published: (2026)
by: Li, Rui, et al.
Published: (2026)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
by: Long, Fuchen, et al.
Published: (2024)
by: Long, Fuchen, et al.
Published: (2024)
InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution
by: Zhang, Ziqing, et al.
Published: (2025)
by: Zhang, Ziqing, et al.
Published: (2025)
VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer
by: Chen, Liyang, et al.
Published: (2023)
by: Chen, Liyang, et al.
Published: (2023)
Generative Photographic Control for Scene-Consistent Video Cinematic Editing
by: Sun, Huiqiang, et al.
Published: (2025)
by: Sun, Huiqiang, et al.
Published: (2025)
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
by: Liu, Ruiying, et al.
Published: (2025)
by: Liu, Ruiying, et al.
Published: (2025)
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
by: Sun, Yang-Tian, et al.
Published: (2025)
by: Sun, Yang-Tian, et al.
Published: (2025)
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Similar Items
-
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025) -
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025) -
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025) -
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025) -
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
by: Zhao, Lei, et al.
Published: (2025)