LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Jianxiong, Chen, Zhaoxi, Liu, Xian, Feng, Jianfeng, Si, Chenyang, Fu, Yanwei, Qiao, Yu, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
Making Your Dreams A Reality: Decoding the Dreams into a Coherent Video Story from fMRI Signals
by: Fu, Yanwei, et al.
Published: (2025)
by: Fu, Yanwei, et al.
Published: (2025)
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing
by: Gao, Jianxiong, et al.
Published: (2025)
by: Gao, Jianxiong, et al.
Published: (2025)
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)
by: Si, Chenyang, et al.
Published: (2025)
MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset
by: Gao, Jianxiong, et al.
Published: (2024)
by: Gao, Jianxiong, et al.
Published: (2024)
FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
MinD-3D: Reconstruct High-quality 3D objects in Human Brain
by: Gao, Jianxiong, et al.
Published: (2023)
by: Gao, Jianxiong, et al.
Published: (2023)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
by: Yang, Ying, et al.
Published: (2026)
by: Yang, Ying, et al.
Published: (2026)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution Detection
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
by: Cao, Yukang, et al.
Published: (2025)
by: Cao, Yukang, et al.
Published: (2025)
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
by: Qiu, Haonan, et al.
Published: (2024)
by: Qiu, Haonan, et al.
Published: (2024)
Generative Gaussian Splatting for Unbounded 3D City Generation
by: Xie, Haozhe, et al.
Published: (2024)
by: Xie, Haozhe, et al.
Published: (2024)
FreeInit: Bridging Initialization Gap in Video Diffusion Models
by: Wu, Tianxing, et al.
Published: (2023)
by: Wu, Tianxing, et al.
Published: (2023)
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
by: Lv, Zhengyao, et al.
Published: (2024)
by: Lv, Zhengyao, et al.
Published: (2024)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
by: Zhang, Haojie, et al.
Published: (2024)
by: Zhang, Haojie, et al.
Published: (2024)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
Compositional Generative Model of Unbounded 4D Cities
by: Xie, Haozhe, et al.
Published: (2025)
by: Xie, Haozhe, et al.
Published: (2025)
LumosFlow: Motion-Guided Long Video Generation
by: Chen, Jiahao, et al.
Published: (2025)
by: Chen, Jiahao, et al.
Published: (2025)
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
by: Xie, Haozhe, et al.
Published: (2023)
by: Xie, Haozhe, et al.
Published: (2023)
Collaborative Multi-Modal Coding for High-Quality 3D Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
PhysX-3D: Physical-Grounded 3D Asset Generation
by: Cao, Ziang, et al.
Published: (2025)
by: Cao, Ziang, et al.
Published: (2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
by: Lin, Hangyu, et al.
Published: (2026)
by: Lin, Hangyu, et al.
Published: (2026)
Hyper-Transformer for Amodal Completion
by: Gao, Jianxiong, et al.
Published: (2024)
by: Gao, Jianxiong, et al.
Published: (2024)
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
by: Liu, Tianqi, et al.
Published: (2025)
by: Liu, Tianqi, et al.
Published: (2025)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
by: Gu, Zekai, et al.
Published: (2025)
by: Gu, Zekai, et al.
Published: (2025)
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
by: Liu, Zixian, et al.
Published: (2026)
by: Liu, Zixian, et al.
Published: (2026)
Modeling Spatiotemporal Neural Frames for High Resolution Brain Dynamic
by: Qu, Wanying, et al.
Published: (2026)
by: Qu, Wanying, et al.
Published: (2026)
3D Scene Generation: A Survey
by: Wen, Beichen, et al.
Published: (2025)
by: Wen, Beichen, et al.
Published: (2025)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
by: Qi, Tianhao, et al.
Published: (2025)
by: Qi, Tianhao, et al.
Published: (2025)
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
LongDiff: Training-Free Long Video Generation in One Go
by: Li, Zhuoling, et al.
Published: (2025)
by: Li, Zhuoling, et al.
Published: (2025)
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
by: Lin, Chin-Yang, et al.
Published: (2025)
by: Lin, Chin-Yang, et al.
Published: (2025)
The Pictorial Cortex: Zero-Shot Cross-Subject fMRI-to-Image Reconstruction via Compositional Latent Modeling
by: Huo, Jingyang, et al.
Published: (2026)
by: Huo, Jingyang, et al.
Published: (2026)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025)
by: Tian, Shulin, et al.
Published: (2025)
Similar Items
-
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
by: Gao, Jianxiong, et al.
Published: (2025) -
Making Your Dreams A Reality: Decoding the Dreams into a Coherent Video Story from fMRI Signals
by: Fu, Yanwei, et al.
Published: (2025) -
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
by: Lv, Zhengyao, et al.
Published: (2025) -
CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing
by: Gao, Jianxiong, et al.
Published: (2025) -
RepVideo: Rethinking Cross-Layer Representation for Video Generation
by: Si, Chenyang, et al.
Published: (2025)