MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Canyu, Liu, Mingyu, Wang, Wen, Chen, Weihua, Wang, Fan, Chen, Hao, Zhang, Bo, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
von: Wu, Weijia, et al.
Veröffentlicht: (2024)
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
von: Ding, Ganggui, et al.
Veröffentlicht: (2024)
von: Ding, Ganggui, et al.
Veröffentlicht: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
von: Li, Liyang, et al.
Veröffentlicht: (2026)
von: Li, Liyang, et al.
Veröffentlicht: (2026)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
von: Zhong, Hao, et al.
Veröffentlicht: (2025)
von: Zhong, Hao, et al.
Veröffentlicht: (2025)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
von: Zhao, Canyu, et al.
Veröffentlicht: (2025)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
von: Xu, Tianshuo, et al.
Veröffentlicht: (2024)
von: Xu, Tianshuo, et al.
Veröffentlicht: (2024)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
von: Zhao, Canyu, et al.
Veröffentlicht: (2026)
von: Zhao, Canyu, et al.
Veröffentlicht: (2026)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
von: Xu, Guangkai, et al.
Veröffentlicht: (2024)
von: Xu, Guangkai, et al.
Veröffentlicht: (2024)
Generative Active Learning for Long-tailed Instance Segmentation
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2024)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
von: Zhu, Muzhi, et al.
Veröffentlicht: (2025)
BrainDreamer: Reasoning-Coherent and Controllable Image Generation from EEG Brain Signals via Language Guidance
von: Wang, Ling, et al.
Veröffentlicht: (2024)
von: Wang, Ling, et al.
Veröffentlicht: (2024)
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
von: Chen, Zhekai, et al.
Veröffentlicht: (2024)
von: Chen, Zhekai, et al.
Veröffentlicht: (2024)
InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation
von: Zhuo, Wenjie, et al.
Veröffentlicht: (2024)
von: Zhuo, Wenjie, et al.
Veröffentlicht: (2024)
LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
von: Cheng, Chong, et al.
Veröffentlicht: (2026)
von: Cheng, Chong, et al.
Veröffentlicht: (2026)
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
von: Huang, Zheng, et al.
Veröffentlicht: (2025)
HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
von: Wang, Boyuan, et al.
Veröffentlicht: (2025)
FusDreamer: Label-efficient Remote Sensing World Model for Multimodal Data Classification
von: Wang, Jinping, et al.
Veröffentlicht: (2025)
von: Wang, Jinping, et al.
Veröffentlicht: (2025)
MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology
von: Yang, Shu, et al.
Veröffentlicht: (2024)
von: Yang, Shu, et al.
Veröffentlicht: (2024)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
von: Wang, Jianghui, et al.
Veröffentlicht: (2023)
von: Wang, Jianghui, et al.
Veröffentlicht: (2023)
DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
von: Zhao, Guosheng, et al.
Veröffentlicht: (2024)
von: Zhao, Guosheng, et al.
Veröffentlicht: (2024)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
Generative Video Matting
von: Ge, Yongtao, et al.
Veröffentlicht: (2025)
von: Ge, Yongtao, et al.
Veröffentlicht: (2025)
FlexiDreamer: Single Image-to-3D Generation with FlexiCubes
von: Zhao, Ruowen, et al.
Veröffentlicht: (2024)
von: Zhao, Ruowen, et al.
Veröffentlicht: (2024)
LumosFlow: Motion-Guided Long Video Generation
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
Diffusion Models are Efficient Data Generators for Human Mesh Recovery
von: Ge, Yongtao, et al.
Veröffentlicht: (2024)
von: Ge, Yongtao, et al.
Veröffentlicht: (2024)
Towards Automated Movie Trailer Generation
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
von: Qiu, Di, et al.
Veröffentlicht: (2024)
von: Qiu, Di, et al.
Veröffentlicht: (2024)
BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation
von: Wang, Yutong, et al.
Veröffentlicht: (2026)
von: Wang, Yutong, et al.
Veröffentlicht: (2026)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaofeng, et al.
Veröffentlicht: (2024)
CityDreamer: Compositional Generative Model of Unbounded 3D Cities
von: Xie, Haozhe, et al.
Veröffentlicht: (2023)
von: Xie, Haozhe, et al.
Veröffentlicht: (2023)
ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation
von: Zhao, Guosheng, et al.
Veröffentlicht: (2025)
von: Zhao, Guosheng, et al.
Veröffentlicht: (2025)
Framer: Interactive Frame Interpolation
von: Wang, Wen, et al.
Veröffentlicht: (2024)
von: Wang, Wen, et al.
Veröffentlicht: (2024)
Object-aware Inversion and Reassembly for Image Editing
von: Yang, Zhen, et al.
Veröffentlicht: (2023)
von: Yang, Zhen, et al.
Veröffentlicht: (2023)
RGM: A Robust Generalizable Matching Model
von: Zhang, Songyan, et al.
Veröffentlicht: (2023)
von: Zhang, Songyan, et al.
Veröffentlicht: (2023)
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
von: Fan, Chengxiang, et al.
Veröffentlicht: (2024)
von: Fan, Chengxiang, et al.
Veröffentlicht: (2024)
Manifold-Aware Point Cloud Completion via Geodesic-Attentive Hierarchical Feature Learning
von: Sun, Jianan, et al.
Veröffentlicht: (2025)
von: Sun, Jianan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
von: Wu, Weijia, et al.
Veröffentlicht: (2024) -
FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
von: Ding, Ganggui, et al.
Veröffentlicht: (2024) -
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
von: Li, Liyang, et al.
Veröffentlicht: (2026) -
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
von: Zhao, Canyu, et al.
Veröffentlicht: (2025) -
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
von: Zhong, Hao, et al.
Veröffentlicht: (2025)