CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Yufan, Guo, Xun, Wang, Yizhi, Fang, Jacob Zhiyuan, Wang, Angtian, Yuan, Shenghai, Yang, Yiding, Liu, Bo, Huang, Haibin, Ma, Chongyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
ATI: Any Trajectory Instruction for Controllable Video Generation
by: Wang, Angtian, et al.
Published: (2025)
by: Wang, Angtian, et al.
Published: (2025)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
by: Zhang, Guofeng, et al.
Published: (2026)
by: Zhang, Guofeng, et al.
Published: (2026)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025)
by: Zhang, Guofeng, et al.
Published: (2025)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
by: Guo, Xun, et al.
Published: (2023)
by: Guo, Xun, et al.
Published: (2023)
Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models
by: Yin, Yuanyang, et al.
Published: (2026)
by: Yin, Yuanyang, et al.
Published: (2026)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
by: Yuan, Shenghai, et al.
Published: (2025)
by: Yuan, Shenghai, et al.
Published: (2025)
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
by: Gu, Yuming, et al.
Published: (2025)
by: Gu, Yuming, et al.
Published: (2025)
FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
by: FSVideo Team, et al.
Published: (2026)
by: FSVideo Team, et al.
Published: (2026)
StoryMem: Multi-shot Long Video Storytelling with Memory
by: Zhang, Kaiwen, et al.
Published: (2025)
by: Zhang, Kaiwen, et al.
Published: (2025)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
by: Zhang, Hongyu, et al.
Published: (2025)
by: Zhang, Hongyu, et al.
Published: (2025)
DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning
by: Guo, Xun, et al.
Published: (2024)
by: Guo, Xun, et al.
Published: (2024)
Backdoor Cleaning without External Guidance in MLLM Fine-tuning
by: Rong, Xuankun, et al.
Published: (2025)
by: Rong, Xuankun, et al.
Published: (2025)
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation
by: Xu, Yijia, et al.
Published: (2026)
by: Xu, Yijia, et al.
Published: (2026)
ViMo: Generating Motions from Casual Videos
by: Qiu, Liangdong, et al.
Published: (2024)
by: Qiu, Liangdong, et al.
Published: (2024)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
Semantic Flow: Learning Semantic Field of Dynamic Scenes from Monocular Videos
by: Tian, Fengrui, et al.
Published: (2024)
by: Tian, Fengrui, et al.
Published: (2024)
InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance
by: Pan, Dongwei, et al.
Published: (2026)
by: Pan, Dongwei, et al.
Published: (2026)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
by: Yang, Shiyuan, et al.
Published: (2024)
by: Yang, Shiyuan, et al.
Published: (2024)
Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
by: Liu, Shengyuan, et al.
Published: (2024)
by: Liu, Shengyuan, et al.
Published: (2024)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Proactive Guidance of Multi-Turn Conversation in Industrial Search
by: Li, Xiaoyu, et al.
Published: (2025)
by: Li, Xiaoyu, et al.
Published: (2025)
UNICBench: UNIfied Counting Benchmark for MLLM
by: Rong, Chenggang, et al.
Published: (2026)
by: Rong, Chenggang, et al.
Published: (2026)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
by: Li, Zongjian, et al.
Published: (2025)
by: Li, Zongjian, et al.
Published: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
Efficient Motion-Aware Video MLLM
by: Zhao, Zijia, et al.
Published: (2025)
by: Zhao, Zijia, et al.
Published: (2025)
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
by: Zhang, Zhenxing, et al.
Published: (2025)
by: Zhang, Zhenxing, et al.
Published: (2025)
PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment
by: Xiong, Zhexiao, et al.
Published: (2026)
by: Xiong, Zhexiao, et al.
Published: (2026)
MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
by: Li, Tianyuan, et al.
Published: (2025)
by: Li, Tianyuan, et al.
Published: (2025)
3D UAV Trajectory Estimation and Classification from Internet Videos via Language Model
by: Lei, Haoxiang, et al.
Published: (2026)
by: Lei, Haoxiang, et al.
Published: (2026)
LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generation
by: Yang, Cunyuan, et al.
Published: (2026)
by: Yang, Cunyuan, et al.
Published: (2026)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
by: Fang, Rongyao, et al.
Published: (2024)
by: Fang, Rongyao, et al.
Published: (2024)
ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding
by: Xia, Tianze, et al.
Published: (2026)
by: Xia, Tianze, et al.
Published: (2026)
Distributed Invariant Kalman Filter for Cooperative Localization using Matrix Lie Groups
by: Zhou, Yizhi, et al.
Published: (2024)
by: Zhou, Yizhi, et al.
Published: (2024)
Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
by: Wang, Shulei, et al.
Published: (2025)
by: Wang, Shulei, et al.
Published: (2025)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
by: Shaulov, Ariel, et al.
Published: (2025)
by: Shaulov, Ariel, et al.
Published: (2025)
Artic: AI-oriented Real-time Communication for MLLM Video Assistant
by: Wu, Jiangkai, et al.
Published: (2026)
by: Wu, Jiangkai, et al.
Published: (2026)
Similar Items
-
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025) -
ATI: Any Trajectory Instruction for Controllable Video Generation
by: Wang, Angtian, et al.
Published: (2025) -
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025) -
HECTOR: Hybrid Editable Compositional Object References for Video Generation
by: Zhang, Guofeng, et al.
Published: (2026) -
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025)