BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yutong, Wang, Yunke, Chen, Xinyuan, Xu, Chang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Automated Movie Trailer Generation
by: Argaw, Dawit Mureja, et al.
Published: (2024)
by: Argaw, Dawit Mureja, et al.
Published: (2024)
Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
by: Zhu, Sidan, et al.
Published: (2025)
by: Zhu, Sidan, et al.
Published: (2025)
PARE: Pruning and Adaptive Routing for Efficient Video Generation
by: Wang, Yutong, et al.
Published: (2026)
by: Wang, Yutong, et al.
Published: (2026)
Movie Trailer Genre Classification Using Multimodal Pretrained Features
by: Sulun, Serkan, et al.
Published: (2024)
by: Sulun, Serkan, et al.
Published: (2024)
Multi-Tailed Vision Transformer for Efficient Inference
by: Wang, Yunke, et al.
Published: (2022)
by: Wang, Yunke, et al.
Published: (2022)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
by: Zhao, Canyu, et al.
Published: (2024)
by: Zhao, Canyu, et al.
Published: (2024)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
by: Wu, Weijia, et al.
Published: (2024)
by: Wu, Weijia, et al.
Published: (2024)
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
Intensive Vision-guided Network for Radiology Report Generation
by: Zheng, Fudan, et al.
Published: (2024)
by: Zheng, Fudan, et al.
Published: (2024)
UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection
by: Chang, Laibin, et al.
Published: (2026)
by: Chang, Laibin, et al.
Published: (2026)
MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
by: Xu, Meng, et al.
Published: (2024)
by: Xu, Meng, et al.
Published: (2024)
CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
by: Zheng, Mingzhe, et al.
Published: (2025)
by: Zheng, Mingzhe, et al.
Published: (2025)
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
Marine Saliency Segmenter: Object-Focused Conditional Diffusion with Region-Level Semantic Knowledge Distillation
by: Chang, Laibin, et al.
Published: (2025)
by: Chang, Laibin, et al.
Published: (2025)
X-Dancer: Expressive Music to Human Dance Video Generation
by: Chen, Zeyuan, et al.
Published: (2025)
by: Chen, Zeyuan, et al.
Published: (2025)
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
by: Li, Shujia, et al.
Published: (2025)
by: Li, Shujia, et al.
Published: (2025)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
MovieTeller: Tool-augmented Movie Synopsis with ID Consistent Progressive Abstraction
by: Li, Yizhi, et al.
Published: (2026)
by: Li, Yizhi, et al.
Published: (2026)
Movie101v2: Improved Movie Narration Benchmark
by: Yue, Zihao, et al.
Published: (2024)
by: Yue, Zihao, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
by: He, Runze, et al.
Published: (2026)
by: He, Runze, et al.
Published: (2026)
Captain Cinema: Towards Short Movie Generation
by: Xiao, Junfei, et al.
Published: (2025)
by: Xiao, Junfei, et al.
Published: (2025)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025)
by: Liu, Shanhui, et al.
Published: (2025)
Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas
by: Bretti, Carlo, et al.
Published: (2024)
by: Bretti, Carlo, et al.
Published: (2024)
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
by: Hao, Yutong, et al.
Published: (2025)
by: Hao, Yutong, et al.
Published: (2025)
Visual Imitation Learning with Calibrated Contrastive Representation
by: Wang, Yunke, et al.
Published: (2024)
by: Wang, Yunke, et al.
Published: (2024)
BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing
by: Wang, Kaiwen, et al.
Published: (2026)
by: Wang, Kaiwen, et al.
Published: (2026)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023)
by: Wang, Jianghui, et al.
Published: (2023)
Finding the Right Moment: Human-Assisted Trailer Creation via Task Composition
by: Papalampidi, Pinelopi, et al.
Published: (2021)
by: Papalampidi, Pinelopi, et al.
Published: (2021)
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025)
by: Faure, Gueter Josmy, et al.
Published: (2025)
MAEDiff: Masked Autoencoder-enhanced Diffusion Models for Unsupervised Anomaly Detection in Brain Images
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
by: Liu, Xinran, et al.
Published: (2026)
by: Liu, Xinran, et al.
Published: (2026)
ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent
by: Wang, Hanyi, et al.
Published: (2026)
by: Wang, Hanyi, et al.
Published: (2026)
Automated Movie Generation via Multi-Agent CoT Planning
by: Wu, Weijia, et al.
Published: (2025)
by: Wu, Weijia, et al.
Published: (2025)
Similar Items
-
Towards Automated Movie Trailer Generation
by: Argaw, Dawit Mureja, et al.
Published: (2024) -
Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
by: Zhu, Sidan, et al.
Published: (2025) -
PARE: Pruning and Adaptive Routing for Efficient Video Generation
by: Wang, Yutong, et al.
Published: (2026) -
Movie Trailer Genre Classification Using Multimodal Pretrained Features
by: Sulun, Serkan, et al.
Published: (2024) -
Multi-Tailed Vision Transformer for Efficient Inference
by: Wang, Yunke, et al.
Published: (2022)