MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chi, Xiaowei, Wang, Yatian, Cheng, Aosong, Fang, Pengjun, Tian, Zeyue, He, Yingqing, Liu, Zhaoyang, Qi, Xingqun, Pan, Jiahao, Zhang, Rongyu, Li, Mengfei, Yuan, Ruibin, Jiang, Yanbing, Xue, Wei, Luo, Wenhan, Chen, Qifeng, Zhang, Shanghang, Liu, Qifeng, Guo, Yike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
von: Tian, Zeyue, et al.
Veröffentlicht: (2024)
von: Tian, Zeyue, et al.
Veröffentlicht: (2024)
EVA: An Embodied World Model for Future Video Anticipation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
LLMs Meet Multimodal Generation and Editing: A Survey
von: He, Yingqing, et al.
Veröffentlicht: (2024)
von: He, Yingqing, et al.
Veröffentlicht: (2024)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
AudioX: A Unified Framework for Anything-to-Audio Generation
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
von: Qi, Xingqun, et al.
Veröffentlicht: (2025)
von: Qi, Xingqun, et al.
Veröffentlicht: (2025)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
von: Fang, Pengjun, et al.
Veröffentlicht: (2026)
von: Fang, Pengjun, et al.
Veröffentlicht: (2026)
CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
von: Qi, Xingqun, et al.
Veröffentlicht: (2024)
von: Qi, Xingqun, et al.
Veröffentlicht: (2024)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
von: Wu, Kam Man, et al.
Veröffentlicht: (2025)
von: Wu, Kam Man, et al.
Veröffentlicht: (2025)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Hengyuan, et al.
Veröffentlicht: (2025)
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
von: Pham, Kien T., et al.
Veröffentlicht: (2025)
von: Pham, Kien T., et al.
Veröffentlicht: (2025)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
TALE: Training-free Cross-domain Image Composition via Adaptive Latent Manipulation and Energy-guided Optimization
von: Pham, Kien T., et al.
Veröffentlicht: (2024)
von: Pham, Kien T., et al.
Veröffentlicht: (2024)
ChatMusician: Understanding and Generating Music Intrinsically with LLM
von: Yuan, Ruibin, et al.
Veröffentlicht: (2024)
von: Yuan, Ruibin, et al.
Veröffentlicht: (2024)
An Automatic Deep Learning Approach for Trailer Generation through Large Language Models
von: Balestri, Roberto, et al.
Veröffentlicht: (2026)
von: Balestri, Roberto, et al.
Veröffentlicht: (2026)
ComposerX: Multi-Agent Symbolic Music Composition with LLMs
von: Deng, Qixin, et al.
Veröffentlicht: (2024)
von: Deng, Qixin, et al.
Veröffentlicht: (2024)
YuE: Scaling Open Foundation Models for Long-Form Music Generation
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas
von: Bretti, Carlo, et al.
Veröffentlicht: (2024)
von: Bretti, Carlo, et al.
Veröffentlicht: (2024)
Trailer Reimagined: An Innovative, Llm-DRiven, Expressive Automated Movie Summary framework (TRAILDREAMS)
von: Balestri, Roberto, et al.
Veröffentlicht: (2026)
von: Balestri, Roberto, et al.
Veröffentlicht: (2026)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
Music Grounding by Short Video
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings
von: Liu, Shimiao, et al.
Veröffentlicht: (2025)
von: Liu, Shimiao, et al.
Veröffentlicht: (2025)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
PiGW: A Plug-in Generative Watermarking Framework
von: Ma, Rui, et al.
Veröffentlicht: (2024)
von: Ma, Rui, et al.
Veröffentlicht: (2024)
Interpretable Zero-shot Referring Expression Comprehension with Query-driven Scene Graphs
von: Wu, Yike, et al.
Veröffentlicht: (2026)
von: Wu, Yike, et al.
Veröffentlicht: (2026)
MusFlow: Multimodal Music Generation via Conditional Flow Matching
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
von: Xu, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Xu, Zhaoyang, et al.
Veröffentlicht: (2026)
Music Arena: Live Evaluation for Text-to-Music
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
Adaptive 3D Mesh Steganography Based on Feature-Preserving Distortion
von: Zhang, Yushu, et al.
Veröffentlicht: (2022)
von: Zhang, Yushu, et al.
Veröffentlicht: (2022)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
von: Yao, Dong, et al.
Veröffentlicht: (2023)
von: Yao, Dong, et al.
Veröffentlicht: (2023)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
Audio-FLAN: A Preliminary Release
von: Xue, Liumeng, et al.
Veröffentlicht: (2025)
von: Xue, Liumeng, et al.
Veröffentlicht: (2025)
MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
von: Wu, Chengyu, et al.
Veröffentlicht: (2024)
von: Wu, Chengyu, et al.
Veröffentlicht: (2024)
Towards Practical Real-Time Low-Latency Music Source Separation
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
von: Tian, Zeyue, et al.
Veröffentlicht: (2024) -
EVA: An Embodied World Model for Future Video Anticipation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024) -
LLMs Meet Multimodal Generation and Editing: A Survey
von: He, Yingqing, et al.
Veröffentlicht: (2024) -
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024) -
AudioX: A Unified Framework for Anything-to-Audio Generation
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)