MV-Crafter: An Intelligent System for Music-guided Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Chuer, Dang, Shengqi, Liu, Yuqi, Zhao, Nanxuan, Shi, Yang, Cao, Nan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChartBlender: An Interactive System for Authoring and Synchronizing Visualization Charts in Video
by: He, Yi, et al.
Published: (2025)
by: He, Yi, et al.
Published: (2025)
On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models
by: Taniguchi, Tadahiro
Published: (2025)
by: Taniguchi, Tadahiro
Published: (2025)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)
by: Fu, Chencan, et al.
Published: (2024)
Save It for the "Hot" Day: An LLM-Empowered Visual Analytics System for Heat Risk Management
by: Li, Haobo, et al.
Published: (2024)
by: Li, Haobo, et al.
Published: (2024)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
by: Qiu, Yue, et al.
Published: (2025)
by: Qiu, Yue, et al.
Published: (2025)
Private Chat in a Public Space of Metaverse Systems
by: Chen, Jiarui, et al.
Published: (2025)
by: Chen, Jiarui, et al.
Published: (2025)
MRATTS: An MR-Based Acupoint Therapy Training System with Real-Time Acupoint Detection and Evaluation Standards
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
by: Huh, Mina, et al.
Published: (2026)
by: Huh, Mina, et al.
Published: (2026)
MULTI-CASE: A Transformer-based Ethics-aware Multimodal Investigative Intelligence Framework
by: Fischer, Maximilian T., et al.
Published: (2024)
by: Fischer, Maximilian T., et al.
Published: (2024)
A Collaborative Extended Reality Prototype for 3D Surgical Planning and Visualization
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
User-Generated Content and Editors in Games: A Comprehensive Survey
by: Liu, Yuyue, et al.
Published: (2024)
by: Liu, Yuyue, et al.
Published: (2024)
DataScout: Automatic Data Fact Retrieval for Statement Augmentation with an LLM-Based Agent
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
by: Wong, Kam Kwai, et al.
Published: (2023)
by: Wong, Kam Kwai, et al.
Published: (2023)
ICE: Interactive 3D Game Character Editing via Dialogue
by: Wu, Haoqian, et al.
Published: (2024)
by: Wu, Haoqian, et al.
Published: (2024)
Secure & Personalized Music-to-Video Generation via CHARCHA
by: Agarwal, Mehul, et al.
Published: (2025)
by: Agarwal, Mehul, et al.
Published: (2025)
IDEA: Augmenting Design Intelligence through Design Space Exploration
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Disc-Cover Complexity Trends in Music Illustrations from Sinatra to Swift
by: Fracaro, Nicolas, et al.
Published: (2025)
by: Fracaro, Nicolas, et al.
Published: (2025)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments
by: Tong, Yuqi, et al.
Published: (2024)
by: Tong, Yuqi, et al.
Published: (2024)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
by: Yeh, Catherine, et al.
Published: (2026)
by: Yeh, Catherine, et al.
Published: (2026)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
by: Chen, Xiaolin, et al.
Published: (2022)
by: Chen, Xiaolin, et al.
Published: (2022)
Language-Guided Multimodal Texture Authoring via Generative Models
by: Qian, Wanli, et al.
Published: (2026)
by: Qian, Wanli, et al.
Published: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Human-Machine Collaboration-Guided Space Design: Combination of Machine Learning Models and Humanistic Design Concepts
by: Yang, Yuxuan
Published: (2025)
by: Yang, Yuxuan
Published: (2025)
Multimodal Infusion Tuning for Large Models
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
PoEmotion: Can AI Utilize Chinese Calligraphy to Express Emotion from Poems?
by: Liu, Tiancheng, et al.
Published: (2025)
by: Liu, Tiancheng, et al.
Published: (2025)
musicolors: Bridging Sound and Visuals For Synesthetic Creative Musical Experience
by: Lee, ChungHa, et al.
Published: (2025)
by: Lee, ChungHa, et al.
Published: (2025)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Video-Mediated Emotion Disclosure: Expressions of Fear, Sadness, and Joy by People with Schizophrenia on YouTube
by: Liu, Jiaying Lizzy, et al.
Published: (2025)
by: Liu, Jiaying Lizzy, et al.
Published: (2025)
VCEMO: Multi-Modal Emotion Recognition for Chinese Voiceprints
by: Tang, Jinghua, et al.
Published: (2024)
by: Tang, Jinghua, et al.
Published: (2024)
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025)
by: Kyaw, Alexander Htet, et al.
Published: (2025)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
by: De La Fuente, Tessa
Published: (2025)
by: De La Fuente, Tessa
Published: (2025)
Unveiling the Visual Rhetoric of Persuasive Cartography: A Case Study of the Design of Octopus Maps
by: Lin, Daocheng, et al.
Published: (2025)
by: Lin, Daocheng, et al.
Published: (2025)
Simulacra Naturae: Generative Ecosystem driven by Agent-Based Simulations and Brain Organoid Collective Intelligence
by: Manoudaki, Nefeli, et al.
Published: (2025)
by: Manoudaki, Nefeli, et al.
Published: (2025)
Similar Items
-
ChartBlender: An Interactive System for Authoring and Synchronizing Visualization Charts in Video
by: He, Yi, et al.
Published: (2025) -
On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models
by: Taniguchi, Tadahiro
Published: (2025) -
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025) -
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024) -
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)