ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zixuan, Tang, Chi-Keung, Tai, Yu-Wing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
ChatCam: Empowering Camera Control through Conversational AI
by: Liu, Xinhang, et al.
Published: (2024)
by: Liu, Xinhang, et al.
Published: (2024)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
C3LLM: Conditional Multimodal Content Generation Using Large Language Models
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
WorldCraft: Photo-Realistic 3D World Creation and Customization via LLM Agents
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRF
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning
by: Zhao, Yubo, et al.
Published: (2025)
by: Zhao, Yubo, et al.
Published: (2025)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
by: Guo, Xinyue, et al.
Published: (2025)
by: Guo, Xinyue, et al.
Published: (2025)
InceptionHuman: Controllable Prompt-to-NeRF for Photorealistic 3D Human Generation
by: Kao, Shiu-hong, et al.
Published: (2023)
by: Kao, Shiu-hong, et al.
Published: (2023)
SANeRF-HQ: Segment Anything for NeRF in High Quality
by: Liu, Yichen, et al.
Published: (2023)
by: Liu, Yichen, et al.
Published: (2023)
ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation
by: Ao, Yuzhuo, et al.
Published: (2026)
by: Ao, Yuzhuo, et al.
Published: (2026)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
by: Cong, Gaoxiang, et al.
Published: (2026)
by: Cong, Gaoxiang, et al.
Published: (2026)
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
by: Guan, Kaisi, et al.
Published: (2025)
by: Guan, Kaisi, et al.
Published: (2025)
Inpaint4DNeRF: Promptable Spatio-Temporal NeRF Inpainting with Generative Diffusion Models
by: Jiang, Han, et al.
Published: (2023)
by: Jiang, Han, et al.
Published: (2023)
VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification
by: Liu, Jianmeng, et al.
Published: (2024)
by: Liu, Jianmeng, et al.
Published: (2024)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
by: Yang, Jianxuan, et al.
Published: (2025)
by: Yang, Jianxuan, et al.
Published: (2025)
Learning to Highlight Audio by Watching Movies
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction
by: Liu, Xinhang, et al.
Published: (2023)
by: Liu, Xinhang, et al.
Published: (2023)
When Vision Speaks for Sound
by: Wen, Xiaofei, et al.
Published: (2026)
by: Wen, Xiaofei, et al.
Published: (2026)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
by: Saad, Mahnoor Fatima, et al.
Published: (2025)
by: Saad, Mahnoor Fatima, et al.
Published: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
by: Dai, Yusheng, et al.
Published: (2026)
by: Dai, Yusheng, et al.
Published: (2026)
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
by: Chen, Zihao, et al.
Published: (2024)
by: Chen, Zihao, et al.
Published: (2024)
DragVideo: Interactive Drag-style Video Editing
by: Deng, Yufan, et al.
Published: (2023)
by: Deng, Yufan, et al.
Published: (2023)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
by: Huang-Menders, Alexander, et al.
Published: (2025)
by: Huang-Menders, Alexander, et al.
Published: (2025)
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
by: Cong, Gaoxiang, et al.
Published: (2025)
by: Cong, Gaoxiang, et al.
Published: (2025)
Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
by: Wang, Jiahua, et al.
Published: (2025)
by: Wang, Jiahua, et al.
Published: (2025)
Semantics-Aware Human Motion Generation from Audio Instructions
by: Wang, Zi-An, et al.
Published: (2025)
by: Wang, Zi-An, et al.
Published: (2025)
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
by: Zhang, Zhedong, et al.
Published: (2025)
by: Zhang, Zhedong, et al.
Published: (2025)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Similar Items
-
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
by: Wang, Zixuan, et al.
Published: (2024) -
ChatCam: Empowering Camera Control through Conversational AI
by: Liu, Xinhang, et al.
Published: (2024) -
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025) -
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025) -
C3LLM: Conditional Multimodal Content Generation Using Large Language Models
by: Wang, Zixuan, et al.
Published: (2024)