MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Haoxuan, Wang, Zihao, Hong, Haorong, Feng, Youwei, Yu, Jiaxin, Diao, Han, Xu, Yunfei, Zhang, Kejun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
di: Huh, Mina, et al.
Pubblicazione: (2026)
di: Huh, Mina, et al.
Pubblicazione: (2026)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024)
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
di: Kahl, Benjamin
Pubblicazione: (2025)
di: Kahl, Benjamin
Pubblicazione: (2025)
Creating Aesthetic Sonifications on the Web with SIREN
di: Peng, Tristan, et al.
Pubblicazione: (2024)
di: Peng, Tristan, et al.
Pubblicazione: (2024)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
di: Cai, Zhuojiang, et al.
Pubblicazione: (2024)
di: Cai, Zhuojiang, et al.
Pubblicazione: (2024)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
di: Hopkins, Torin, et al.
Pubblicazione: (2026)
di: Hopkins, Torin, et al.
Pubblicazione: (2026)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
di: Kim, Yonghyun, et al.
Pubblicazione: (2025)
di: Kim, Yonghyun, et al.
Pubblicazione: (2025)
Efficient Video to Audio Mapper with Visual Scene Detection
di: Yi, Mingjing, et al.
Pubblicazione: (2024)
di: Yi, Mingjing, et al.
Pubblicazione: (2024)
Personality-Enhanced Multimodal Depression Detection in the Elderly
di: Wang, Honghong, et al.
Pubblicazione: (2025)
di: Wang, Honghong, et al.
Pubblicazione: (2025)
Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
di: Xie, Siyi, et al.
Pubblicazione: (2025)
di: Xie, Siyi, et al.
Pubblicazione: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
StereoFoley: Object-Aware Stereo Audio Generation from Video
di: Karchkhadze, Tornike, et al.
Pubblicazione: (2025)
di: Karchkhadze, Tornike, et al.
Pubblicazione: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
di: He, Mao-Kui, et al.
Pubblicazione: (2024)
di: He, Mao-Kui, et al.
Pubblicazione: (2024)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance Generation
di: Yang, Kaixing, et al.
Pubblicazione: (2025)
di: Yang, Kaixing, et al.
Pubblicazione: (2025)
MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
di: Yao, Dong, et al.
Pubblicazione: (2023)
di: Yao, Dong, et al.
Pubblicazione: (2023)
Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations
di: Wachter, Maximilian, et al.
Pubblicazione: (2026)
di: Wachter, Maximilian, et al.
Pubblicazione: (2026)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
di: Lei, Ke, et al.
Pubblicazione: (2026)
di: Lei, Ke, et al.
Pubblicazione: (2026)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
di: Sulun, Serkan, et al.
Pubblicazione: (2025)
di: Sulun, Serkan, et al.
Pubblicazione: (2025)
DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
di: Hong, Fa-Ting, et al.
Pubblicazione: (2024)
Long-Form Text-to-Music Generation with Adaptive Prompts: A Case Study in Tabletop Role-Playing Games Soundtracks
di: Marra, Felipe, et al.
Pubblicazione: (2024)
di: Marra, Felipe, et al.
Pubblicazione: (2024)
AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
di: Jiang, Yi-Lin, et al.
Pubblicazione: (2024)
di: Jiang, Yi-Lin, et al.
Pubblicazione: (2024)
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
di: Bryan-Kinns, Nick, et al.
Pubblicazione: (2024)
di: Bryan-Kinns, Nick, et al.
Pubblicazione: (2024)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
di: Selvamani, Shaja Arul, et al.
Pubblicazione: (2025)
di: Selvamani, Shaja Arul, et al.
Pubblicazione: (2025)
Towards Reliable Large Audio Language Model
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
Workflow-Based Evaluation of Music Generation Systems
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
di: Peng, Jing, et al.
Pubblicazione: (2026)
di: Peng, Jing, et al.
Pubblicazione: (2026)
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
di: Yang, Kaixing, et al.
Pubblicazione: (2025)
di: Yang, Kaixing, et al.
Pubblicazione: (2025)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
di: Kim, Haven, et al.
Pubblicazione: (2025)
di: Kim, Haven, et al.
Pubblicazione: (2025)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
di: Gu, Ke, et al.
Pubblicazione: (2025)
di: Gu, Ke, et al.
Pubblicazione: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
di: Cui, Yang, et al.
Pubblicazione: (2025)
di: Cui, Yang, et al.
Pubblicazione: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
Building Audio-Visual Digital Twins with Smartphones
di: Lan, Zitong, et al.
Pubblicazione: (2025)
di: Lan, Zitong, et al.
Pubblicazione: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
di: Gong, Jingyao
Pubblicazione: (2026)
di: Gong, Jingyao
Pubblicazione: (2026)
Documenti analoghi
-
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
di: Huh, Mina, et al.
Pubblicazione: (2026) -
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025) -
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024) -
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025) -
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
di: Kahl, Benjamin
Pubblicazione: (2025)