A Survey on World Models Grounded in Acoustic Physical Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xiaoliang, Chang, Le, Yu, Xin, Huang, Yunhe, Tu, Xianling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
von: Khushiyant, et al.
Veröffentlicht: (2026)
von: Khushiyant, et al.
Veröffentlicht: (2026)
Quantum-Enhanced Analysis and Grading of Vocal Performance
von: Agarwal, Rohan
Veröffentlicht: (2025)
von: Agarwal, Rohan
Veröffentlicht: (2025)
Prevailing Research Areas for Music AI in the Era of Foundation Models
von: Wei, Megan, et al.
Veröffentlicht: (2024)
von: Wei, Megan, et al.
Veröffentlicht: (2024)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
von: Richter-Powell, Jessie, et al.
Veröffentlicht: (2025)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
von: Dhiman, Jai
Veröffentlicht: (2026)
von: Dhiman, Jai
Veröffentlicht: (2026)
Guitar Tone Morphing by Diffusion-based Model
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Automatic Album Sequencing
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
von: Herrmann, Vincent, et al.
Veröffentlicht: (2024)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
von: Aristorenas, Aris J.
Veröffentlicht: (2024)
von: Aristorenas, Aris J.
Veröffentlicht: (2024)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
von: He, Zhanhong, et al.
Veröffentlicht: (2025)
von: He, Zhanhong, et al.
Veröffentlicht: (2025)
GraFPrint: A GNN-Based Approach for Audio Identification
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
Adaptable Symbolic Music Infilling with MIDI-RWKV
von: Zhou-Zheng, Christian, et al.
Veröffentlicht: (2025)
von: Zhou-Zheng, Christian, et al.
Veröffentlicht: (2025)
Passive Underwater Acoustic Signal Separation based on Feature Decoupling Dual-path Network
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
von: Combes, Paolo, et al.
Veröffentlicht: (2025)
von: Combes, Paolo, et al.
Veröffentlicht: (2025)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Generation of Musical Timbres using a Text-Guided Diffusion Model
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
von: Tsangko, Iosif, et al.
Veröffentlicht: (2025)
von: Tsangko, Iosif, et al.
Veröffentlicht: (2025)
Evaluating Model-Agnostic Meta-Learning on MetaWorld ML10 Benchmark: Fast Adaptation in Robotic Manipulation Tasks
von: Atamuradov, Sanjar
Veröffentlicht: (2025)
von: Atamuradov, Sanjar
Veröffentlicht: (2025)
Unified speech and gesture synthesis using flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025)
von: Cauzinille, Jules, et al.
Veröffentlicht: (2025)
Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
von: Wu, Junyan, et al.
Veröffentlicht: (2024)
von: Wu, Junyan, et al.
Veröffentlicht: (2024)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
von: Maben, Leander Melroy, et al.
Veröffentlicht: (2025)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
von: Cui, Hanfang, et al.
Veröffentlicht: (2025)
von: Cui, Hanfang, et al.
Veröffentlicht: (2025)
Analyzing the Impact of Multimodal Perception on Sample Complexity and Optimization Landscapes in Imitation Learning
von: Abuelsamen, Luai, et al.
Veröffentlicht: (2025)
von: Abuelsamen, Luai, et al.
Veröffentlicht: (2025)
BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
von: Park, Taesoo, et al.
Veröffentlicht: (2025)
von: Park, Taesoo, et al.
Veröffentlicht: (2025)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
von: Sharma, Aditya, et al.
Veröffentlicht: (2025)
von: Sharma, Aditya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025) -
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025) -
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024) -
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023) -
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)