Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
Fuente:
arXiv
Guardado en:
| Autores principales: | Richter-Powell, Jessie, Torralba, Antonio, Lorraine, Jonathan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automatic Album Sequencing
por: Herrmann, Vincent, et al.
Publicado: (2024)
por: Herrmann, Vincent, et al.
Publicado: (2024)
A Framework for Multimodal Medical Image Interaction
por: Schütz, Laura, et al.
Publicado: (2024)
por: Schütz, Laura, et al.
Publicado: (2024)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
por: Woodard, Brandon, et al.
Publicado: (2025)
por: Woodard, Brandon, et al.
Publicado: (2025)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
por: Cho, Hyunsung, et al.
Publicado: (2024)
por: Cho, Hyunsung, et al.
Publicado: (2024)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
por: Khushiyant, et al.
Publicado: (2026)
por: Khushiyant, et al.
Publicado: (2026)
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
por: Chopra, Anuradha, et al.
Publicado: (2025)
por: Chopra, Anuradha, et al.
Publicado: (2025)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
por: Dhiman, Jai
Publicado: (2026)
por: Dhiman, Jai
Publicado: (2026)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
por: Samsurya, Bilkent
Publicado: (2025)
por: Samsurya, Bilkent
Publicado: (2025)
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
por: Combes, Paolo, et al.
Publicado: (2025)
por: Combes, Paolo, et al.
Publicado: (2025)
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
por: Khah, Arman Nik, et al.
Publicado: (2026)
por: Khah, Arman Nik, et al.
Publicado: (2026)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
por: Melechovsky, Jan, et al.
Publicado: (2025)
por: Melechovsky, Jan, et al.
Publicado: (2025)
Music Tempo Estimation on Solo Instrumental Performance
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
por: Mehta, Shivam, et al.
Publicado: (2024)
por: Mehta, Shivam, et al.
Publicado: (2024)
GraFPrint: A GNN-Based Approach for Audio Identification
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
Adaptable Symbolic Music Infilling with MIDI-RWKV
por: Zhou-Zheng, Christian, et al.
Publicado: (2025)
por: Zhou-Zheng, Christian, et al.
Publicado: (2025)
Dichotic harmony for the musical practice
por: Madgazin, Vadim R.
Publicado: (2010)
por: Madgazin, Vadim R.
Publicado: (2010)
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
por: Yi, Yungang, et al.
Publicado: (2024)
por: Yi, Yungang, et al.
Publicado: (2024)
Quantum-Enhanced Analysis and Grading of Vocal Performance
por: Agarwal, Rohan
Publicado: (2025)
por: Agarwal, Rohan
Publicado: (2025)
BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
por: Park, Taesoo, et al.
Publicado: (2025)
por: Park, Taesoo, et al.
Publicado: (2025)
Prevailing Research Areas for Music AI in the Era of Foundation Models
por: Wei, Megan, et al.
Publicado: (2024)
por: Wei, Megan, et al.
Publicado: (2024)
Matcha-TTS: A fast TTS architecture with conditional flow matching
por: Mehta, Shivam, et al.
Publicado: (2023)
por: Mehta, Shivam, et al.
Publicado: (2023)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
por: Aristorenas, Aris J.
Publicado: (2024)
por: Aristorenas, Aris J.
Publicado: (2024)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
por: Mahfi, Muntahi Safwan, et al.
Publicado: (2025)
por: Mahfi, Muntahi Safwan, et al.
Publicado: (2025)
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
por: Tsangko, Iosif, et al.
Publicado: (2025)
por: Tsangko, Iosif, et al.
Publicado: (2025)
Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training
por: Park, Jisang, et al.
Publicado: (2024)
por: Park, Jisang, et al.
Publicado: (2024)
Sonify Anything: Towards Context-Aware Sonic Interactions in AR
por: Schütz, Laura, et al.
Publicado: (2025)
por: Schütz, Laura, et al.
Publicado: (2025)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
The evolution of inharmonicity and noisiness in contemporary popular music
por: Deruty, Emmanuel, et al.
Publicado: (2024)
por: Deruty, Emmanuel, et al.
Publicado: (2024)
Understanding the Algorithm Behind Audio Key Detection
por: Silva, Henrique Perez G.
Publicado: (2025)
por: Silva, Henrique Perez G.
Publicado: (2025)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
por: McKinnon, Max, et al.
Publicado: (2026)
por: McKinnon, Max, et al.
Publicado: (2026)
Generation of Musical Timbres using a Text-Guided Diffusion Model
por: Yuan, Weixuan, et al.
Publicado: (2025)
por: Yuan, Weixuan, et al.
Publicado: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
por: Li, Pengcheng, et al.
Publicado: (2024)
por: Li, Pengcheng, et al.
Publicado: (2024)
Taming Audio VAEs via Target-KL Regularization
por: Seetharaman, Prem, et al.
Publicado: (2026)
por: Seetharaman, Prem, et al.
Publicado: (2026)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
por: Bhadra, Dipayan, et al.
Publicado: (2025)
por: Bhadra, Dipayan, et al.
Publicado: (2025)
If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
por: Nunes, Caio, et al.
Publicado: (2025)
por: Nunes, Caio, et al.
Publicado: (2025)
Ejemplares similares
-
Automatic Album Sequencing
por: Herrmann, Vincent, et al.
Publicado: (2024) -
A Framework for Multimodal Medical Image Interaction
por: Schütz, Laura, et al.
Publicado: (2024) -
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
por: Woodard, Brandon, et al.
Publicado: (2025) -
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
por: Cho, Hyunsung, et al.
Publicado: (2024) -
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
por: Khushiyant, et al.
Publicado: (2026)