PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
Fuente:
arXiv
Guardado en:
| Autores principales: | Vosoughi, Ali, Zang, Yongyi, Yang, Qihui, Paek, Nathan, Leistikow, Randal, Xu, Chenliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
por: Yang, Qihui, et al.
Publicado: (2025)
por: Yang, Qihui, et al.
Publicado: (2025)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
por: Paek, Nathan, et al.
Publicado: (2025)
por: Paek, Nathan, et al.
Publicado: (2025)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
por: Zhu, Di, et al.
Publicado: (2026)
por: Zhu, Di, et al.
Publicado: (2026)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
por: Dhiman, Jai
Publicado: (2026)
por: Dhiman, Jai
Publicado: (2026)
Refining music sample identification with a self-supervised graph neural network
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
Efficient Vocal Source Separation Through Windowed Sink Attention
por: Benetatos, Christodoulos, et al.
Publicado: (2025)
por: Benetatos, Christodoulos, et al.
Publicado: (2025)
Quantum-Enhanced Analysis and Grading of Vocal Performance
por: Agarwal, Rohan
Publicado: (2025)
por: Agarwal, Rohan
Publicado: (2025)
GraFPrint: A GNN-Based Approach for Audio Identification
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
por: Aristorenas, Aris J.
Publicado: (2024)
por: Aristorenas, Aris J.
Publicado: (2024)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
Step-Audio-R1 Technical Report
por: Tian, Fei, et al.
Publicado: (2025)
por: Tian, Fei, et al.
Publicado: (2025)
Online Correction of Dispersion Error in 2D Waveguide Meshes
por: Fontana, Federico, et al.
Publicado: (2000)
por: Fontana, Federico, et al.
Publicado: (2000)
Adaptable Symbolic Music Infilling with MIDI-RWKV
por: Zhou-Zheng, Christian, et al.
Publicado: (2025)
por: Zhou-Zheng, Christian, et al.
Publicado: (2025)
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
por: Cauzinille, Jules, et al.
Publicado: (2025)
por: Cauzinille, Jules, et al.
Publicado: (2025)
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
por: Wang, Zihao, et al.
Publicado: (2025)
por: Wang, Zihao, et al.
Publicado: (2025)
Dichotic harmony for the musical practice
por: Madgazin, Vadim R.
Publicado: (2010)
por: Madgazin, Vadim R.
Publicado: (2010)
Machine learning based animal emotion classification using audio signals
por: Slobodian, Mariia, et al.
Publicado: (2025)
por: Slobodian, Mariia, et al.
Publicado: (2025)
Masked Contrastive Pre-Training Improves Music Audio Key Detection
por: Yonay, Ori, et al.
Publicado: (2026)
por: Yonay, Ori, et al.
Publicado: (2026)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026)
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026)
Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
por: Geleta, Margarita, et al.
Publicado: (2025)
por: Geleta, Margarita, et al.
Publicado: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
por: Marmoret, Axel
Publicado: (2026)
por: Marmoret, Axel
Publicado: (2026)
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
por: Tsangko, Iosif, et al.
Publicado: (2025)
por: Tsangko, Iosif, et al.
Publicado: (2025)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
por: Khushiyant, et al.
Publicado: (2026)
por: Khushiyant, et al.
Publicado: (2026)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
por: Richter-Powell, Jessie, et al.
Publicado: (2025)
por: Richter-Powell, Jessie, et al.
Publicado: (2025)
BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
por: Park, Taesoo, et al.
Publicado: (2025)
por: Park, Taesoo, et al.
Publicado: (2025)
Embodied Exploration of Latent Spaces and Explainable AI
por: Wilson, Elizabeth, et al.
Publicado: (2024)
por: Wilson, Elizabeth, et al.
Publicado: (2024)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
por: Mehta, Shivam, et al.
Publicado: (2024)
por: Mehta, Shivam, et al.
Publicado: (2024)
Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training
por: Park, Jisang, et al.
Publicado: (2024)
por: Park, Jisang, et al.
Publicado: (2024)
Signal-Theoretic Characterization of Waveguide Mesh Geometries for Models of Two--Dimensional Wave Propagation in Elastic Media
por: Fontana, Federico, et al.
Publicado: (2001)
por: Fontana, Federico, et al.
Publicado: (2001)
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
por: Qian, Lekai, et al.
Publicado: (2026)
por: Qian, Lekai, et al.
Publicado: (2026)
The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
por: Chen, Kuan-Yu, et al.
Publicado: (2026)
por: Chen, Kuan-Yu, et al.
Publicado: (2026)
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
por: Khah, Arman Nik, et al.
Publicado: (2026)
por: Khah, Arman Nik, et al.
Publicado: (2026)
Automatic Album Sequencing
por: Herrmann, Vincent, et al.
Publicado: (2024)
por: Herrmann, Vincent, et al.
Publicado: (2024)
If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
por: Nunes, Caio, et al.
Publicado: (2025)
por: Nunes, Caio, et al.
Publicado: (2025)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
por: Shu, Hongzhi, et al.
Publicado: (2024)
por: Shu, Hongzhi, et al.
Publicado: (2024)
Generation of Musical Timbres using a Text-Guided Diffusion Model
por: Yuan, Weixuan, et al.
Publicado: (2025)
por: Yuan, Weixuan, et al.
Publicado: (2025)
Ejemplares similares
-
FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
por: Yang, Qihui, et al.
Publicado: (2025) -
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
por: Paek, Nathan, et al.
Publicado: (2025) -
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
por: Bhattacharjee, Aditya, et al.
Publicado: (2025) -
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
por: Zhu, Di, et al.
Publicado: (2026) -
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
por: Dhiman, Jai
Publicado: (2026)