Learning Spatially-Aware Language and Audio Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Devnani, Bhavika, Seto, Skyler, Aldeneh, Zakaria, Toso, Alessandro, Menyaylenko, Elena, Theobald, Barry-John, Sheaffer, Jonathan, Sarabia, Miguel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024)
por: Chen, Li-Wei, et al.
Publicado: (2024)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
por: Jung, Jee-weon, et al.
Publicado: (2024)
por: Jung, Jee-weon, et al.
Publicado: (2024)
StereoFoley: Object-Aware Stereo Audio Generation from Video
por: Karchkhadze, Tornike, et al.
Publicado: (2025)
por: Karchkhadze, Tornike, et al.
Publicado: (2025)
SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
por: Hu, Jinbo, et al.
Publicado: (2025)
por: Hu, Jinbo, et al.
Publicado: (2025)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
por: Wada, Aogu, et al.
Publicado: (2025)
por: Wada, Aogu, et al.
Publicado: (2025)
Binamix -- A Python Library for Generating Binaural Audio Datasets
por: Barry, Dan, et al.
Publicado: (2025)
por: Barry, Dan, et al.
Publicado: (2025)
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
por: Panah, Davoud Shariat, et al.
Publicado: (2025)
por: Panah, Davoud Shariat, et al.
Publicado: (2025)
Can Large Language Models Understand Spatial Audio?
por: Tang, Changli, et al.
Publicado: (2024)
por: Tang, Changli, et al.
Publicado: (2024)
NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
por: Ragano, Alessandro, et al.
Publicado: (2023)
por: Ragano, Alessandro, et al.
Publicado: (2023)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
por: Wu, Shih-Lun, et al.
Publicado: (2023)
por: Wu, Shih-Lun, et al.
Publicado: (2023)
Universal Spatial Audio Transcoder
por: Sagasti, Amaia, et al.
Publicado: (2024)
por: Sagasti, Amaia, et al.
Publicado: (2024)
Novel-View Acoustic Synthesis from 3D Reconstructed Rooms
por: Ahn, Byeongjoo, et al.
Publicado: (2023)
por: Ahn, Byeongjoo, et al.
Publicado: (2023)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
por: Tailleur, Modan, et al.
Publicado: (2024)
por: Tailleur, Modan, et al.
Publicado: (2024)
Quantifying Spatial Audio Quality Impairment
por: Watcharasupat, Karn N., et al.
Publicado: (2023)
por: Watcharasupat, Karn N., et al.
Publicado: (2023)
Attention-Based Audio Embeddings for Query-by-Example
por: Singh, Anup, et al.
Publicado: (2022)
por: Singh, Anup, et al.
Publicado: (2022)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
por: Shan, Weiqiao, et al.
Publicado: (2025)
por: Shan, Weiqiao, et al.
Publicado: (2025)
AudioSpa: Spatializing Sound Events with Text
por: Feng, Linfeng, et al.
Publicado: (2025)
por: Feng, Linfeng, et al.
Publicado: (2025)
Region-Specific Audio Tagging for Spatial Sound
por: Zhao, Jinzheng, et al.
Publicado: (2025)
por: Zhao, Jinzheng, et al.
Publicado: (2025)
Pengi: An Audio Language Model for Audio Tasks
por: Deshmukh, Soham, et al.
Publicado: (2023)
por: Deshmukh, Soham, et al.
Publicado: (2023)
Past, Present, and Future of Spatial Audio and Room Acoustics
por: Koyama, Shoichi, et al.
Publicado: (2025)
por: Koyama, Shoichi, et al.
Publicado: (2025)
Towards Spatial Audio Understanding via Question Answering
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
ASAudio: A Survey of Advanced Spatial Audio Research
por: Zhu, Zhiyuan, et al.
Publicado: (2025)
por: Zhu, Zhiyuan, et al.
Publicado: (2025)
Measuring Audio Prompt Adherence with Distribution-based Embedding Distances
por: Grachten, Maarten
Publicado: (2024)
por: Grachten, Maarten
Publicado: (2024)
Spatial-Aware Conditioned Fusion for Audio-Visual Navigation
por: Wu, Shaohang, et al.
Publicado: (2026)
por: Wu, Shaohang, et al.
Publicado: (2026)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
por: Primus, Paul, et al.
Publicado: (2024)
por: Primus, Paul, et al.
Publicado: (2024)
The Extended SONICOM HRTF Dataset and Spatial Audio Metrics Toolbox
por: Poole, Katarina C., et al.
Publicado: (2025)
por: Poole, Katarina C., et al.
Publicado: (2025)
PAM: Prompting Audio-Language Models for Audio Quality Assessment
por: Deshmukh, Soham, et al.
Publicado: (2024)
por: Deshmukh, Soham, et al.
Publicado: (2024)
Continuous Audio Language Models
por: Rouard, Simon, et al.
Publicado: (2025)
por: Rouard, Simon, et al.
Publicado: (2025)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
por: Chu, Annie, et al.
Publicado: (2024)
por: Chu, Annie, et al.
Publicado: (2024)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
por: Xue, Jinlong, et al.
Publicado: (2024)
por: Xue, Jinlong, et al.
Publicado: (2024)
Audio Inpainting in Time-Frequency Domain with Phase-Aware Prior
por: Balušík, Peter, et al.
Publicado: (2026)
por: Balušík, Peter, et al.
Publicado: (2026)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
por: Dinkel, Heinrich, et al.
Publicado: (2026)
por: Dinkel, Heinrich, et al.
Publicado: (2026)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
por: Wang, Yuanyuan, et al.
Publicado: (2024)
por: Wang, Yuanyuan, et al.
Publicado: (2024)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
por: Kushwaha, Saksham Singh, et al.
Publicado: (2024)
por: Kushwaha, Saksham Singh, et al.
Publicado: (2024)
VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
por: Ostan, Paolo, et al.
Publicado: (2025)
por: Ostan, Paolo, et al.
Publicado: (2025)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
por: Xiao, Feiyang, et al.
Publicado: (2024)
por: Xiao, Feiyang, et al.
Publicado: (2024)
Ejemplares similares
-
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
por: Aldeneh, Zakaria, et al.
Publicado: (2024) -
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
por: Aldeneh, Zakaria, et al.
Publicado: (2024) -
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024) -
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
por: Aldeneh, Zakaria, et al.
Publicado: (2024) -
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
por: Jung, Jee-weon, et al.
Publicado: (2024)