Towards Multimodal Query-Based Spatial Audio Source Extraction
Fuente:
arXiv
Guardado en:
| Autores principales: | Yu, Chenxin, Ma, Hao, Li, Xu, Zhang, Xiao-Lei, Shao, Mingjie, Zhang, Chi, Li, Xuelong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Eliminating Quantization Errors in Classification-Based Sound Source Localization
por: Feng, Linfeng, et al.
Publicado: (2023)
por: Feng, Linfeng, et al.
Publicado: (2023)
CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction
por: Ma, Hao, et al.
Publicado: (2024)
por: Ma, Hao, et al.
Publicado: (2024)
AudioSpa: Spatializing Sound Events with Text
por: Feng, Linfeng, et al.
Publicado: (2025)
por: Feng, Linfeng, et al.
Publicado: (2025)
Language-Queried Target Sound Extraction Without Parallel Training Data
por: Ma, Hao, et al.
Publicado: (2024)
por: Ma, Hao, et al.
Publicado: (2024)
High-Fidelity Generative Audio Compression at 0.275kbps
por: Ma, Hao, et al.
Publicado: (2026)
por: Ma, Hao, et al.
Publicado: (2026)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
por: Ma, Hao, et al.
Publicado: (2025)
por: Ma, Hao, et al.
Publicado: (2025)
Diffusion-Based Adversarial Purification for Speaker Verification
por: Bai, Yibo, et al.
Publicado: (2023)
por: Bai, Yibo, et al.
Publicado: (2023)
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization
por: Li, Xueqing, et al.
Publicado: (2025)
por: Li, Xueqing, et al.
Publicado: (2025)
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
por: Li, Longhao, et al.
Publicado: (2026)
por: Li, Longhao, et al.
Publicado: (2026)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
por: Xiao, Feiyang, et al.
Publicado: (2024)
por: Xiao, Feiyang, et al.
Publicado: (2024)
Multi-View Based Audio Visual Target Speaker Extraction
por: Yang, Peijun, et al.
Publicado: (2026)
por: Yang, Peijun, et al.
Publicado: (2026)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
por: Lei, Ke, et al.
Publicado: (2026)
por: Lei, Ke, et al.
Publicado: (2026)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
por: Yin, Han, et al.
Publicado: (2024)
por: Yin, Han, et al.
Publicado: (2024)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
por: Jing, Ruihao, et al.
Publicado: (2025)
por: Jing, Ruihao, et al.
Publicado: (2025)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
por: Saijo, Kohei, et al.
Publicado: (2024)
por: Saijo, Kohei, et al.
Publicado: (2024)
Online Audio-Visual Autoregressive Speaker Extraction
por: Pan, Zexu, et al.
Publicado: (2025)
por: Pan, Zexu, et al.
Publicado: (2025)
Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio Source Separation
por: Bai, Ye, et al.
Publicado: (2024)
por: Bai, Ye, et al.
Publicado: (2024)
Can Large Language Models Understand Spatial Audio?
por: Tang, Changli, et al.
Publicado: (2024)
por: Tang, Changli, et al.
Publicado: (2024)
Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
por: Chen, Zhang, et al.
Publicado: (2026)
por: Chen, Zhang, et al.
Publicado: (2026)
A Knowledge-Driven Approach to Target Speech Extraction in the Presence of Background Sound Effects for Cinematic Audio Source Separation (CASS)
por: Ho, Chun-wei, et al.
Publicado: (2026)
por: Ho, Chun-wei, et al.
Publicado: (2026)
Deep Learning Based Stage-wise Two-dimensional Speaker Localization with Large Ad-hoc Microphone Arrays
por: Liu, Shupei, et al.
Publicado: (2022)
por: Liu, Shupei, et al.
Publicado: (2022)
Towards Weakly Supervised Text-to-Audio Grounding
por: Xu, Xuenan, et al.
Publicado: (2024)
por: Xu, Xuenan, et al.
Publicado: (2024)
$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
por: Zhu, Boyu, et al.
Publicado: (2025)
por: Zhu, Boyu, et al.
Publicado: (2025)
Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
por: Zhang, Xiangyu, et al.
Publicado: (2024)
por: Zhang, Xiangyu, et al.
Publicado: (2024)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
por: Hai, Jiarui, et al.
Publicado: (2024)
por: Hai, Jiarui, et al.
Publicado: (2024)
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
por: Li, Zixuan, et al.
Publicado: (2025)
por: Li, Zixuan, et al.
Publicado: (2025)
Attention-Based Audio Embeddings for Query-by-Example
por: Singh, Anup, et al.
Publicado: (2022)
por: Singh, Anup, et al.
Publicado: (2022)
Towards Neural Audio Codec Source Parsing
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
por: Zhao, Lei, et al.
Publicado: (2025)
por: Zhao, Lei, et al.
Publicado: (2025)
Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database
por: Yaffe, Danielle, et al.
Publicado: (2025)
por: Yaffe, Danielle, et al.
Publicado: (2025)
ASAudio: A Survey of Advanced Spatial Audio Research
por: Zhu, Zhiyuan, et al.
Publicado: (2025)
por: Zhu, Zhiyuan, et al.
Publicado: (2025)
Towards Spatial Audio Understanding via Question Answering
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
por: Li, Chenxing, et al.
Publicado: (2024)
por: Li, Chenxing, et al.
Publicado: (2024)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
por: Shetu, Shrishti Saha, et al.
Publicado: (2025)
Spatial Audio Signal Enhancement: A Multi-output MVDR Method in The Spherical Harmonic-domain
por: Zhang, Huawei, et al.
Publicado: (2024)
por: Zhang, Huawei, et al.
Publicado: (2024)
Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction
por: Pan, Zexu, et al.
Publicado: (2026)
por: Pan, Zexu, et al.
Publicado: (2026)
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
por: Geng, Xuelong, et al.
Publicado: (2024)
por: Geng, Xuelong, et al.
Publicado: (2024)
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
por: Greif, Jonathan, et al.
Publicado: (2024)
por: Greif, Jonathan, et al.
Publicado: (2024)
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
por: Shi, Runwu, et al.
Publicado: (2025)
por: Shi, Runwu, et al.
Publicado: (2025)
Ejemplares similares
-
Eliminating Quantization Errors in Classification-Based Sound Source Localization
por: Feng, Linfeng, et al.
Publicado: (2023) -
CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction
por: Ma, Hao, et al.
Publicado: (2024) -
AudioSpa: Spatializing Sound Events with Text
por: Feng, Linfeng, et al.
Publicado: (2025) -
Language-Queried Target Sound Extraction Without Parallel Training Data
por: Ma, Hao, et al.
Publicado: (2024) -
High-Fidelity Generative Audio Compression at 0.275kbps
por: Ma, Hao, et al.
Publicado: (2026)