AudioSpa: Spatializing Sound Events with Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Linfeng, Zhao, Lei, Zhu, Boyu, Zhang, Xiao-Lei, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
Deep Learning Based Stage-wise Two-dimensional Speaker Localization with Large Ad-hoc Microphone Arrays
von: Liu, Shupei, et al.
Veröffentlicht: (2022)
von: Liu, Shupei, et al.
Veröffentlicht: (2022)
High-Fidelity Generative Audio Compression at 0.275kbps
von: Ma, Hao, et al.
Veröffentlicht: (2026)
von: Ma, Hao, et al.
Veröffentlicht: (2026)
Diffusion-Based Adversarial Purification for Speaker Verification
von: Bai, Yibo, et al.
Veröffentlicht: (2023)
von: Bai, Yibo, et al.
Veröffentlicht: (2023)
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
Eliminating Quantization Errors in Classification-Based Sound Source Localization
von: Feng, Linfeng, et al.
Veröffentlicht: (2023)
von: Feng, Linfeng, et al.
Veröffentlicht: (2023)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
ASAudio: A Survey of Advanced Spatial Audio Research
von: Zhu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhiyuan, et al.
Veröffentlicht: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression Localization
von: Zhang, Xueping, et al.
Veröffentlicht: (2025)
von: Zhang, Xueping, et al.
Veröffentlicht: (2025)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
von: Wang, Pingjie, et al.
Veröffentlicht: (2024)
von: Wang, Pingjie, et al.
Veröffentlicht: (2024)
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
FSD50K-Solo: Automated Curation of Single-Source Sound Events
von: Yang, Ningyuan, et al.
Veröffentlicht: (2026)
von: Yang, Ningyuan, et al.
Veröffentlicht: (2026)
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
First-Shot Unsupervised Anomalous Sound Detection With Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event Detection
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
Unified Audio Event Detection
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Sound Event Bounding Boxes
von: Ebbers, Janek, et al.
Veröffentlicht: (2024)
von: Ebbers, Janek, et al.
Veröffentlicht: (2024)
Stereo Audio Rendering for Personal Sound Zones Using a Binaural Spatially Adaptive Neural Network (BSANN)
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025) -
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024) -
Deep Learning Based Stage-wise Two-dimensional Speaker Localization with Large Ad-hoc Microphone Arrays
von: Liu, Shupei, et al.
Veröffentlicht: (2022) -
High-Fidelity Generative Audio Compression at 0.275kbps
von: Ma, Hao, et al.
Veröffentlicht: (2026) -
Diffusion-Based Adversarial Purification for Speaker Verification
von: Bai, Yibo, et al.
Veröffentlicht: (2023)