Gespeichert in:
| Hauptverfasser: | Sagasti, Amaia, Scaini, Davide, Arteaga, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2405.04471 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Quantifying Spatial Audio Quality Impairment
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023)
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
Deep learning based spatial aliasing reduction in beamforming for audio capture
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
Can Large Language Models Understand Spatial Audio?
von: Tang, Changli, et al.
Veröffentlicht: (2024)
von: Tang, Changli, et al.
Veröffentlicht: (2024)
Past, Present, and Future of Spatial Audio and Room Acoustics
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025)
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025)
Towards Spatial Audio Understanding via Question Answering
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
ASAudio: A Survey of Advanced Spatial Audio Research
von: Zhu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhiyuan, et al.
Veröffentlicht: (2025)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
von: Ronchini, Francesca, et al.
Veröffentlicht: (2020)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2020)
The Extended SONICOM HRTF Dataset and Spatial Audio Metrics Toolbox
von: Poole, Katarina C., et al.
Veröffentlicht: (2025)
von: Poole, Katarina C., et al.
Veröffentlicht: (2025)
SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
von: Hu, Jinbo, et al.
Veröffentlicht: (2025)
von: Hu, Jinbo, et al.
Veröffentlicht: (2025)
Room Impulse Response Generation Conditioned on Acoustic Parameters
von: Arellano, Silvia, et al.
Veröffentlicht: (2025)
von: Arellano, Silvia, et al.
Veröffentlicht: (2025)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
VR-PTOLEMAIC: A Virtual Environment for the Perceptual Testing of Spatial Audio Algorithms
von: Ostan, Paolo, et al.
Veröffentlicht: (2025)
von: Ostan, Paolo, et al.
Veröffentlicht: (2025)
UniSep: Universal Target Audio Separation with Language Models at Scale
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
Exploring the Potential of Data-Driven Spatial Audio Enhancement Using a Single-Channel Model
von: Santos, Arthur N. dos, et al.
Veröffentlicht: (2024)
von: Santos, Arthur N. dos, et al.
Veröffentlicht: (2024)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
von: Lee, Dongheon, et al.
Veröffentlicht: (2024)
von: Lee, Dongheon, et al.
Veröffentlicht: (2024)
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Room Impulse Response Synthesis via Differentiable Feedback Delay Networks for Efficient Spatial Audio Rendering
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
Stereo Audio Rendering for Personal Sound Zones Using a Binaural Spatially Adaptive Neural Network (BSANN)
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
Streaming Audio Transformers for Online Audio Tagging
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Pengi: An Audio Language Model for Audio Tasks
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
von: Chen, Yu, et al.
Veröffentlicht: (2025)
von: Chen, Yu, et al.
Veröffentlicht: (2025)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
Speaker Distance Estimation in Enclosures from Single-Channel Audio
von: Neri, Michael, et al.
Veröffentlicht: (2024)
von: Neri, Michael, et al.
Veröffentlicht: (2024)
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
PAM: Prompting Audio-Language Models for Audio Quality Assessment
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?
von: Namballa, Richa, et al.
Veröffentlicht: (2025)
von: Namballa, Richa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025) -
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023) -
Quantifying Spatial Audio Quality Impairment
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2023) -
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025) -
Deep learning based spatial aliasing reduction in beamforming for audio capture
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)