Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
Fuente:
arXiv
Guardado en:
| Autores principales: | Heikkinen, Mikko, Politis, Archontis, Drossos, Konstantinos, Virtanen, Tuomas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Omnidirectional: Neural Ambisonics Encoding for Arbitrary Microphone Directivity Patterns using Cross-Attention
por: Heikkinen, Mikko, et al.
Publicado: (2026)
por: Heikkinen, Mikko, et al.
Publicado: (2026)
Neural Ambisonics encoding for compact irregular microphone arrays
por: Heikkinen, Mikko, et al.
Publicado: (2024)
por: Heikkinen, Mikko, et al.
Publicado: (2024)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
por: Luong, Diep, et al.
Publicado: (2025)
por: Luong, Diep, et al.
Publicado: (2025)
Automatic Contextual Audio Denoising
por: Luong, Diep, et al.
Publicado: (2026)
por: Luong, Diep, et al.
Publicado: (2026)
Multi-Utterance Speech Separation and Association Trained on Short Segments
por: Wang, Yuzhu, et al.
Publicado: (2025)
por: Wang, Yuzhu, et al.
Publicado: (2025)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
por: Wang, Yuzhu, et al.
Publicado: (2025)
por: Wang, Yuzhu, et al.
Publicado: (2025)
Moving Speaker Separation via Parallel Spectral-Spatial Processing
por: Wang, Yuzhu, et al.
Publicado: (2026)
por: Wang, Yuzhu, et al.
Publicado: (2026)
Score-informed Music Source Separation: Improving Synthetic-to-real Generalization in Classical Music
por: Tunturi, Eetu, et al.
Publicado: (2025)
por: Tunturi, Eetu, et al.
Publicado: (2025)
Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
por: Drossos, Konstantinos, et al.
Publicado: (2025)
por: Drossos, Konstantinos, et al.
Publicado: (2025)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
por: Dai, Wang, et al.
Publicado: (2025)
por: Dai, Wang, et al.
Publicado: (2025)
Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
por: Politis, Archontis
Publicado: (2024)
por: Politis, Archontis
Publicado: (2024)
Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications
por: Diaz-Guerra, David, et al.
Publicado: (2023)
por: Diaz-Guerra, David, et al.
Publicado: (2023)
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
por: Neri, Michael, et al.
Publicado: (2026)
por: Neri, Michael, et al.
Publicado: (2026)
Speaker Distance Estimation in Enclosures from Single-Channel Audio
por: Neri, Michael, et al.
Publicado: (2024)
por: Neri, Michael, et al.
Publicado: (2024)
Representation Learning for Audio Privacy Preservation using Source Separation and Robust Adversarial Learning
por: Luong, Diep, et al.
Publicado: (2023)
por: Luong, Diep, et al.
Publicado: (2023)
Adversarial Representation Learning for Robust Privacy Preservation in Audio
por: Gharib, Shayan, et al.
Publicado: (2023)
por: Gharib, Shayan, et al.
Publicado: (2023)
SynthSOD: Developing an Heterogeneous Dataset for Orchestra Music Source Separation
por: Garcia-Martinez, Jaime, et al.
Publicado: (2024)
por: Garcia-Martinez, Jaime, et al.
Publicado: (2024)
Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
por: Silaev, Mikhail, et al.
Publicado: (2026)
por: Silaev, Mikhail, et al.
Publicado: (2026)
Sound Event Detection and Localization with Distance Estimation
por: Krause, Daniel Aleksander, et al.
Publicado: (2024)
por: Krause, Daniel Aleksander, et al.
Publicado: (2024)
Towards Spatial Audio Understanding via Question Answering
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
Multi-Channel Replay Speech Detection using Acoustic Maps
por: Neri, Michael, et al.
Publicado: (2026)
por: Neri, Michael, et al.
Publicado: (2026)
Inter-Speaker Relative Cues for Two-Stage Text-Guided Target Speech Extraction
por: Dai, Wang, et al.
Publicado: (2026)
por: Dai, Wang, et al.
Publicado: (2026)
From Weak to Strong Sound Event Labels using Adaptive Change-Point Detection and Active Learning
por: Martinsson, John, et al.
Publicado: (2024)
por: Martinsson, John, et al.
Publicado: (2024)
Multi-label Zero-Shot Audio Classification with Temporal Attention
por: Dogan, Duygu, et al.
Publicado: (2024)
por: Dogan, Duygu, et al.
Publicado: (2024)
Array-Aware Ambisonics and HRTF Encoding for Binaural Reproduction With Wearable Arrays
por: Gayer, Yhonatan, et al.
Publicado: (2025)
por: Gayer, Yhonatan, et al.
Publicado: (2025)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
por: Zhang, Shiqi, et al.
Publicado: (2025)
por: Zhang, Shiqi, et al.
Publicado: (2025)
A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays
por: Miotello, Federico, et al.
Publicado: (2024)
por: Miotello, Federico, et al.
Publicado: (2024)
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
por: Dai, Wang, et al.
Publicado: (2024)
por: Dai, Wang, et al.
Publicado: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
por: Moritz, Martin, et al.
Publicado: (2024)
por: Moritz, Martin, et al.
Publicado: (2024)
Ambisonizer: Neural Upmixing as Spherical Harmonics Generation
por: Zang, Yongyi, et al.
Publicado: (2024)
por: Zang, Yongyi, et al.
Publicado: (2024)
Compression of Higher Order Ambisonics with Multichannel RVQGAN
por: Hirvonen, Toni, et al.
Publicado: (2024)
por: Hirvonen, Toni, et al.
Publicado: (2024)
Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
por: Cohen, Ohad, et al.
Publicado: (2024)
por: Cohen, Ohad, et al.
Publicado: (2024)
Residual Learning for Neural Ambisonics Encoders
por: Deppisch, Thomas, et al.
Publicado: (2026)
por: Deppisch, Thomas, et al.
Publicado: (2026)
Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array
por: Qiao, Yue, et al.
Publicado: (2024)
por: Qiao, Yue, et al.
Publicado: (2024)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
por: Wen, Wen, et al.
Publicado: (2024)
por: Wen, Wen, et al.
Publicado: (2024)
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
por: Ronchini, Francesca, et al.
Publicado: (2024)
por: Ronchini, Francesca, et al.
Publicado: (2024)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
por: Neri, Michael, et al.
Publicado: (2025)
por: Neri, Michael, et al.
Publicado: (2025)
Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier
por: Dumpis, Martynas, et al.
Publicado: (2026)
por: Dumpis, Martynas, et al.
Publicado: (2026)
Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
por: Hakala, Aapo, et al.
Publicado: (2025)
por: Hakala, Aapo, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond Omnidirectional: Neural Ambisonics Encoding for Arbitrary Microphone Directivity Patterns using Cross-Attention
por: Heikkinen, Mikko, et al.
Publicado: (2026) -
Neural Ambisonics encoding for compact irregular microphone arrays
por: Heikkinen, Mikko, et al.
Publicado: (2024) -
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
por: Luong, Diep, et al.
Publicado: (2025) -
Automatic Contextual Audio Denoising
por: Luong, Diep, et al.
Publicado: (2026) -
Multi-Utterance Speech Separation and Association Trained on Short Segments
por: Wang, Yuzhu, et al.
Publicado: (2025)