Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Runwu, Itoyama, Katsutoshi, Nakadai, Kazuhiro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Can all variations within the unified mask-based beamformer framework achieve identical peak extraction performance?
por: Hiroe, Atsuo, et al.
Publicado: (2024)
por: Hiroe, Atsuo, et al.
Publicado: (2024)
Distance Based Single-Channel Target Speech Extraction
por: Shi, Runwu, et al.
Publicado: (2024)
por: Shi, Runwu, et al.
Publicado: (2024)
Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
por: Shi, Runwu, et al.
Publicado: (2025)
por: Shi, Runwu, et al.
Publicado: (2025)
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
por: Shi, Runwu, et al.
Publicado: (2025)
por: Shi, Runwu, et al.
Publicado: (2025)
Live Vocal Extraction from K-pop Performances
por: Kim, Yujin, et al.
Publicado: (2025)
por: Kim, Yujin, et al.
Publicado: (2025)
Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
por: Wang, Jiang, et al.
Publicado: (2025)
por: Wang, Jiang, et al.
Publicado: (2025)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
por: Maghsoudi, Maryam, et al.
Publicado: (2026)
por: Maghsoudi, Maryam, et al.
Publicado: (2026)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
por: Sato, Hiroshi, et al.
Publicado: (2025)
por: Sato, Hiroshi, et al.
Publicado: (2025)
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
por: Shi, Runwu, et al.
Publicado: (2025)
por: Shi, Runwu, et al.
Publicado: (2025)
Weakly Supervised Detection and Temporal Localization of Whale Calls in Long-Duration Bioacoustic Data
por: Nihal, Ragib Amin, et al.
Publicado: (2025)
por: Nihal, Ragib Amin, et al.
Publicado: (2025)
An Efficient GPU-based Implementation for Noise Robust Sound Source Localization
por: Lin, Zirui, et al.
Publicado: (2025)
por: Lin, Zirui, et al.
Publicado: (2025)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
por: Premananth, Gowtham, et al.
Publicado: (2024)
por: Premananth, Gowtham, et al.
Publicado: (2024)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
por: Zaiem, Salah, et al.
Publicado: (2023)
por: Zaiem, Salah, et al.
Publicado: (2023)
Binaural Selective Attention Model for Target Speaker Extraction
por: Meng, Hanyu, et al.
Publicado: (2024)
por: Meng, Hanyu, et al.
Publicado: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
por: Wilkins, Julia, et al.
Publicado: (2024)
por: Wilkins, Julia, et al.
Publicado: (2024)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
por: Wilkins, Julia, et al.
Publicado: (2025)
por: Wilkins, Julia, et al.
Publicado: (2025)
Online Similarity-and-Independence-Aware Beamformer for Low-latency Target Sound Extraction
por: Hiroe, Atsuo
Publicado: (2023)
por: Hiroe, Atsuo
Publicado: (2023)
Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
por: Shore, Noah
Publicado: (2025)
por: Shore, Noah
Publicado: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
por: Serre, Thomas, et al.
Publicado: (2026)
por: Serre, Thomas, et al.
Publicado: (2026)
Phase-Based Signal Representations for Scattering
por: Haider, Daniel, et al.
Publicado: (2022)
por: Haider, Daniel, et al.
Publicado: (2022)
Blind Source Separation of Radar Signals in Time Domain Using Deep Learning
por: Hinderer, Sven
Publicado: (2025)
por: Hinderer, Sven
Publicado: (2025)
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
por: Phan, Dang Thoai, et al.
Publicado: (2025)
por: Phan, Dang Thoai, et al.
Publicado: (2025)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
por: Sato, Hiroshi, et al.
Publicado: (2024)
por: Sato, Hiroshi, et al.
Publicado: (2024)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
por: Gu, Yicheng, et al.
Publicado: (2024)
por: Gu, Yicheng, et al.
Publicado: (2024)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
por: Ziogas, Ioannis, et al.
Publicado: (2024)
por: Ziogas, Ioannis, et al.
Publicado: (2024)
Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection
por: Liao, Yuan, et al.
Publicado: (2025)
por: Liao, Yuan, et al.
Publicado: (2025)
Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
por: Abbott, Leigh, et al.
Publicado: (2024)
por: Abbott, Leigh, et al.
Publicado: (2024)
A Machine Hearing System for Robust Cough Detection Based on a High-Level Representation of Band-Specific Audio Features
por: Monge-Alvarez, Jesús, et al.
Publicado: (2024)
por: Monge-Alvarez, Jesús, et al.
Publicado: (2024)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
por: Chen, Xiaodan, et al.
Publicado: (2025)
por: Chen, Xiaodan, et al.
Publicado: (2025)
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
por: Wang, Bo, et al.
Publicado: (2024)
por: Wang, Bo, et al.
Publicado: (2024)
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
por: Zhu, Haolin, et al.
Publicado: (2024)
por: Zhu, Haolin, et al.
Publicado: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
por: Shetu, Shrishti Saha, et al.
Publicado: (2024)
Audio Compression using Periodic Gabor with Biorthogonal Exchange: Implementation Using the Zak Transform
por: Alimi, Roger, et al.
Publicado: (2025)
por: Alimi, Roger, et al.
Publicado: (2025)
Using Neurogram Similarity Index Measure (NSIM) to Model Hearing Loss and Cochlear Neural Degeneration
por: Cheema, Ahsan J., et al.
Publicado: (2025)
por: Cheema, Ahsan J., et al.
Publicado: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
por: Aihara, Ryo, et al.
Publicado: (2025)
por: Aihara, Ryo, et al.
Publicado: (2025)
Multiple Mobile Target Detection and Tracking in Active Sonar Array Using a Track-Before-Detect Approach
por: Abu, Avi, et al.
Publicado: (2024)
por: Abu, Avi, et al.
Publicado: (2024)
Learning Perceptually Relevant Temporal Envelope Morphing
por: Dixit, Satvik, et al.
Publicado: (2025)
por: Dixit, Satvik, et al.
Publicado: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
por: Haeb-Umbach, Reinhold, et al.
Publicado: (2025)
por: Haeb-Umbach, Reinhold, et al.
Publicado: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
por: Zhang, Wangyou, et al.
Publicado: (2025)
por: Zhang, Wangyou, et al.
Publicado: (2025)
Ejemplares similares
-
Can all variations within the unified mask-based beamformer framework achieve identical peak extraction performance?
por: Hiroe, Atsuo, et al.
Publicado: (2024) -
Distance Based Single-Channel Target Speech Extraction
por: Shi, Runwu, et al.
Publicado: (2024) -
Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
por: Shi, Runwu, et al.
Publicado: (2025) -
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
por: Shi, Runwu, et al.
Publicado: (2025) -
Live Vocal Extraction from K-pop Performances
por: Kim, Yujin, et al.
Publicado: (2025)