Rank-based loss for learning hierarchical representations
Fuente:
arXiv
Guardado en:
| Autores principales: | Nolasco, Ines, Stowell, Dan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2021
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Acoustic identification of individual animals with hierarchical contrastive learning
por: Nolasco, Ines, et al.
Publicado: (2024)
por: Nolasco, Ines, et al.
Publicado: (2024)
Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
por: Ghani, Burooj, et al.
Publicado: (2024)
por: Ghani, Burooj, et al.
Publicado: (2024)
Computational bioacoustics with deep learning: a review and roadmap
por: Stowell, Dan
Publicado: (2021)
por: Stowell, Dan
Publicado: (2021)
InsectSet459: an open dataset of insect sounds for bioacoustic machine learning
por: Faiß, Marius, et al.
Publicado: (2025)
por: Faiß, Marius, et al.
Publicado: (2025)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
por: Pepino, Leonardo, et al.
Publicado: (2023)
por: Pepino, Leonardo, et al.
Publicado: (2023)
Transformation of audio embeddings into interpretable, concept-based representations
por: Zhang, Alice, et al.
Publicado: (2025)
por: Zhang, Alice, et al.
Publicado: (2025)
The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
por: Dabike, Gerardo Roa, et al.
Publicado: (2024)
por: Dabike, Gerardo Roa, et al.
Publicado: (2024)
Automatic acoustic detection of birds through deep learning: the first Bird Audio Detection challenge
por: Stowell, Dan, et al.
Publicado: (2018)
por: Stowell, Dan, et al.
Publicado: (2018)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
por: Sadok, Samir, et al.
Publicado: (2023)
por: Sadok, Samir, et al.
Publicado: (2023)
Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers
por: Souza, Gabriel, et al.
Publicado: (2024)
por: Souza, Gabriel, et al.
Publicado: (2024)
Single-channel speech enhancement using learnable loss mixup
por: Chang, Oscar, et al.
Publicado: (2023)
por: Chang, Oscar, et al.
Publicado: (2023)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
por: Kim, Yunsik, et al.
Publicado: (2025)
por: Kim, Yunsik, et al.
Publicado: (2025)
Adaptive Representations of Sound for Automatic Insect Recognition
por: Faiß, Marius, et al.
Publicado: (2023)
por: Faiß, Marius, et al.
Publicado: (2023)
Late fusion ensembles for speech recognition on diverse input audio representations
por: Jezidžić, Marin, et al.
Publicado: (2024)
por: Jezidžić, Marin, et al.
Publicado: (2024)
Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
por: Laakkonen, Janne, et al.
Publicado: (2025)
por: Laakkonen, Janne, et al.
Publicado: (2025)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
por: Maharana, Sarthak Kumar, et al.
Publicado: (2023)
por: Maharana, Sarthak Kumar, et al.
Publicado: (2023)
Improved symbolic drum style classification with grammar-based hierarchical representations
por: Géré, Léo, et al.
Publicado: (2024)
por: Géré, Léo, et al.
Publicado: (2024)
Selfsupervised learning for pathological speech detection
por: Sheikh, Shakeel Ahmad
Publicado: (2024)
por: Sheikh, Shakeel Ahmad
Publicado: (2024)
Self-supervised learning of speech representations with Dutch archival data
por: Vaessen, Nik, et al.
Publicado: (2025)
por: Vaessen, Nik, et al.
Publicado: (2025)
Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
por: Bonafos, Guillem, et al.
Publicado: (2024)
por: Bonafos, Guillem, et al.
Publicado: (2024)
A contrastive-learning approach for auditory attention detection
por: Bajestan, Seyed Ali Alavi, et al.
Publicado: (2024)
por: Bajestan, Seyed Ali Alavi, et al.
Publicado: (2024)
Generalizable speech deepfake detection via meta-learned LoRA
por: Laakkonen, Janne, et al.
Publicado: (2025)
por: Laakkonen, Janne, et al.
Publicado: (2025)
A low latency attention module for streaming self-supervised speech representation learning
por: Ma, Jianbo, et al.
Publicado: (2023)
por: Ma, Jianbo, et al.
Publicado: (2023)
Gradient-based Optimisation of Modulation Effects
por: Carson, Alistair, et al.
Publicado: (2026)
por: Carson, Alistair, et al.
Publicado: (2026)
An Analysis of the Variance of Diffusion-based Speech Enhancement
por: Lay, Bunlong, et al.
Publicado: (2024)
por: Lay, Bunlong, et al.
Publicado: (2024)
Towards Audio Codec-based Speech Separation
por: Yip, Jia Qi, et al.
Publicado: (2024)
por: Yip, Jia Qi, et al.
Publicado: (2024)
A Concept-based approach to Voice Disorder Detection
por: Ghia, Davide, et al.
Publicado: (2025)
por: Ghia, Davide, et al.
Publicado: (2025)
HRTF Estimation using a Score-based Prior
por: Thuillier, Etienne, et al.
Publicado: (2024)
por: Thuillier, Etienne, et al.
Publicado: (2024)
Speech Enhancement and Dereverberation with Diffusion-based Generative Models
por: Richter, Julius, et al.
Publicado: (2022)
por: Richter, Julius, et al.
Publicado: (2022)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
por: Ulgen, Ismail Rasim, et al.
Publicado: (2026)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2026)
Audio-based automatic mating success prediction of giant pandas
por: Yan, WeiRan, et al.
Publicado: (2019)
por: Yan, WeiRan, et al.
Publicado: (2019)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
por: Lee, Philip H., et al.
Publicado: (2024)
por: Lee, Philip H., et al.
Publicado: (2024)
Wind Noise Reduction with a Diffusion-based Stochastic Regeneration Model
por: Lemercier, Jean-Marie, et al.
Publicado: (2023)
por: Lemercier, Jean-Marie, et al.
Publicado: (2023)
Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription
por: Yeung, Michael, et al.
Publicado: (2025)
por: Yeung, Michael, et al.
Publicado: (2025)
Clustering-based hard negative sampling for supervised contrastive speaker verification
por: Masztalski, Piotr, et al.
Publicado: (2025)
por: Masztalski, Piotr, et al.
Publicado: (2025)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
por: Primus, Paul, et al.
Publicado: (2024)
por: Primus, Paul, et al.
Publicado: (2024)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
por: Gebhard, Alexander, et al.
Publicado: (2023)
por: Gebhard, Alexander, et al.
Publicado: (2023)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
por: Zeineldeen, Mohammad, et al.
Publicado: (2023)
por: Zeineldeen, Mohammad, et al.
Publicado: (2023)
Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
por: Nie, Jingping, et al.
Publicado: (2024)
por: Nie, Jingping, et al.
Publicado: (2024)
Dementia classification from spontaneous speech using wrapper-based feature selection
por: Niemelä, Marko, et al.
Publicado: (2025)
por: Niemelä, Marko, et al.
Publicado: (2025)
Ejemplares similares
-
Acoustic identification of individual animals with hierarchical contrastive learning
por: Nolasco, Ines, et al.
Publicado: (2024) -
Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
por: Ghani, Burooj, et al.
Publicado: (2024) -
Computational bioacoustics with deep learning: a review and roadmap
por: Stowell, Dan
Publicado: (2021) -
InsectSet459: an open dataset of insect sounds for bioacoustic machine learning
por: Faiß, Marius, et al.
Publicado: (2025) -
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
por: Pepino, Leonardo, et al.
Publicado: (2023)