Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing
Fuente:
arXiv
Guardado en:
| Autores principales: | Perera, David, Letzelter, Victor, Mariotte, Théo, Cortés, Adrien, Chen, Mickael, Essid, Slim, Richard, Gaël |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multiple Choice Learning for Efficient Speech Separation with Many Speakers
por: Perera, David, et al.
Publicado: (2024)
por: Perera, David, et al.
Publicado: (2024)
Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
por: Gagnere, Antonin, et al.
Publicado: (2025)
por: Gagnere, Antonin, et al.
Publicado: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
por: Zaiem, Salah, et al.
Publicado: (2024)
por: Zaiem, Salah, et al.
Publicado: (2024)
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
por: Gagnere, Antonin, et al.
Publicado: (2024)
por: Gagnere, Antonin, et al.
Publicado: (2024)
SALT: Standardized Audio event Label Taxonomy
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)
Winner-takes-all learners are geometry-aware conditional density estimators
por: Letzelter, Victor, et al.
Publicado: (2024)
por: Letzelter, Victor, et al.
Publicado: (2024)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
por: Olvera, Michel, et al.
Publicado: (2024)
por: Olvera, Michel, et al.
Publicado: (2024)
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
por: Serre, Thomas, et al.
Publicado: (2024)
por: Serre, Thomas, et al.
Publicado: (2024)
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
por: Mariotte, Théo, et al.
Publicado: (2024)
por: Mariotte, Théo, et al.
Publicado: (2024)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
por: Serre, Thomas, et al.
Publicado: (2026)
por: Serre, Thomas, et al.
Publicado: (2026)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
por: Berger, Clémentine, et al.
Publicado: (2025)
por: Berger, Clémentine, et al.
Publicado: (2025)
Online speaker diarization of meetings guided by speech separation
por: Gruttadauria, Elio, et al.
Publicado: (2024)
por: Gruttadauria, Elio, et al.
Publicado: (2024)
ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
por: Mariotte, Theo, et al.
Publicado: (2024)
por: Mariotte, Theo, et al.
Publicado: (2024)
Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
por: Lebourdais, Martin, et al.
Publicado: (2024)
por: Lebourdais, Martin, et al.
Publicado: (2024)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
por: Berger, Clémentine, et al.
Publicado: (2025)
por: Berger, Clémentine, et al.
Publicado: (2025)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
por: Zaiem, Salah, et al.
Publicado: (2023)
por: Zaiem, Salah, et al.
Publicado: (2023)
WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
por: Baoueb, Teysir, et al.
Publicado: (2024)
por: Baoueb, Teysir, et al.
Publicado: (2024)
An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
por: Malard, Hugo, et al.
Publicado: (2024)
por: Malard, Hugo, et al.
Publicado: (2024)
Point Processes and spatial statistics in time-frequency analysis
por: Pascal, Barbara, et al.
Publicado: (2024)
por: Pascal, Barbara, et al.
Publicado: (2024)
An Explainable Proxy Model for Multiabel Audio Segmentation
por: Mariotte, Théo, et al.
Publicado: (2024)
por: Mariotte, Théo, et al.
Publicado: (2024)
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
por: Raymondaud, Quentin, et al.
Publicado: (2024)
por: Raymondaud, Quentin, et al.
Publicado: (2024)
Sparse Autoencoders Make Audio Foundation Models more Explainable
por: Mariotte, Théo, et al.
Publicado: (2025)
por: Mariotte, Théo, et al.
Publicado: (2025)
VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
por: Koriyama, Tomoki
Publicado: (2024)
por: Koriyama, Tomoki
Publicado: (2024)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
por: Kameoka, Hirokazu, et al.
Publicado: (2020)
por: Kameoka, Hirokazu, et al.
Publicado: (2020)
Singer Identity Representation Learning using Self-Supervised Techniques
por: Torres, Bernardo, et al.
Publicado: (2024)
por: Torres, Bernardo, et al.
Publicado: (2024)
Structure-informed Positional Encoding for Music Generation
por: Agarwal, Manvi, et al.
Publicado: (2024)
por: Agarwal, Manvi, et al.
Publicado: (2024)
Déréverbération non-supervisée de la parole par modèle hybride
por: Bahrman, Louis, et al.
Publicado: (2025)
por: Bahrman, Louis, et al.
Publicado: (2025)
6KSFx Synth Dataset
por: Garcia, Nelly, et al.
Publicado: (2025)
por: Garcia, Nelly, et al.
Publicado: (2025)
Research on the Acoustic Emission Source Localization Methodology in Composite Materials based on Artificial Intelligence
por: Won, Jongick, et al.
Publicado: (2024)
por: Won, Jongick, et al.
Publicado: (2024)
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
por: Ando, Angelika, et al.
Publicado: (2025)
por: Ando, Angelika, et al.
Publicado: (2025)
Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
por: Miara, Victor, et al.
Publicado: (2024)
por: Miara, Victor, et al.
Publicado: (2024)
Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
por: Giniès, Benoît, et al.
Publicado: (2025)
por: Giniès, Benoît, et al.
Publicado: (2025)
Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2
por: Bousquet, Pierre-Michel, et al.
Publicado: (2024)
por: Bousquet, Pierre-Michel, et al.
Publicado: (2024)
Spatial Reverberation and Dereverberation using an Acoustic Multiple-Input Multiple-Output System
por: Morgenstern, Hai, et al.
Publicado: (2024)
por: Morgenstern, Hai, et al.
Publicado: (2024)
Speech dereverberation constrained on room impulse response characteristics
por: Bahrman, Louis, et al.
Publicado: (2024)
por: Bahrman, Louis, et al.
Publicado: (2024)
On Speaker Attribution with SURT
por: Raj, Desh, et al.
Publicado: (2024)
por: Raj, Desh, et al.
Publicado: (2024)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
por: Stourbe, Theophile, et al.
Publicado: (2024)
por: Stourbe, Theophile, et al.
Publicado: (2024)
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
por: Duret, Jarod, et al.
Publicado: (2024)
por: Duret, Jarod, et al.
Publicado: (2024)
Multilingual Turn-taking Prediction Using Voice Activity Projection
por: Inoue, Koji, et al.
Publicado: (2024)
por: Inoue, Koji, et al.
Publicado: (2024)
Visual Cues Support Robust Turn-taking Prediction in Noise
por: Russell, Sam O'Connor, et al.
Publicado: (2025)
por: Russell, Sam O'Connor, et al.
Publicado: (2025)
Ejemplares similares
-
Multiple Choice Learning for Efficient Speech Separation with Many Speakers
por: Perera, David, et al.
Publicado: (2024) -
Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
por: Gagnere, Antonin, et al.
Publicado: (2025) -
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
por: Zaiem, Salah, et al.
Publicado: (2024) -
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
por: Gagnere, Antonin, et al.
Publicado: (2024) -
SALT: Standardized Audio event Label Taxonomy
por: Stamatiadis, Paraskevas, et al.
Publicado: (2024)