Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
Fuente:
arXiv
Salvato in:
| Autori principali: | van Dalen, Rogier C., Zhang, Shucong, Parcollet, Titouan, Bhattacharya, Sourav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
di: Zhang, Shucong, et al.
Pubblicazione: (2025)
di: Zhang, Shucong, et al.
Pubblicazione: (2025)
Linear-Complexity Self-Supervised Learning for Speech Processing
di: Zhang, Shucong, et al.
Pubblicazione: (2024)
di: Zhang, Shucong, et al.
Pubblicazione: (2024)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
di: Parcollet, Titouan, et al.
Pubblicazione: (2025)
di: Parcollet, Titouan, et al.
Pubblicazione: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
di: Parcollet, Titouan, et al.
Pubblicazione: (2024)
di: Parcollet, Titouan, et al.
Pubblicazione: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
di: Duret, Jarod, et al.
Pubblicazione: (2024)
di: Duret, Jarod, et al.
Pubblicazione: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
di: Whetten, Ryan, et al.
Pubblicazione: (2025)
di: Whetten, Ryan, et al.
Pubblicazione: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
di: Zaiem, Salah, et al.
Pubblicazione: (2024)
di: Zaiem, Salah, et al.
Pubblicazione: (2024)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
di: Whetten, Ryan, et al.
Pubblicazione: (2024)
di: Whetten, Ryan, et al.
Pubblicazione: (2024)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
di: Zaiem, Salah, et al.
Pubblicazione: (2023)
di: Zaiem, Salah, et al.
Pubblicazione: (2023)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
di: Falai, Alessio, et al.
Pubblicazione: (2025)
di: Falai, Alessio, et al.
Pubblicazione: (2025)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
di: Whetten, Ryan, et al.
Pubblicazione: (2026)
di: Whetten, Ryan, et al.
Pubblicazione: (2026)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
di: van Rensburg, Kyle Janse, et al.
Pubblicazione: (2026)
di: van Rensburg, Kyle Janse, et al.
Pubblicazione: (2026)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
di: Carbonneau, Marc-André, et al.
Pubblicazione: (2025)
di: Carbonneau, Marc-André, et al.
Pubblicazione: (2025)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
di: Noroozi, Vahid, et al.
Pubblicazione: (2024)
di: Noroozi, Vahid, et al.
Pubblicazione: (2024)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
di: Mi, Jinyi, et al.
Pubblicazione: (2024)
di: Mi, Jinyi, et al.
Pubblicazione: (2024)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
di: Wang, Kyra, et al.
Pubblicazione: (2024)
di: Wang, Kyra, et al.
Pubblicazione: (2024)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
di: Lu, Cheng, et al.
Pubblicazione: (2024)
di: Lu, Cheng, et al.
Pubblicazione: (2024)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
di: Lin, Weiwei, et al.
Pubblicazione: (2024)
di: Lin, Weiwei, et al.
Pubblicazione: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
di: Sakuma, Asahi, et al.
Pubblicazione: (2025)
di: Sakuma, Asahi, et al.
Pubblicazione: (2025)
Measuring Entrainment in Spontaneous Code-switched Speech
di: Bhattacharya, Debasmita, et al.
Pubblicazione: (2023)
di: Bhattacharya, Debasmita, et al.
Pubblicazione: (2023)
Investigation of Speaker Representation for Target-Speaker Speech Processing
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
Robustness of Speech Separation Models for Similar-pitch Speakers
di: Lay, Bunlong, et al.
Pubblicazione: (2024)
di: Lay, Bunlong, et al.
Pubblicazione: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
di: Roll, Nathan, et al.
Pubblicazione: (2025)
di: Roll, Nathan, et al.
Pubblicazione: (2025)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
di: Park, Taejin, et al.
Pubblicazione: (2024)
di: Park, Taejin, et al.
Pubblicazione: (2024)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
di: Sun, Zhaokai, et al.
Pubblicazione: (2025)
di: Sun, Zhaokai, et al.
Pubblicazione: (2025)
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
di: Hentschel, Michael, et al.
Pubblicazione: (2024)
di: Hentschel, Michael, et al.
Pubblicazione: (2024)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
di: Sun, Guangzhi, et al.
Pubblicazione: (2022)
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
di: Rutowski, Tomasz, et al.
Pubblicazione: (2024)
di: Rutowski, Tomasz, et al.
Pubblicazione: (2024)
REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation
di: Hirschkind, Nameer, et al.
Pubblicazione: (2025)
di: Hirschkind, Nameer, et al.
Pubblicazione: (2025)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
di: Chang, Heng-Jui, et al.
Pubblicazione: (2024)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2024)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
di: Shah, Muhammad A., et al.
Pubblicazione: (2024)
di: Shah, Muhammad A., et al.
Pubblicazione: (2024)
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
di: Huo, Mingyue, et al.
Pubblicazione: (2026)
di: Huo, Mingyue, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
di: Parcollet, Titouan, et al.
Pubblicazione: (2023) -
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
di: Zhang, Shucong, et al.
Pubblicazione: (2025) -
Linear-Complexity Self-Supervised Learning for Speech Processing
di: Zhang, Shucong, et al.
Pubblicazione: (2024) -
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
di: Tseng, Yuan, et al.
Pubblicazione: (2025) -
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
di: Parcollet, Titouan, et al.
Pubblicazione: (2025)