A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language
Fuente:
arXiv
Saved in:
| Main Authors: | Bañeras-Roux, Thibault, Rouvier, Mickael, Wottawa, Jane, Dufour, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics
by: Roux, Thibault Bañeras, et al.
Published: (2026)
by: Roux, Thibault Bañeras, et al.
Published: (2026)
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
by: Bañeras-Roux, Thibault, et al.
Published: (2026)
A Benchmark of French ASR Systems Based on Error Severity
by: Tholly, Antoine, et al.
Published: (2025)
by: Tholly, Antoine, et al.
Published: (2025)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
by: Labrak, Yanis, et al.
Published: (2025)
by: Labrak, Yanis, et al.
Published: (2025)
How Important Is Tokenization in French Medical Masked Language Models?
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks
by: Labrak, Yanis, et al.
Published: (2023)
by: Labrak, Yanis, et al.
Published: (2023)
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
by: Raymondaud, Quentin, et al.
Published: (2024)
by: Raymondaud, Quentin, et al.
Published: (2024)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
by: Eeckt, Steven Vander, et al.
Published: (2021)
by: Eeckt, Steven Vander, et al.
Published: (2021)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
by: Luu, Nam, et al.
Published: (2025)
by: Luu, Nam, et al.
Published: (2025)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
by: Rufai, Amina Mardiyyah, et al.
Published: (2020)
by: Rufai, Amina Mardiyyah, et al.
Published: (2020)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
by: Agro, Maha Tufail, et al.
Published: (2025)
by: Agro, Maha Tufail, et al.
Published: (2025)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
by: He, Xinlu, et al.
Published: (2025)
by: He, Xinlu, et al.
Published: (2025)
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
by: Abdullah, Abdulhady Abas, et al.
Published: (2024)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025)
by: Storey, Edward, et al.
Published: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
by: Hono, Yukiya, et al.
Published: (2023)
by: Hono, Yukiya, et al.
Published: (2023)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
by: Shakeel, Muhammad, et al.
Published: (2024)
by: Shakeel, Muhammad, et al.
Published: (2024)
Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2
by: Bousquet, Pierre-Michel, et al.
Published: (2024)
by: Bousquet, Pierre-Michel, et al.
Published: (2024)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
by: Duret, Jarod, et al.
Published: (2024)
by: Duret, Jarod, et al.
Published: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
by: Wang, Yujin, et al.
Published: (2022)
by: Wang, Yujin, et al.
Published: (2022)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
by: Higuchi, Yosuke, et al.
Published: (2023)
by: Higuchi, Yosuke, et al.
Published: (2023)
WST: Weakly Supervised Transducer for Automatic Speech Recognition
by: Gao, Dongji, et al.
Published: (2025)
by: Gao, Dongji, et al.
Published: (2025)
End-to-end Speech Recognition with similar length speech and text
by: Fan, Peng, et al.
Published: (2025)
by: Fan, Peng, et al.
Published: (2025)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
by: Zhang, Linhao, et al.
Published: (2025)
by: Zhang, Linhao, et al.
Published: (2025)
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
by: Le, Phuong-Hang, et al.
Published: (2026)
by: Le, Phuong-Hang, et al.
Published: (2026)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
by: Kocour, Martin, et al.
Published: (2025)
by: Kocour, Martin, et al.
Published: (2025)
HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
by: Wei, Xuchen, et al.
Published: (2025)
by: Wei, Xuchen, et al.
Published: (2025)
An End-to-End Speech Summarization Using Large Language Model
by: Shang, Hengchao, et al.
Published: (2024)
by: Shang, Hengchao, et al.
Published: (2024)
Automatic Speech Recognition for the Ika Language
by: Nzenwata, Uchenna, et al.
Published: (2024)
by: Nzenwata, Uchenna, et al.
Published: (2024)
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition
by: Cui, Xiaodong, et al.
Published: (2024)
by: Cui, Xiaodong, et al.
Published: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
by: Liu, Heyang, et al.
Published: (2024)
by: Liu, Heyang, et al.
Published: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
by: Farhadipour, Aref, et al.
Published: (2023)
by: Farhadipour, Aref, et al.
Published: (2023)
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
by: Belouadi, Jonas, et al.
Published: (2022)
by: Belouadi, Jonas, et al.
Published: (2022)
A Case Study on Filtering for End-to-End Speech Translation
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
by: Alam, Md Mahfuz Ibn, et al.
Published: (2024)
Pushing the Limits of Zero-shot End-to-End Speech Translation
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
Similar Items
-
Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition
by: Bañeras-Roux, Thibault, et al.
Published: (2026) -
A Paradigm for Interpreting Metrics and Identifying Critical Errors in Automatic Speech Recognition
by: Bañeras-Roux, Thibault, et al.
Published: (2026) -
HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics
by: Roux, Thibault Bañeras, et al.
Published: (2026) -
Evaluation of Automatic Speech Recognition Using Generative Large Language Models
by: Bañeras-Roux, Thibault, et al.
Published: (2026) -
A Benchmark of French ASR Systems Based on Error Severity
by: Tholly, Antoine, et al.
Published: (2025)