A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bañeras-Roux, Thibault, Rouvier, Mickael, Wottawa, Jane, Dufour, Richard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913090754314240
author Bañeras-Roux, Thibault
Rouvier, Mickael
Wottawa, Jane
Dufour, Richard
author_facet Bañeras-Roux, Thibault
Rouvier, Mickael
Wottawa, Jane
Dufour, Richard
contents The performance of end-to-end automatic speech recognition (ASR) systems enables their increasing integration into numerous applications. While there are various benefits to such speech-to-text systems, the choice of hyperparameters and models plays a crucial role in their performance. Typically, these choices are determined by considering only the character (CER) and/or word error rate (WER) metrics. However, it has been shown in several studies that these metrics are largely incomplete and fail to adequately describe the downstream application of automatic transcripts. In this paper, we conduct a qualitative study on the French language that investigates the impact of subword tokenization algorithms and self-supervised learning models from different linguistic and acoustic perspectives, using a comprehensive set of evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03696
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language
Bañeras-Roux, Thibault
Rouvier, Mickael
Wottawa, Jane
Dufour, Richard
Computation and Language
The performance of end-to-end automatic speech recognition (ASR) systems enables their increasing integration into numerous applications. While there are various benefits to such speech-to-text systems, the choice of hyperparameters and models plays a crucial role in their performance. Typically, these choices are determined by considering only the character (CER) and/or word error rate (WER) metrics. However, it has been shown in several studies that these metrics are largely incomplete and fail to adequately describe the downstream application of automatic transcripts. In this paper, we conduct a qualitative study on the French language that investigates the impact of subword tokenization algorithms and self-supervised learning models from different linguistic and acoustic perspectives, using a comprehensive set of evaluation metrics.
title A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language
topic Computation and Language
url https://arxiv.org/abs/2605.03696