LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Parcollet, Titouan, Nguyen, Ha, Evain, Solene, Boito, Marcely Zanon, Pupier, Adrien, Mdhaffar, Salima, Le, Hang, Alisamir, Sina, Tomashenko, Natalia, Dinarelli, Marco, Zhang, Shucong, Allauzen, Alexandre, Coavoux, Maximin, Esteve, Yannick, Rouvier, Mickael, Goulian, Jerome, Lecouteux, Benjamin, Portet, Francois, Rossato, Solange, Ringeval, Fabien, Schwab, Didier, Besacier, Laurent |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What has LeBenchmark Learnt about French Syntax?
by: Dugonjić, Zdravko, et al.
Published: (2024)
by: Dugonjić, Zdravko, et al.
Published: (2024)
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech
by: Pupier, Adrien, et al.
Published: (2024)
by: Pupier, Adrien, et al.
Published: (2024)
NAVER LABS Europe Submission to the Instruction-following Track
by: Lee, Beomseok, et al.
Published: (2025)
by: Lee, Beomseok, et al.
Published: (2025)
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024)
by: Boito, Marcely Zanon, et al.
Published: (2024)
Open Implementation and Study of BEST-RQ for Speech Processing
by: Whetten, Ryan, et al.
Published: (2024)
by: Whetten, Ryan, et al.
Published: (2024)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
by: Whetten, Ryan, et al.
Published: (2026)
by: Whetten, Ryan, et al.
Published: (2026)
SpeechMapper: Speech-to-text Embedding Projector for LLMs
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
by: Le, Phuong-Hang, et al.
Published: (2026)
by: Le, Phuong-Hang, et al.
Published: (2026)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
by: Whetten, Ryan, et al.
Published: (2024)
by: Whetten, Ryan, et al.
Published: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
by: Whetten, Ryan, et al.
Published: (2025)
by: Whetten, Ryan, et al.
Published: (2025)
ding-01 :ARG0: An AMR Corpus for Spontaneous French Dialogue
by: Kang, Jeongwoo, et al.
Published: (2025)
by: Kang, Jeongwoo, et al.
Published: (2025)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
by: Parcollet, Titouan, et al.
Published: (2025)
by: Parcollet, Titouan, et al.
Published: (2025)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
by: Zhang, Shucong, et al.
Published: (2025)
by: Zhang, Shucong, et al.
Published: (2025)
Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
by: Parcollet, Titouan, et al.
Published: (2024)
by: Parcollet, Titouan, et al.
Published: (2024)
Linear-Complexity Self-Supervised Learning for Speech Processing
by: Zhang, Shucong, et al.
Published: (2024)
by: Zhang, Shucong, et al.
Published: (2024)
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
by: Parcollet, Titouan, et al.
Published: (2023)
by: Parcollet, Titouan, et al.
Published: (2023)
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
by: Herron, Felix, et al.
Published: (2026)
by: Herron, Felix, et al.
Published: (2026)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
by: van Dalen, Rogier C., et al.
Published: (2025)
by: van Dalen, Rogier C., et al.
Published: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
by: Parcollet, Titouan, et al.
Published: (2026)
by: Parcollet, Titouan, et al.
Published: (2026)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
by: Ferraz, Thomas Palmeira, et al.
Published: (2023)
by: Ferraz, Thomas Palmeira, et al.
Published: (2023)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
by: Tseng, Yuan, et al.
Published: (2025)
by: Tseng, Yuan, et al.
Published: (2025)
Reassessing Graph Linearization for Sequence-to-sequence AMR Parsing: On the Advantages and Limitations of Triple-Based Encoding
by: Kang, Jeongwoo, et al.
Published: (2025)
by: Kang, Jeongwoo, et al.
Published: (2025)
Should Cross-Lingual AMR Parsing go Meta? An Empirical Assessment of Meta-Learning and Joint Learning AMR Parsing
by: Kang, Jeongwoo, et al.
Published: (2024)
by: Kang, Jeongwoo, et al.
Published: (2024)
Responsible Benchmarking of Fairness for Automatic Speech Recognition
by: Herron, Felix, et al.
Published: (2026)
by: Herron, Felix, et al.
Published: (2026)
Where Do Self-Supervised Speech Models Become Unfair?
by: Herron, Felix, et al.
Published: (2026)
by: Herron, Felix, et al.
Published: (2026)
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
by: Zhou, Yongxin, et al.
Published: (2023)
by: Zhou, Yongxin, et al.
Published: (2023)
Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals
by: Zhou, Yongxin, et al.
Published: (2023)
by: Zhou, Yongxin, et al.
Published: (2023)
Automated Clinical Report Generation for Remote Cognitive Remediation: Comparing Knowledge-Engineered Templates and LLMs in Low-Resource Settings
by: Zhou, Yongxin, et al.
Published: (2026)
by: Zhou, Yongxin, et al.
Published: (2026)
Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder
by: Bouziane, Maryem, et al.
Published: (2026)
by: Bouziane, Maryem, et al.
Published: (2026)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
by: Zaiem, Salah, et al.
Published: (2024)
by: Zaiem, Salah, et al.
Published: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
by: Duret, Jarod, et al.
Published: (2024)
by: Duret, Jarod, et al.
Published: (2024)
TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
by: Bougares, Fethi, et al.
Published: (2025)
by: Bougares, Fethi, et al.
Published: (2025)
Using Multimodal and Language-Agnostic Sentence Embeddings for Abstractive Summarization
by: Chellaf, Chaimae, et al.
Published: (2026)
by: Chellaf, Chaimae, et al.
Published: (2026)
SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
by: Mdhaffar, Salima, et al.
Published: (2024)
by: Mdhaffar, Salima, et al.
Published: (2024)
ADI-20: Arabic Dialect Identification dataset and models
by: Elleuch, Haroun, et al.
Published: (2025)
by: Elleuch, Haroun, et al.
Published: (2025)
Marco Estratégico sobre Bosques Mediterráneos y Declaración de Tlemcen
by: Besacier, C
Published: (2011)
by: Besacier, C
Published: (2011)
Explorando las oportunidades de REDD+ en el Mediterráneo - un proyecto regional financiado por el Fondo Francés para el Medio Ambiente Mundial (FFEM)
by: Besacier, C
Published: (2014)
by: Besacier, C
Published: (2014)
ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks
by: Elleuch, Haroun, et al.
Published: (2025)
by: Elleuch, Haroun, et al.
Published: (2025)
SENSE models: an open source solution for multilingual and multimodal semantic-based tasks
by: Mdhaffar, Salima, et al.
Published: (2025)
by: Mdhaffar, Salima, et al.
Published: (2025)
Similar Items
-
What has LeBenchmark Learnt about French Syntax?
by: Dugonjić, Zdravko, et al.
Published: (2024) -
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech
by: Pupier, Adrien, et al.
Published: (2024) -
NAVER LABS Europe Submission to the Instruction-following Track
by: Lee, Beomseok, et al.
Published: (2025) -
mHuBERT-147: A Compact Multilingual HuBERT Model
by: Boito, Marcely Zanon, et al.
Published: (2024) -
Open Implementation and Study of BEST-RQ for Speech Processing
by: Whetten, Ryan, et al.
Published: (2024)