An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Whetten, Ryan, Parcollet, Titouan, Moumen, Adel, Dinarelli, Marco, Estève, Yannick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
von: Whetten, Ryan, et al.
Veröffentlicht: (2026)
Open Implementation and Study of BEST-RQ for Speech Processing
von: Whetten, Ryan, et al.
Veröffentlicht: (2024)
von: Whetten, Ryan, et al.
Veröffentlicht: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
ProGRes: Prompted Generative Rescoring on ASR n-Best
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
von: Sekkat, Chloé, et al.
Veröffentlicht: (2024)
von: Sekkat, Chloé, et al.
Veröffentlicht: (2024)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
von: Mdhaffar, Salima, et al.
Veröffentlicht: (2024)
von: Mdhaffar, Salima, et al.
Veröffentlicht: (2024)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
von: Parcollet, Titouan, et al.
Veröffentlicht: (2024)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2024)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
von: Papi, Sara, et al.
Veröffentlicht: (2023)
von: Papi, Sara, et al.
Veröffentlicht: (2023)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Simultaneous Speech-to-Speech Translation Without Aligned Data
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
von: van Dalen, Rogier C., et al.
Veröffentlicht: (2025)
von: van Dalen, Rogier C., et al.
Veröffentlicht: (2025)
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
von: Battenberg, Eric, et al.
Veröffentlicht: (2024)
von: Battenberg, Eric, et al.
Veröffentlicht: (2024)
Error Analysis in a Modular Meeting Transcription System
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
von: He, Mutian, et al.
Veröffentlicht: (2024)
von: He, Mutian, et al.
Veröffentlicht: (2024)
WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
von: Kang, Yongqi, et al.
Veröffentlicht: (2025)
von: Kang, Yongqi, et al.
Veröffentlicht: (2025)
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
von: Nachmani, Eliya, et al.
Veröffentlicht: (2023)
von: Nachmani, Eliya, et al.
Veröffentlicht: (2023)
Linear-Complexity Self-Supervised Learning for Speech Processing
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
von: Zhang, Shucong, et al.
Veröffentlicht: (2024)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
Is Attention always needed? A Case Study on Language Identification from Speech
von: Mandal, Atanu, et al.
Veröffentlicht: (2021)
von: Mandal, Atanu, et al.
Veröffentlicht: (2021)
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
von: AbdulKader, Ahmad, et al.
Veröffentlicht: (2017)
von: AbdulKader, Ahmad, et al.
Veröffentlicht: (2017)
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
von: Allouche, Itai, et al.
Veröffentlicht: (2025)
von: Allouche, Itai, et al.
Veröffentlicht: (2025)
Voice Impression Control in Zero-Shot TTS
von: Fujita, Kenichi, et al.
Veröffentlicht: (2025)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2025)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2022)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2022)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025) -
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
von: Whetten, Ryan, et al.
Veröffentlicht: (2026) -
Open Implementation and Study of BEST-RQ for Speech Processing
von: Whetten, Ryan, et al.
Veröffentlicht: (2024) -
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024) -
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)