M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yufeng, Raj, Desh, Lin, Ju, Moritz, Niko, Jia, Junteng, Keren, Gil, Lakomkin, Egor, Huang, Yiteng, Donley, Jacob, Mahadeokar, Jay, Kalinli, Ozlem |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faster Speech-LLaMA Inference with Multi-token Prediction
by: Raj, Desh, et al.
Published: (2024)
by: Raj, Desh, et al.
Published: (2024)
Efficient Streaming LLM for Speech Recognition
by: Jia, Junteng, et al.
Published: (2024)
by: Jia, Junteng, et al.
Published: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
by: Fathullah, Yassir, et al.
Published: (2023)
by: Fathullah, Yassir, et al.
Published: (2023)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
by: Kang, Wonjune, et al.
Published: (2024)
by: Kang, Wonjune, et al.
Published: (2024)
Can Speech LLMs Think while Listening?
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
by: Yeh, Sung-Lin, et al.
Published: (2026)
by: Yeh, Sung-Lin, et al.
Published: (2026)
Token-Weighted RNN-T for Learning from Flawed Data
by: Keren, Gil, et al.
Published: (2024)
by: Keren, Gil, et al.
Published: (2024)
Open Implementation and Study of BEST-RQ for Speech Processing
by: Whetten, Ryan, et al.
Published: (2024)
by: Whetten, Ryan, et al.
Published: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
Effective internal language model training and fusion for factorized transducer model
by: Guo, Jinxi, et al.
Published: (2024)
by: Guo, Jinxi, et al.
Published: (2024)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
by: Baumann, Ilja, et al.
Published: (2025)
by: Baumann, Ilja, et al.
Published: (2025)
Conversational Speech Naturalness Predictor
by: Xu, Anfeng, et al.
Published: (2026)
by: Xu, Anfeng, et al.
Published: (2026)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
by: Xie, Jiamin, et al.
Published: (2023)
by: Xie, Jiamin, et al.
Published: (2023)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
by: Yang, Yufeng, et al.
Published: (2025)
by: Yang, Yufeng, et al.
Published: (2025)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
by: Whetten, Ryan, et al.
Published: (2024)
by: Whetten, Ryan, et al.
Published: (2024)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
by: Deng, Keqi, et al.
Published: (2024)
by: Deng, Keqi, et al.
Published: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
by: Zhao, Jinzheng, et al.
Published: (2024)
by: Zhao, Jinzheng, et al.
Published: (2024)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
by: Raj, Desh
Published: (2024)
by: Raj, Desh
Published: (2024)
BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation
by: Bagat, Raphaël, et al.
Published: (2025)
by: Bagat, Raphaël, et al.
Published: (2025)
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
by: Lin, Ju, et al.
Published: (2024)
by: Lin, Ju, et al.
Published: (2024)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
by: Ma, Yingyi, et al.
Published: (2024)
by: Ma, Yingyi, et al.
Published: (2024)
RQ: 1960 to 1985.
by: Krieger, Tillie
Published: (1985)
by: Krieger, Tillie
Published: (1985)
MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Towards measuring fairness in speech recognition: Fair-Speech dataset
by: Veliche, Irina-Elena, et al.
Published: (2024)
by: Veliche, Irina-Elena, et al.
Published: (2024)
On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech Enhancement
by: Hsieh, Tsun-An, et al.
Published: (2024)
by: Hsieh, Tsun-An, et al.
Published: (2024)
BIOFERTILIZERS: SUSTAINABLE SOLUTIONS FOR AGRICULTURAL CROP PRODUCTIVITY
by: Aradhana Dohroo*1, et al.
Published: (2025)
by: Aradhana Dohroo*1, et al.
Published: (2025)
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
by: Han, Minglun, et al.
Published: (2024)
by: Han, Minglun, et al.
Published: (2024)
Ambisonics Encoding For Arbitrary Microphone Arrays Incorporating Residual Channels For Binaural Reproduction
by: Gayer, Yhonatan, et al.
Published: (2024)
by: Gayer, Yhonatan, et al.
Published: (2024)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
by: Jiang, Liuyuan, et al.
Published: (2025)
by: Jiang, Liuyuan, et al.
Published: (2025)
Comparative energy metrics and annual efficiency analyses of CPC‐ETC integrated single slope solar desalting unit
by: Ajay Raj Singh, et al.
Published: (2024)
by: Ajay Raj Singh, et al.
Published: (2024)
Closure to the PRISM equation derived from nonlinear response theory
by: Donley, James P.
Published: (2024)
by: Donley, James P.
Published: (2024)
Report of the Faculty Research Initiative Grant: Learning-Disabled Students and Academic Library Services.
by: Donley, Mary, et al.
Published: (1990)
by: Donley, Mary, et al.
Published: (1990)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
by: Ranjan, Desh, et al.
Published: (2024)
by: Ranjan, Desh, et al.
Published: (2024)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
by: Hong, Joanna, et al.
Published: (2025)
by: Hong, Joanna, et al.
Published: (2025)
Ara-Best-RQ: Multi Dialectal Arabic SSL
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
What Do Speech Foundation Models Not Learn About Speech?
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
Recent Developments in Research on Antisense: A Sensible Modality for Designing Next‐Generation Super Vegetables and a Paradigm Shift Towards Gene‐Editing Tools
by: Srishti, et al.
Published: (2025)
by: Srishti, et al.
Published: (2025)
MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
by: Gordon, Ofir, et al.
Published: (2025)
by: Gordon, Ofir, et al.
Published: (2025)
RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation
by: Chan, Chi-Min, et al.
Published: (2024)
by: Chan, Chi-Min, et al.
Published: (2024)
Similar Items
-
Faster Speech-LLaMA Inference with Multi-token Prediction
by: Raj, Desh, et al.
Published: (2024) -
Efficient Streaming LLM for Speech Recognition
by: Jia, Junteng, et al.
Published: (2024) -
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
by: Zhou, Wei, et al.
Published: (2024) -
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
by: Fathullah, Yassir, et al.
Published: (2023) -
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
by: Kang, Wonjune, et al.
Published: (2024)