Large Language Model Guided Decoding for Self-Supervised Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cohen, Eyal, Raj, Bhiksha, Keshet, Joseph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
von: Rousso, Rotem, et al.
Veröffentlicht: (2024)
von: Rousso, Rotem, et al.
Veröffentlicht: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
Keyword-Guided Adaptation of Automatic Speech Recognition
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024)
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
Domain Adaptation for Contrastive Audio-Language Models
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
von: Wan, Genshun, et al.
Veröffentlicht: (2026)
von: Wan, Genshun, et al.
Veröffentlicht: (2026)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Human Voice is Unique
von: Singh, Rita, et al.
Veröffentlicht: (2025)
von: Singh, Rita, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
von: Benita, Roi, et al.
Veröffentlicht: (2023)
von: Benita, Roi, et al.
Veröffentlicht: (2023)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
Rethinking Mamba in Speech Processing by Self-Supervised Models
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
Privacy-oriented manipulation of speaker representations
von: Teixeira, Francisco, et al.
Veröffentlicht: (2023)
von: Teixeira, Francisco, et al.
Veröffentlicht: (2023)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Drax: Speech Recognition with Discrete Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
von: Aldarmaki, Ibrahim, et al.
Veröffentlicht: (2024)
von: Aldarmaki, Ibrahim, et al.
Veröffentlicht: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
von: Yanuka, Moran, et al.
Veröffentlicht: (2025)
von: Yanuka, Moran, et al.
Veröffentlicht: (2025)
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms
von: Konan, Joseph, et al.
Veröffentlicht: (2023)
von: Konan, Joseph, et al.
Veröffentlicht: (2023)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
von: C, Shiva Kumar, et al.
Veröffentlicht: (2025)
AURA Score: A Metric For Holistic Audio Question Answering Evaluation
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
von: Dixit, Satvik, et al.
Veröffentlicht: (2025)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
von: Serrand, Coralie, et al.
Veröffentlicht: (2025)
von: Serrand, Coralie, et al.
Veröffentlicht: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
von: Girish, et al.
Veröffentlicht: (2026)
von: Girish, et al.
Veröffentlicht: (2026)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
von: Rousso, Rotem, et al.
Veröffentlicht: (2024) -
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025) -
Keyword-Guided Adaptation of Automatic Speech Recognition
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024) -
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024) -
Domain Adaptation for Contrastive Audio-Language Models
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)