NLE: Non-autoregressive LLM-based ASR by Transcript Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dekel, Avihu, Thomas, Samuel, Fukada, Takashi, Saon, George |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026)
von: Saon, George, et al.
Veröffentlicht: (2026)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
von: Shechtman, Slava, et al.
Veröffentlicht: (2024)
von: Shechtman, Slava, et al.
Veröffentlicht: (2024)
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)
A Non-autoregressive Model for Joint STT and TTS
von: Sunder, Vishal, et al.
Veröffentlicht: (2025)
von: Sunder, Vishal, et al.
Veröffentlicht: (2025)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
von: Saon, George, et al.
Veröffentlicht: (2025)
von: Saon, George, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Align-Consistency: Improving Non-autoregressive and Semi-supervised ASR with Consistency Regularization
von: Huang, Wanting, et al.
Veröffentlicht: (2026)
von: Huang, Wanting, et al.
Veröffentlicht: (2026)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
von: Gupta, Ankit, et al.
Veröffentlicht: (2024)
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Spoken question answering for visual queries
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
von: Yen, Hao, et al.
Veröffentlicht: (2026)
von: Yen, Hao, et al.
Veröffentlicht: (2026)
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning
von: Kong, YuXiang, et al.
Veröffentlicht: (2025)
von: Kong, YuXiang, et al.
Veröffentlicht: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
An investigation of modularity for noise robustness in conformer-based ASR
von: de Gibson, Louise Coppieters, et al.
Veröffentlicht: (2024)
von: de Gibson, Louise Coppieters, et al.
Veröffentlicht: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
von: Joshi, Sakshi, et al.
Veröffentlicht: (2025)
von: Joshi, Sakshi, et al.
Veröffentlicht: (2025)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
von: Li, Li, et al.
Veröffentlicht: (2026)
von: Li, Li, et al.
Veröffentlicht: (2026)
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
von: Gómez-Zaragozá, Lucía, et al.
Veröffentlicht: (2023)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
von: Xi, Yu, et al.
Veröffentlicht: (2025)
von: Xi, Yu, et al.
Veröffentlicht: (2025)
Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
Progressive unsupervised domain adaptation for ASR using ensemble models and multi-stage training
von: Ahmad, Rehan, et al.
Veröffentlicht: (2024)
von: Ahmad, Rehan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026) -
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026) -
Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
von: Shechtman, Slava, et al.
Veröffentlicht: (2024) -
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026) -
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
von: Fujita, Yuya, et al.
Veröffentlicht: (2024)