NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Lin, Yen-Ting, Chen, Zhehuai, Zelasko, Piotr, Wan, Zhen, Yang, Xuesong, Chen, Zih-Ching, Puvvada, Krishna C, Fu, Szu-Wei, Hu, Ke, Chiu, Jun Wei, Balam, Jagadeesh, Ginsburg, Boris, Wang, Yu-Chiang Frank, Yang, Chao-Han Huck |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
por: Chen, Zhehuai, et al.
Publicado: (2024)
por: Chen, Zhehuai, et al.
Publicado: (2024)
Chain-of-Thought Prompting for Speech Translation
por: Hu, Ke, et al.
Publicado: (2024)
por: Hu, Ke, et al.
Publicado: (2024)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
por: Lu, Ke-Han, et al.
Publicado: (2024)
por: Lu, Ke-Han, et al.
Publicado: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
por: Puvvada, Krishna C., et al.
Publicado: (2024)
por: Puvvada, Krishna C., et al.
Publicado: (2024)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
por: Burchi, Maxime, et al.
Publicado: (2024)
por: Burchi, Maxime, et al.
Publicado: (2024)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
por: Hu, Ke, et al.
Publicado: (2025)
por: Hu, Ke, et al.
Publicado: (2025)
EMMeTT: Efficient Multimodal Machine Translation Training
por: Żelasko, Piotr, et al.
Publicado: (2024)
por: Żelasko, Piotr, et al.
Publicado: (2024)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
por: Hu, Ke, et al.
Publicado: (2025)
por: Hu, Ke, et al.
Publicado: (2025)
Training and Inference Efficiency of Encoder-Decoder Speech Models
por: Żelasko, Piotr, et al.
Publicado: (2025)
por: Żelasko, Piotr, et al.
Publicado: (2025)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
por: Noroozi, Vahid, et al.
Publicado: (2024)
por: Noroozi, Vahid, et al.
Publicado: (2024)
Anticipating Future with Large Language Model for Simultaneous Machine Translation
por: Ouyang, Siqi, et al.
Publicado: (2024)
por: Ouyang, Siqi, et al.
Publicado: (2024)
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
por: Sekoyan, Monica, et al.
Publicado: (2025)
por: Sekoyan, Monica, et al.
Publicado: (2025)
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend
por: Jukić, Ante, et al.
Publicado: (2024)
por: Jukić, Ante, et al.
Publicado: (2024)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
por: Yang, Chao-Han Huck, et al.
Publicado: (2024)
por: Yang, Chao-Han Huck, et al.
Publicado: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
por: Wang, Weiqing, et al.
Publicado: (2024)
por: Wang, Weiqing, et al.
Publicado: (2024)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
por: Huang, Sung-Feng, et al.
Publicado: (2025)
por: Huang, Sung-Feng, et al.
Publicado: (2025)
Schrödinger Bridge for Generative Speech Enhancement
por: Jukić, Ante, et al.
Publicado: (2024)
por: Jukić, Ante, et al.
Publicado: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
por: Park, Taejin, et al.
Publicado: (2024)
por: Park, Taejin, et al.
Publicado: (2024)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
por: Huang, He, et al.
Publicado: (2024)
por: Huang, He, et al.
Publicado: (2024)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
por: Lu, Ke-Han, et al.
Publicado: (2024)
por: Lu, Ke-Han, et al.
Publicado: (2024)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
por: Noroozi, Vahid, et al.
Publicado: (2023)
por: Noroozi, Vahid, et al.
Publicado: (2023)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
por: Yang, Chih-Kai, et al.
Publicado: (2025)
por: Yang, Chih-Kai, et al.
Publicado: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
por: Dhawan, Kunal, et al.
Publicado: (2024)
por: Dhawan, Kunal, et al.
Publicado: (2024)
Expanding the Utilization of Pharmacological Treatments for Alcohol Use Disorder: Reflections on a Swedish Nationwide Study
por: Szu‐Chieh Chiu, et al.
Publicado: (2025)
por: Szu‐Chieh Chiu, et al.
Publicado: (2025)
GetBatch: Distributed Multi-Object Retrieval for ML Data Loading
por: Aizman, Alex, et al.
Publicado: (2026)
por: Aizman, Alex, et al.
Publicado: (2026)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
por: Koluguri, Nithin Rao, et al.
Publicado: (2024)
por: Koluguri, Nithin Rao, et al.
Publicado: (2024)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
por: Wang, Yun, et al.
Publicado: (2025)
por: Wang, Yun, et al.
Publicado: (2025)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
por: Zhao, Yushu, et al.
Publicado: (2025)
por: Zhao, Yushu, et al.
Publicado: (2025)
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
por: Medennikov, Ivan, et al.
Publicado: (2025)
por: Medennikov, Ivan, et al.
Publicado: (2025)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
por: Chu, Kexin, et al.
Publicado: (2025)
por: Chu, Kexin, et al.
Publicado: (2025)
Dynamic Latent Separation for Deep Learning
por: Tuan, Yi-Lin, et al.
Publicado: (2022)
por: Tuan, Yi-Lin, et al.
Publicado: (2022)
Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
por: Feng, Bo-Han, et al.
Publicado: (2025)
por: Feng, Bo-Han, et al.
Publicado: (2025)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
por: Chen, Shengzhuang, et al.
Publicado: (2025)
por: Chen, Shengzhuang, et al.
Publicado: (2025)
The First Drop of Ink: Nonlinear Impact of Misleading Information in Long-Context Reasoning
por: Gao, Muhan, et al.
Publicado: (2026)
por: Gao, Muhan, et al.
Publicado: (2026)
Modeling Missing at Random Neuropsychological Test Scores Using a Mixture of Binomial Product Experts
por: Suen, Daniel, et al.
Publicado: (2023)
por: Suen, Daniel, et al.
Publicado: (2023)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
por: He, Yifei, et al.
Publicado: (2025)
por: He, Yifei, et al.
Publicado: (2025)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
por: Yoon, Youngsik, et al.
Publicado: (2026)
por: Yoon, Youngsik, et al.
Publicado: (2026)
Ejemplares similares
-
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
por: Chen, Zhehuai, et al.
Publicado: (2024) -
Chain-of-Thought Prompting for Speech Translation
por: Hu, Ke, et al.
Publicado: (2024) -
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
por: Peng, Yifan, et al.
Publicado: (2024) -
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
por: Lu, Ke-Han, et al.
Publicado: (2024) -
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
por: Puvvada, Krishna C., et al.
Publicado: (2024)