Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Kudlur, Manjunath, King, Evan, Wang, James, Warden, Pete |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
di: King, Evan, et al.
Pubblicazione: (2025)
di: King, Evan, et al.
Pubblicazione: (2025)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
RO-N3WS: Enhancing Generalization in Low-Resource ASR with Diverse Romanian Speech Benchmarks
di: Diaconu, Alexandra, et al.
Pubblicazione: (2026)
di: Diaconu, Alexandra, et al.
Pubblicazione: (2026)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
di: Torgashov, Nikita, et al.
Pubblicazione: (2025)
di: Torgashov, Nikita, et al.
Pubblicazione: (2025)
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
di: Bai, Junwen, et al.
Pubblicazione: (2024)
di: Bai, Junwen, et al.
Pubblicazione: (2024)
Coupling Speech Encoders with Downstream Text Models
di: Chelba, Ciprian, et al.
Pubblicazione: (2024)
di: Chelba, Ciprian, et al.
Pubblicazione: (2024)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
di: Bai, Richard He, et al.
Pubblicazione: (2025)
di: Bai, Richard He, et al.
Pubblicazione: (2025)
Text-only adaptation in LLM-based ASR through text denoising
di: Carofilis, Andrés, et al.
Pubblicazione: (2026)
di: Carofilis, Andrés, et al.
Pubblicazione: (2026)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025)
di: Ngo, Huong, et al.
Pubblicazione: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
Uni-ASR: Unified LLM-Based Architecture for Non-Streaming and Streaming Automatic Speech Recognition
di: Xia, Yinfeng, et al.
Pubblicazione: (2026)
di: Xia, Yinfeng, et al.
Pubblicazione: (2026)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
di: Fan, Xulin, et al.
Pubblicazione: (2026)
di: Fan, Xulin, et al.
Pubblicazione: (2026)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
di: Vieting, Peter, et al.
Pubblicazione: (2023)
di: Vieting, Peter, et al.
Pubblicazione: (2023)
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
di: Yang, Xusheng, et al.
Pubblicazione: (2025)
di: Yang, Xusheng, et al.
Pubblicazione: (2025)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
di: Wang, Charles L., et al.
Pubblicazione: (2026)
di: Wang, Charles L., et al.
Pubblicazione: (2026)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
Unified Learnable 2D Convolutional Feature Extraction for ASR
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Semantic Codebooks as Effective Priors for Neural Speech Compression
di: Bai, Liuyang, et al.
Pubblicazione: (2025)
di: Bai, Liuyang, et al.
Pubblicazione: (2025)
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
di: Ferreira, Alef Iury Siqueira, et al.
Pubblicazione: (2025)
di: Ferreira, Alef Iury Siqueira, et al.
Pubblicazione: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Anatomy of Industrial Scale Multilingual ASR
di: Ramirez, Francis McCann, et al.
Pubblicazione: (2024)
di: Ramirez, Francis McCann, et al.
Pubblicazione: (2024)
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
di: He, Jiajun, et al.
Pubblicazione: (2025)
di: He, Jiajun, et al.
Pubblicazione: (2025)
Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
di: Rosin, Theresa Pekarek, et al.
Pubblicazione: (2025)
di: Rosin, Theresa Pekarek, et al.
Pubblicazione: (2025)
SALSA: Speedy ASR-LLM Synchronous Aggregation
di: Mittal, Ashish, et al.
Pubblicazione: (2024)
di: Mittal, Ashish, et al.
Pubblicazione: (2024)
Revisiting ASR Error Correction with Specialized Models
di: Gu, Zijin, et al.
Pubblicazione: (2024)
di: Gu, Zijin, et al.
Pubblicazione: (2024)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
Unsupervised ASR via Cross-Lingual Pseudo-Labeling
di: Likhomanenko, Tatiana, et al.
Pubblicazione: (2023)
di: Likhomanenko, Tatiana, et al.
Pubblicazione: (2023)
Federated Learning of Large ASR Models in the Real World
di: Xiao, Yonghui, et al.
Pubblicazione: (2024)
di: Xiao, Yonghui, et al.
Pubblicazione: (2024)
Energy-Based Models with Applications to Speech and Language Processing
di: Ou, Zhijian
Pubblicazione: (2024)
di: Ou, Zhijian
Pubblicazione: (2024)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
di: Mdhaffar, Salima, et al.
Pubblicazione: (2024)
di: Mdhaffar, Salima, et al.
Pubblicazione: (2024)
DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement
di: Sun, Tao, et al.
Pubblicazione: (2024)
di: Sun, Tao, et al.
Pubblicazione: (2024)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
di: Tankasala, Srinath, et al.
Pubblicazione: (2023)
di: Tankasala, Srinath, et al.
Pubblicazione: (2023)
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
di: Galimzianov, Dmitrii, et al.
Pubblicazione: (2024)
di: Galimzianov, Dmitrii, et al.
Pubblicazione: (2024)
Transformer-based Model for ASR N-Best Rescoring and Rewriting
di: Kang, Iwen E., et al.
Pubblicazione: (2024)
di: Kang, Iwen E., et al.
Pubblicazione: (2024)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
di: Lei, Zhihong, et al.
Pubblicazione: (2024)
di: Lei, Zhihong, et al.
Pubblicazione: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
di: Yang, Tzu-Ting, et al.
Pubblicazione: (2024)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
di: Lou, Yuxuan, et al.
Pubblicazione: (2026)
di: Lou, Yuxuan, et al.
Pubblicazione: (2026)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
di: Xie, Yuan, et al.
Pubblicazione: (2026)
di: Xie, Yuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
di: King, Evan, et al.
Pubblicazione: (2025) -
Moonshine: Speech Recognition for Live Transcription and Voice Commands
di: Jeffries, Nat, et al.
Pubblicazione: (2024) -
RO-N3WS: Enhancing Generalization in Low-Resource ASR with Diverse Romanian Speech Benchmarks
di: Diaconu, Alexandra, et al.
Pubblicazione: (2026) -
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
di: Torgashov, Nikita, et al.
Pubblicazione: (2025) -
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
di: Bai, Junwen, et al.
Pubblicazione: (2024)