Moonshine: Speech Recognition for Live Transcription and Voice Commands
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910659800727552 |
|---|---|
| author | Jeffries, Nat King, Evan Kudlur, Manjunath Nicholson, Guy Wang, James Warden, Pete |
| author_facet | Jeffries, Nat King, Evan Kudlur, Manjunath Nicholson, Guy Wang, James Warden, Pete |
| contents | This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing. Moonshine is based on an encoder-decoder transformer architecture and employs Rotary Position Embedding (RoPE) instead of traditional absolute position embeddings. The model is trained on speech segments of various lengths, but without using zero-padding, leading to greater efficiency for the encoder during inference time. When benchmarked against OpenAI's Whisper tiny-en, Moonshine Tiny demonstrates a 5x reduction in compute requirements for transcribing a 10-second speech segment while incurring no increase in word error rates across standard evaluation datasets. These results highlight Moonshine's potential for real-time and resource-constrained applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_15608 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Moonshine: Speech Recognition for Live Transcription and Voice Commands Jeffries, Nat King, Evan Kudlur, Manjunath Nicholson, Guy Wang, James Warden, Pete Sound Computation and Language Machine Learning Audio and Speech Processing This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing. Moonshine is based on an encoder-decoder transformer architecture and employs Rotary Position Embedding (RoPE) instead of traditional absolute position embeddings. The model is trained on speech segments of various lengths, but without using zero-padding, leading to greater efficiency for the encoder during inference time. When benchmarked against OpenAI's Whisper tiny-en, Moonshine Tiny demonstrates a 5x reduction in compute requirements for transcribing a 10-second speech segment while incurring no increase in word error rates across standard evaluation datasets. These results highlight Moonshine's potential for real-time and resource-constrained applications. |
| title | Moonshine: Speech Recognition for Live Transcription and Voice Commands |
| topic | Sound Computation and Language Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.15608 |