Moonshine: Speech Recognition for Live Transcription and Voice Commands

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeffries, Nat, King, Evan, Kudlur, Manjunath, Nicholson, Guy, Wang, James, Warden, Pete
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910659800727552
author Jeffries, Nat
King, Evan
Kudlur, Manjunath
Nicholson, Guy
Wang, James
Warden, Pete
author_facet Jeffries, Nat
King, Evan
Kudlur, Manjunath
Nicholson, Guy
Wang, James
Warden, Pete
contents This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing. Moonshine is based on an encoder-decoder transformer architecture and employs Rotary Position Embedding (RoPE) instead of traditional absolute position embeddings. The model is trained on speech segments of various lengths, but without using zero-padding, leading to greater efficiency for the encoder during inference time. When benchmarked against OpenAI's Whisper tiny-en, Moonshine Tiny demonstrates a 5x reduction in compute requirements for transcribing a 10-second speech segment while incurring no increase in word error rates across standard evaluation datasets. These results highlight Moonshine's potential for real-time and resource-constrained applications.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Moonshine: Speech Recognition for Live Transcription and Voice Commands
Jeffries, Nat
King, Evan
Kudlur, Manjunath
Nicholson, Guy
Wang, James
Warden, Pete
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing. Moonshine is based on an encoder-decoder transformer architecture and employs Rotary Position Embedding (RoPE) instead of traditional absolute position embeddings. The model is trained on speech segments of various lengths, but without using zero-padding, leading to greater efficiency for the encoder during inference time. When benchmarked against OpenAI's Whisper tiny-en, Moonshine Tiny demonstrates a 5x reduction in compute requirements for transcribing a 10-second speech segment while incurring no increase in word error rates across standard evaluation datasets. These results highlight Moonshine's potential for real-time and resource-constrained applications.
title Moonshine: Speech Recognition for Live Transcription and Voice Commands
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.15608