MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hyeon, Sieun, Jung, Kyudan, Won, Jaehee, Kim, Nam-Joon, Ryu, Hyun Gon, Lee, Hyuk-Jae, Do, Jaeyoung
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916684731777024
author Hyeon, Sieun
Jung, Kyudan
Won, Jaehee
Kim, Nam-Joon
Ryu, Hyun Gon
Lee, Hyuk-Jae
Do, Jaeyoung
author_facet Hyeon, Sieun
Jung, Kyudan
Won, Jaehee
Kim, Nam-Joon
Ryu, Hyun Gon
Lee, Hyuk-Jae
Do, Jaeyoung
contents In various academic and professional settings, such as mathematics lectures or research presentations, it is often necessary to convey mathematical expressions orally. However, reading mathematical expressions aloud without accompanying visuals can significantly hinder comprehension, especially for those who are hearing-impaired or rely on subtitles due to language barriers. For instance, when a presenter reads Euler's Formula, current Automatic Speech Recognition (ASR) models often produce a verbose and error-prone textual description (e.g., e to the power of i x equals cosine of x plus i $\textit{side}$ of x), instead of the concise $\LaTeX{}$ format (i.e., $ e^{ix} = \cos(x) + i\sin(x) $), which hampers clear understanding and communication. To address this issue, we introduce MathSpeech, a novel pipeline that integrates ASR models with small Language Models (sLMs) to correct errors in mathematical expressions and accurately convert spoken expressions into structured $\LaTeX{}$ representations. Evaluated on a new dataset derived from lecture recordings, MathSpeech demonstrates $\LaTeX{}$ generation capabilities comparable to leading commercial Large Language Models (LLMs), while leveraging fine-tuned small language models of only 120M parameters. Specifically, in terms of CER, BLEU, and ROUGE scores for $\LaTeX{}$ translation, MathSpeech demonstrated significantly superior capabilities compared to GPT-4o. We observed a decrease in CER from 0.390 to 0.298, and higher ROUGE/BLEU scores compared to GPT-4o.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15655
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
Hyeon, Sieun
Jung, Kyudan
Won, Jaehee
Kim, Nam-Joon
Ryu, Hyun Gon
Lee, Hyuk-Jae
Do, Jaeyoung
Computation and Language
Artificial Intelligence
In various academic and professional settings, such as mathematics lectures or research presentations, it is often necessary to convey mathematical expressions orally. However, reading mathematical expressions aloud without accompanying visuals can significantly hinder comprehension, especially for those who are hearing-impaired or rely on subtitles due to language barriers. For instance, when a presenter reads Euler's Formula, current Automatic Speech Recognition (ASR) models often produce a verbose and error-prone textual description (e.g., e to the power of i x equals cosine of x plus i $\textit{side}$ of x), instead of the concise $\LaTeX{}$ format (i.e., $ e^{ix} = \cos(x) + i\sin(x) $), which hampers clear understanding and communication. To address this issue, we introduce MathSpeech, a novel pipeline that integrates ASR models with small Language Models (sLMs) to correct errors in mathematical expressions and accurately convert spoken expressions into structured $\LaTeX{}$ representations. Evaluated on a new dataset derived from lecture recordings, MathSpeech demonstrates $\LaTeX{}$ generation capabilities comparable to leading commercial Large Language Models (LLMs), while leveraging fine-tuned small language models of only 120M parameters. Specifically, in terms of CER, BLEU, and ROUGE scores for $\LaTeX{}$ translation, MathSpeech demonstrated significantly superior capabilities compared to GPT-4o. We observed a decrease in CER from 0.390 to 0.298, and higher ROUGE/BLEU scores compared to GPT-4o.
title MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.15655