Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Hao, Stahlberg, Felix, Kumar, Shankar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929685020016640
author Zhang, Hao
Stahlberg, Felix
Kumar, Shankar
author_facet Zhang, Hao
Stahlberg, Felix
Kumar, Shankar
contents Large Language Models (LLMs) excel at rewriting tasks such as text style transfer and grammatical error correction. While there is considerable overlap between the inputs and outputs in these tasks, the decoding cost still increases with output length, regardless of the amount of overlap. By leveraging the overlap between the input and the output, Kaneko and Okazaki (2023) proposed model-agnostic edit span representations to compress the rewrites to save computation. They reported an output length reduction rate of nearly 80% with minimal accuracy impact in four rewriting tasks. In this paper, we propose alternative edit phrase representations inspired by phrase-based statistical machine translation. We systematically compare our phrasal representations with their span representations. We apply the LLM rewriting model to the task of Automatic Speech Recognition (ASR) post editing and show that our target-phrase-only edit representation has the best efficiency-accuracy trade-off. On the LibriSpeech test set, our method closes 50-60% of the WER gap between the edit span model and the full rewrite model while losing only 10-20% of the length reduction rate of the edit span model.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing
Zhang, Hao
Stahlberg, Felix
Kumar, Shankar
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) excel at rewriting tasks such as text style transfer and grammatical error correction. While there is considerable overlap between the inputs and outputs in these tasks, the decoding cost still increases with output length, regardless of the amount of overlap. By leveraging the overlap between the input and the output, Kaneko and Okazaki (2023) proposed model-agnostic edit span representations to compress the rewrites to save computation. They reported an output length reduction rate of nearly 80% with minimal accuracy impact in four rewriting tasks. In this paper, we propose alternative edit phrase representations inspired by phrase-based statistical machine translation. We systematically compare our phrasal representations with their span representations. We apply the LLM rewriting model to the task of Automatic Speech Recognition (ASR) post editing and show that our target-phrase-only edit representation has the best efficiency-accuracy trade-off. On the LibriSpeech test set, our method closes 50-60% of the WER gap between the edit span model and the full rewrite model while losing only 10-20% of the length reduction rate of the edit span model.
title Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.13831