Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vejsiu, Luan, Zheng, Qianyu, Chen, Haoxuan, Han, Yizhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911161482477568
author Vejsiu, Luan
Zheng, Qianyu
Chen, Haoxuan
Han, Yizhou
author_facet Vejsiu, Luan
Zheng, Qianyu
Chen, Haoxuan
Han, Yizhou
contents Despite ASR technology being full-scale adopted by industry and for large portions of the population, ASR systems often have errors that require editors to post-edit text quality. While LLMs are powerful post-editing tools, baseline full rewrite models have inference inefficiencies because they often generate the same redundant text over and over again. Compact edit representations have existed but often lack the efficacy and context required for optimal accuracy. This paper introduces CEGER (Context-Enhanced Granular Edit Representation), a compact edit representation that was generated for highly accurate, efficient ASR post-editing. CEGER allows LLMs to generate a sequence of structured, fine-grained, contextually rich commands to modify the original ASR output. A separate expansion module deterministically reconstructs the corrected text based on the commands. Extensive experiments on the LibriSpeech dataset that were conducted, CEGER achieves state-of-the-art accuracy, achieving the lowest word error rate (WER) versus full rewrite and prior compact representations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14263
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing
Vejsiu, Luan
Zheng, Qianyu
Chen, Haoxuan
Han, Yizhou
Computation and Language
Sound
Audio and Speech Processing
Despite ASR technology being full-scale adopted by industry and for large portions of the population, ASR systems often have errors that require editors to post-edit text quality. While LLMs are powerful post-editing tools, baseline full rewrite models have inference inefficiencies because they often generate the same redundant text over and over again. Compact edit representations have existed but often lack the efficacy and context required for optimal accuracy. This paper introduces CEGER (Context-Enhanced Granular Edit Representation), a compact edit representation that was generated for highly accurate, efficient ASR post-editing. CEGER allows LLMs to generate a sequence of structured, fine-grained, contextually rich commands to modify the original ASR output. A separate expansion module deterministically reconstructs the corrected text based on the commands. Extensive experiments on the LibriSpeech dataset that were conducted, CEGER achieves state-of-the-art accuracy, achieving the lowest word error rate (WER) versus full rewrite and prior compact representations.
title Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.14263