Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Chan-Jan, Chen, Yi-Chang, Liao, Feng-Ting, Ho, Pei-Chen, Wang, Yu-Hsiang, Hsu, Po-Chun, Shiu, Da-shan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910999241555968
author Hsu, Chan-Jan
Chen, Yi-Chang
Liao, Feng-Ting
Ho, Pei-Chen
Wang, Yu-Hsiang
Hsu, Po-Chun
Shiu, Da-shan
author_facet Hsu, Chan-Jan
Chen, Yi-Chang
Liao, Feng-Ting
Ho, Pei-Chen
Wang, Yu-Hsiang
Hsu, Po-Chun
Shiu, Da-shan
contents We propose "Generative Fusion Decoding" (GFD), a novel shallow fusion framework designed to integrate large language models (LLMs) into cross-modal text recognition systems for automatic speech recognition (ASR) and optical character recognition (OCR). We derive the necessary formulations to enable GFD to operate across mismatched token spaces of different models by calculating likelihood at the byte level, thereby enabling seamless fusion and synchronous progression during the decoding process. GFD is plug-and-play by design, making it readily compatible with various auto-regressive models without the need for any re-training. GFD proves effective for general ASR and OCR tasks through intermediate and frequent interactions with LLMs, surpassing cascaded methods in English and Mandarin benchmarks. In addition, GFD transfers in-context learning abilities of LLMs and allows for adaptive ASR in instruction-aware and long-context settings, yielding significant WER reductions of up to 17.7\%.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14259
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR
Hsu, Chan-Jan
Chen, Yi-Chang
Liao, Feng-Ting
Ho, Pei-Chen
Wang, Yu-Hsiang
Hsu, Po-Chun
Shiu, Da-shan
Computation and Language
Artificial Intelligence
We propose "Generative Fusion Decoding" (GFD), a novel shallow fusion framework designed to integrate large language models (LLMs) into cross-modal text recognition systems for automatic speech recognition (ASR) and optical character recognition (OCR). We derive the necessary formulations to enable GFD to operate across mismatched token spaces of different models by calculating likelihood at the byte level, thereby enabling seamless fusion and synchronous progression during the decoding process. GFD is plug-and-play by design, making it readily compatible with various auto-regressive models without the need for any re-training. GFD proves effective for general ASR and OCR tasks through intermediate and frequent interactions with LLMs, surpassing cascaded methods in English and Mandarin benchmarks. In addition, GFD transfers in-context learning abilities of LLMs and allows for adaptive ASR in instruction-aware and long-context settings, yielding significant WER reductions of up to 17.7\%.
title Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.14259