PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Jiajun, Sawada, Naoki, Miyazaki, Koichi, Toda, Tomoki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911138501885952
author He, Jiajun
Sawada, Naoki
Miyazaki, Koichi
Toda, Tomoki
author_facet He, Jiajun
Sawada, Naoki
Miyazaki, Koichi
Toda, Tomoki
contents Automatic speech recognition (ASR) systems struggle with domain-specific named entities, especially homophones. Contextual ASR improves recognition but often fails to capture fine-grained phoneme variations due to limited entity diversity. Moreover, prior methods treat entities as independent tokens, leading to incomplete multi-token biasing. To address these issues, we propose Phoneme-Augmented Robust Contextual ASR via COntrastive entity disambiguation (PARCO), which integrates phoneme-aware encoding, contrastive entity disambiguation, entity-level supervision, and hierarchical entity filtering. These components enhance phonetic discrimination, ensure complete entity retrieval, and reduce false positives under uncertainty. Experiments show that PARCO achieves CER of 4.22% on Chinese AISHELL-1 and WER of 11.14% on English DATA2 under 1,000 distractors, significantly outperforming baselines. PARCO also demonstrates robust gains on out-of-domain datasets like THCHS-30 and LibriSpeech.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
He, Jiajun
Sawada, Naoki
Miyazaki, Koichi
Toda, Tomoki
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Automatic speech recognition (ASR) systems struggle with domain-specific named entities, especially homophones. Contextual ASR improves recognition but often fails to capture fine-grained phoneme variations due to limited entity diversity. Moreover, prior methods treat entities as independent tokens, leading to incomplete multi-token biasing. To address these issues, we propose Phoneme-Augmented Robust Contextual ASR via COntrastive entity disambiguation (PARCO), which integrates phoneme-aware encoding, contrastive entity disambiguation, entity-level supervision, and hierarchical entity filtering. These components enhance phonetic discrimination, ensure complete entity retrieval, and reduce false positives under uncertainty. Experiments show that PARCO achieves CER of 4.22% on Chinese AISHELL-1 and WER of 11.14% on English DATA2 under 1,000 distractors, significantly outperforming baselines. PARCO also demonstrates robust gains on out-of-domain datasets like THCHS-30 and LibriSpeech.
title PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2509.04357