Joint decoding method for controllable contextual speech recognition based on Speech LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Yangui, Peng, Jing, Xi, Yu, Li, Xu, Li, Haoyu, Zhang, Chengwei, Zhong, Guohui, Yu, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913985732804608
author Fang, Yangui
Peng, Jing
Xi, Yu
Li, Xu
Li, Haoyu
Zhang, Chengwei
Zhong, Guohui
Yu, Kai
author_facet Fang, Yangui
Peng, Jing
Xi, Yu
Li, Xu
Li, Haoyu
Zhang, Chengwei
Zhong, Guohui
Yu, Kai
contents Contextual speech recognition refers to the ability to identify preferences for specific content based on contextual information. Recently, leveraging the contextual understanding capabilities of Speech LLM to achieve contextual biasing by injecting contextual information through prompts have emerged as a research hotspot.However, the direct information injection method via prompts relies on the internal attention mechanism of the model, making it impossible to explicitly control the extent of information injection. To address this limitation, we propose a joint decoding method to control the contextual information. This approach enables explicit control over the injected contextual information and achieving superior recognition performance. Additionally, Our method can also be used for sensitive word suppression recognition.Furthermore, experimental results show that even Speech LLM not pre-trained on long contextual data can acquire long contextual capabilities through our method.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Joint decoding method for controllable contextual speech recognition based on Speech LLM
Fang, Yangui
Peng, Jing
Xi, Yu
Li, Xu
Li, Haoyu
Zhang, Chengwei
Zhong, Guohui
Yu, Kai
Audio and Speech Processing
Contextual speech recognition refers to the ability to identify preferences for specific content based on contextual information. Recently, leveraging the contextual understanding capabilities of Speech LLM to achieve contextual biasing by injecting contextual information through prompts have emerged as a research hotspot.However, the direct information injection method via prompts relies on the internal attention mechanism of the model, making it impossible to explicitly control the extent of information injection. To address this limitation, we propose a joint decoding method to control the contextual information. This approach enables explicit control over the injected contextual information and achieving superior recognition performance. Additionally, Our method can also be used for sensitive word suppression recognition.Furthermore, experimental results show that even Speech LLM not pre-trained on long contextual data can acquire long contextual capabilities through our method.
title Joint decoding method for controllable contextual speech recognition based on Speech LLM
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.08585