Joint decoding method for controllable contextual speech recognition based on Speech LLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913985732804608 |
|---|---|
| author | Fang, Yangui Peng, Jing Xi, Yu Li, Xu Li, Haoyu Zhang, Chengwei Zhong, Guohui Yu, Kai |
| author_facet | Fang, Yangui Peng, Jing Xi, Yu Li, Xu Li, Haoyu Zhang, Chengwei Zhong, Guohui Yu, Kai |
| contents | Contextual speech recognition refers to the ability to identify preferences for specific content based on contextual information. Recently, leveraging the contextual understanding capabilities of Speech LLM to achieve contextual biasing by injecting contextual information through prompts have emerged as a research hotspot.However, the direct information injection method via prompts relies on the internal attention mechanism of the model, making it impossible to explicitly control the extent of information injection. To address this limitation, we propose a joint decoding method to control the contextual information. This approach enables explicit control over the injected contextual information and achieving superior recognition performance. Additionally, Our method can also be used for sensitive word suppression recognition.Furthermore, experimental results show that even Speech LLM not pre-trained on long contextual data can acquire long contextual capabilities through our method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_08585 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Joint decoding method for controllable contextual speech recognition based on Speech LLM Fang, Yangui Peng, Jing Xi, Yu Li, Xu Li, Haoyu Zhang, Chengwei Zhong, Guohui Yu, Kai Audio and Speech Processing Contextual speech recognition refers to the ability to identify preferences for specific content based on contextual information. Recently, leveraging the contextual understanding capabilities of Speech LLM to achieve contextual biasing by injecting contextual information through prompts have emerged as a research hotspot.However, the direct information injection method via prompts relies on the internal attention mechanism of the model, making it impossible to explicitly control the extent of information injection. To address this limitation, we propose a joint decoding method to control the contextual information. This approach enables explicit control over the injected contextual information and achieving superior recognition performance. Additionally, Our method can also be used for sensitive word suppression recognition.Furthermore, experimental results show that even Speech LLM not pre-trained on long contextual data can acquire long contextual capabilities through our method. |
| title | Joint decoding method for controllable contextual speech recognition based on Speech LLM |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.08585 |