In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914505873686528 |
|---|---|
| author | Fan, Xulin Sunder, Vishal Thomas, Samuel Hasegawa-Johnson, Mark Kingsbury, Brian Saon, George |
| author_facet | Fan, Xulin Sunder, Vishal Thomas, Samuel Hasegawa-Johnson, Mark Kingsbury, Brian Saon, George |
| contents | Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and multimodal synchronization, yet it is often handled by external alignment tools. In this work, we extend an existing speech-aware language model to predict timestamps directly alongside transcripts. We introduce a set of novel lightweight training strategies that improve alignment robustness while preserving recognition quality. Experiments across multiple datasets show that these strategies not only enhance timestamp accuracy, but also yield gains in overall ASR performance. Together, they demonstrate an efficient and unified approach to speech recognition with precise timestamp prediction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_22817 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions Fan, Xulin Sunder, Vishal Thomas, Samuel Hasegawa-Johnson, Mark Kingsbury, Brian Saon, George Audio and Speech Processing Computation and Language Machine Learning Sound Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and multimodal synchronization, yet it is often handled by external alignment tools. In this work, we extend an existing speech-aware language model to predict timestamps directly alongside transcripts. We introduce a set of novel lightweight training strategies that improve alignment robustness while preserving recognition quality. Experiments across multiple datasets show that these strategies not only enhance timestamp accuracy, but also yield gains in overall ASR performance. Together, they demonstrate an efficient and unified approach to speech recognition with precise timestamp prediction. |
| title | In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions |
| topic | Audio and Speech Processing Computation and Language Machine Learning Sound |
| url | https://arxiv.org/abs/2604.22817 |