In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Xulin, Sunder, Vishal, Thomas, Samuel, Hasegawa-Johnson, Mark, Kingsbury, Brian, Saon, George
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914505873686528
author Fan, Xulin
Sunder, Vishal
Thomas, Samuel
Hasegawa-Johnson, Mark
Kingsbury, Brian
Saon, George
author_facet Fan, Xulin
Sunder, Vishal
Thomas, Samuel
Hasegawa-Johnson, Mark
Kingsbury, Brian
Saon, George
contents Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and multimodal synchronization, yet it is often handled by external alignment tools. In this work, we extend an existing speech-aware language model to predict timestamps directly alongside transcripts. We introduce a set of novel lightweight training strategies that improve alignment robustness while preserving recognition quality. Experiments across multiple datasets show that these strategies not only enhance timestamp accuracy, but also yield gains in overall ASR performance. Together, they demonstrate an efficient and unified approach to speech recognition with precise timestamp prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22817
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
Fan, Xulin
Sunder, Vishal
Thomas, Samuel
Hasegawa-Johnson, Mark
Kingsbury, Brian
Saon, George
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and multimodal synchronization, yet it is often handled by external alignment tools. In this work, we extend an existing speech-aware language model to predict timestamps directly alongside transcripts. We introduce a set of novel lightweight training strategies that improve alignment robustness while preserving recognition quality. Experiments across multiple datasets show that these strategies not only enhance timestamp accuracy, but also yield gains in overall ASR performance. Together, they demonstrate an efficient and unified approach to speech recognition with precise timestamp prediction.
title In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2604.22817