Saved in:
Bibliographic Details
Main Authors: Chang, Kai-Wei, Chen, Wei-Chih, Hu, En-Pei, Lee, Hung-yi, Glass, James
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.22267
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910215221280768
author Chang, Kai-Wei
Chen, Wei-Chih
Hu, En-Pei
Lee, Hung-yi
Glass, James
author_facet Chang, Kai-Wei
Chen, Wei-Chih
Hu, En-Pei
Lee, Hung-yi
Glass, James
contents We introduce TiCo, a time-controllable spoken dialogue model (SDM) that follows time-constrained instructions (e.g., "Please generate a response lasting about 15 seconds") and generates spoken responses with controllable duration. This capability is valuable for real-world spoken language systems such as voice assistants and interactive agents, where controlling response duration can improve interaction quality. However, despite their strong ability to generate natural spoken responses, existing models lack time awareness and struggle to follow duration-related instructions. To systematically evaluate this, we introduce TiCo-Bench, the first benchmark for time-controllable instruction following in SDMs, on which existing open-source and commercial models frequently fail to satisfy explicit time constraints. TiCo addresses this limitation by enabling an SDM to estimate elapsed speaking time during generation through Spoken Time Markers (STM) (e.g., <10.6 seconds>). These markers help the model maintain awareness of time and adjust the remaining content to meet the target duration. TiCo is post-trained efficiently without question-answer paired data, relying on self-generation and reinforcement learning with verifiable reward. Experimental results show that TiCo reduces duration error by 2.7x over its backbone and 1.6x over the strongest baseline, while preserving response quality.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22267
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TiCo: Time-Controllable Spoken Dialogue Model
Chang, Kai-Wei
Chen, Wei-Chih
Hu, En-Pei
Lee, Hung-yi
Glass, James
Computation and Language
Artificial Intelligence
Audio and Speech Processing
We introduce TiCo, a time-controllable spoken dialogue model (SDM) that follows time-constrained instructions (e.g., "Please generate a response lasting about 15 seconds") and generates spoken responses with controllable duration. This capability is valuable for real-world spoken language systems such as voice assistants and interactive agents, where controlling response duration can improve interaction quality. However, despite their strong ability to generate natural spoken responses, existing models lack time awareness and struggle to follow duration-related instructions. To systematically evaluate this, we introduce TiCo-Bench, the first benchmark for time-controllable instruction following in SDMs, on which existing open-source and commercial models frequently fail to satisfy explicit time constraints. TiCo addresses this limitation by enabling an SDM to estimate elapsed speaking time during generation through Spoken Time Markers (STM) (e.g., <10.6 seconds>). These markers help the model maintain awareness of time and adjust the remaining content to meet the target duration. TiCo is post-trained efficiently without question-answer paired data, relying on self-generation and reinforcement learning with verifiable reward. Experimental results show that TiCo reduces duration error by 2.7x over its backbone and 1.6x over the strongest baseline, while preserving response quality.
title TiCo: Time-Controllable Spoken Dialogue Model
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2603.22267