TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Shashi, Madikeri, Srikanth, Villatoro-Tello, Esaú, Burdisso, Sergio, Rangappa, Pradeep, Carofilis, Andrés, Motlicek, Petr, Pandia, Karthik, Venkatesan, Shankar, Hacioğlu, Kadri, Stolcke, Andreas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908506006749184
author Kumar, Shashi
Madikeri, Srikanth
Villatoro-Tello, Esaú
Burdisso, Sergio
Rangappa, Pradeep
Carofilis, Andrés
Motlicek, Petr
Pandia, Karthik
Venkatesan, Shankar
Hacioğlu, Kadri
Stolcke, Andreas
author_facet Kumar, Shashi
Madikeri, Srikanth
Villatoro-Tello, Esaú
Burdisso, Sergio
Rangappa, Pradeep
Carofilis, Andrés
Motlicek, Petr
Pandia, Karthik
Venkatesan, Shankar
Hacioğlu, Kadri
Stolcke, Andreas
contents Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19856
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
Kumar, Shashi
Madikeri, Srikanth
Villatoro-Tello, Esaú
Burdisso, Sergio
Rangappa, Pradeep
Carofilis, Andrés
Motlicek, Petr
Pandia, Karthik
Venkatesan, Shankar
Hacioğlu, Kadri
Stolcke, Andreas
Computation and Language
Audio and Speech Processing
Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance.
title TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2508.19856