TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908506006749184 |
|---|---|
| author | Kumar, Shashi Madikeri, Srikanth Villatoro-Tello, Esaú Burdisso, Sergio Rangappa, Pradeep Carofilis, Andrés Motlicek, Petr Pandia, Karthik Venkatesan, Shankar Hacioğlu, Kadri Stolcke, Andreas |
| author_facet | Kumar, Shashi Madikeri, Srikanth Villatoro-Tello, Esaú Burdisso, Sergio Rangappa, Pradeep Carofilis, Andrés Motlicek, Petr Pandia, Karthik Venkatesan, Shankar Hacioğlu, Kadri Stolcke, Andreas |
| contents | Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_19856 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation Kumar, Shashi Madikeri, Srikanth Villatoro-Tello, Esaú Burdisso, Sergio Rangappa, Pradeep Carofilis, Andrés Motlicek, Petr Pandia, Karthik Venkatesan, Shankar Hacioğlu, Kadri Stolcke, Andreas Computation and Language Audio and Speech Processing Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets and scale effectively. We propose TokenVerse++, which introduces learnable vectors in the acoustic embedding space of the XLSR-Transducer ASR model for dynamic task activation. This core mechanism enables training with utterances labeled for only a subset of tasks, a key advantage over TokenVerse. We demonstrate this by successfully integrating a dataset with partial labels, specifically for ASR and an additional task, language identification, improving overall performance. TokenVerse++ achieves results on par with or exceeding TokenVerse across multiple tasks, establishing it as a more practical multitask alternative without sacrificing ASR performance. |
| title | TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.19856 |