Semiparametric Token-Sequence Co-Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Hyunji, Kim, Doyoung, Jun, Jihoon, Joo, Sejune, Jang, Joel, On, Kyoung-Woon, Seo, Minjoon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913264508600320
author Lee, Hyunji
Kim, Doyoung
Jun, Jihoon
Joo, Sejune
Jang, Joel
On, Kyoung-Woon
Seo, Minjoon
author_facet Lee, Hyunji
Kim, Doyoung
Jun, Jihoon
Joo, Sejune
Jang, Joel
On, Kyoung-Woon
Seo, Minjoon
contents In this work, we introduce a semiparametric token-sequence co-supervision training method. It trains a language model by simultaneously leveraging supervision from the traditional next token prediction loss which is calculated over the parametric token embedding space and the next sequence prediction loss which is calculated over the nonparametric sequence embedding space. The nonparametric sequence embedding space is constructed by a separate language model tasked to condense an input text into a single representative embedding. Our experiments demonstrate that a model trained via both supervisions consistently surpasses models trained via each supervision independently. Analysis suggests that this co-supervision encourages a broader generalization capability across the model. Especially, the robustness of parametric token space which is established during the pretraining step tends to effectively enhance the stability of nonparametric sequence embedding space, a new space established by another language model.
format Preprint
id arxiv_https___arxiv_org_abs_2403_09024
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semiparametric Token-Sequence Co-Supervision
Lee, Hyunji
Kim, Doyoung
Jun, Jihoon
Joo, Sejune
Jang, Joel
On, Kyoung-Woon
Seo, Minjoon
Computation and Language
Artificial Intelligence
In this work, we introduce a semiparametric token-sequence co-supervision training method. It trains a language model by simultaneously leveraging supervision from the traditional next token prediction loss which is calculated over the parametric token embedding space and the next sequence prediction loss which is calculated over the nonparametric sequence embedding space. The nonparametric sequence embedding space is constructed by a separate language model tasked to condense an input text into a single representative embedding. Our experiments demonstrate that a model trained via both supervisions consistently surpasses models trained via each supervision independently. Analysis suggests that this co-supervision encourages a broader generalization capability across the model. Especially, the robustness of parametric token space which is established during the pretraining step tends to effectively enhance the stability of nonparametric sequence embedding space, a new space established by another language model.
title Semiparametric Token-Sequence Co-Supervision
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.09024