Massively Multilingual Joint Segmentation and Glossing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ginn, Michael, Tjuatja, Lindia, Rice, Enora, Marashian, Ali, Valentini, Maria, Xu, Jasmine, Neubig, Graham, Palmer, Alexis
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914274462400512
author Ginn, Michael
Tjuatja, Lindia
Rice, Enora
Marashian, Ali
Valentini, Maria
Xu, Jasmine
Neubig, Graham
Palmer, Alexis
author_facet Ginn, Michael
Tjuatja, Lindia
Rice, Enora
Marashian, Ali
Valentini, Maria
Xu, Jasmine
Neubig, Graham
Palmer, Alexis
contents Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossLM achieve high scores on glossing benchmarks, user studies with linguists have found critical barriers to the usefulness of such models in real-world scenarios. In particular, existing models typically generate morpheme-level glosses but assign them to whole words without predicting the actual morpheme boundaries, making the predictions less interpretable and thus untrustworthy to human annotators. We conduct the first study on neural models that jointly predict interlinear glosses and the corresponding morphological segmentation from raw text. We run experiments to determine the optimal way to train models that balance segmentation and glossing accuracy, as well as the alignment between the two tasks. We extend the training corpus of GlossLM and pretrain PolyGloss, a family of seq2seq multilingual models for joint segmentation and glossing that outperforms GlossLM on glossing and beats various open-source LLMs on segmentation, glossing, and alignment. In addition, we demonstrate that PolyGloss can be quickly adapted to a new dataset via low-rank adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10925
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Massively Multilingual Joint Segmentation and Glossing
Ginn, Michael
Tjuatja, Lindia
Rice, Enora
Marashian, Ali
Valentini, Maria
Xu, Jasmine
Neubig, Graham
Palmer, Alexis
Computation and Language
Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossLM achieve high scores on glossing benchmarks, user studies with linguists have found critical barriers to the usefulness of such models in real-world scenarios. In particular, existing models typically generate morpheme-level glosses but assign them to whole words without predicting the actual morpheme boundaries, making the predictions less interpretable and thus untrustworthy to human annotators. We conduct the first study on neural models that jointly predict interlinear glosses and the corresponding morphological segmentation from raw text. We run experiments to determine the optimal way to train models that balance segmentation and glossing accuracy, as well as the alignment between the two tasks. We extend the training corpus of GlossLM and pretrain PolyGloss, a family of seq2seq multilingual models for joint segmentation and glossing that outperforms GlossLM on glossing and beats various open-source LLMs on segmentation, glossing, and alignment. In addition, we demonstrate that PolyGloss can be quickly adapted to a new dataset via low-rank adaptation.
title Massively Multilingual Joint Segmentation and Glossing
topic Computation and Language
url https://arxiv.org/abs/2601.10925