A Joint Multitask Model for Morpho-Syntactic Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Inostroza, Demian, Mistica, Mel, Vylomova, Ekaterina, Guest, Chris, Kurniawan, Kemal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918127485321216
author Inostroza, Demian
Mistica, Mel
Vylomova, Ekaterina
Guest, Chris
Kurniawan, Kemal
author_facet Inostroza, Demian
Mistica, Mel
Vylomova, Ekaterina
Guest, Chris
Kurniawan, Kemal
contents We present a joint multitask model for the UniDive 2025 Morpho-Syntactic Parsing shared task, where systems predict both morphological and syntactic analyses following novel UD annotation scheme. Our system uses a shared XLM-RoBERTa encoder with three specialized decoders for content word identification, dependency parsing, and morphosyntactic feature prediction. Our model achieves the best overall performance on the shared task's leaderboard covering nine typologically diverse languages, with an average MSLAS score of 78.7 percent, LAS of 80.1 percent, and Feats F1 of 90.3 percent. Our ablation studies show that matching the task's gold tokenization and content word identification are crucial to model performance. Error analysis reveals that our model struggles with core grammatical cases (particularly Nom-Acc) and nominal features across languages.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Joint Multitask Model for Morpho-Syntactic Parsing
Inostroza, Demian
Mistica, Mel
Vylomova, Ekaterina
Guest, Chris
Kurniawan, Kemal
Computation and Language
We present a joint multitask model for the UniDive 2025 Morpho-Syntactic Parsing shared task, where systems predict both morphological and syntactic analyses following novel UD annotation scheme. Our system uses a shared XLM-RoBERTa encoder with three specialized decoders for content word identification, dependency parsing, and morphosyntactic feature prediction. Our model achieves the best overall performance on the shared task's leaderboard covering nine typologically diverse languages, with an average MSLAS score of 78.7 percent, LAS of 80.1 percent, and Feats F1 of 90.3 percent. Our ablation studies show that matching the task's gold tokenization and content word identification are crucial to model performance. Error analysis reveals that our model struggles with core grammatical cases (particularly Nom-Acc) and nominal features across languages.
title A Joint Multitask Model for Morpho-Syntactic Parsing
topic Computation and Language
url https://arxiv.org/abs/2508.14307