Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Kevin, Buzaaba, Happy, Fellbaum, Christiane
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915978133110784
author Guan, Kevin
Buzaaba, Happy
Fellbaum, Christiane
author_facet Guan, Kevin
Buzaaba, Happy
Fellbaum, Christiane
contents Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers -- the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT -- across ten typologically diverse languages, with a focus on low-resource African languages. We find that the Biaffine LSTM consistently outperforms transformer models in low-resource regimes, with transformers recovering their advantage as training data increases. The crossover falls within a resource range typical of treebanks for under-resourced languages. Morphological complexity (measured via MATTR) emerges as a significant secondary predictor of transformers' relative disadvantage after controlling for corpus size. These results indicate that the Biaffine LSTM may be better suited for syntactic tool development in low-resource regimes until sufficient annotated data is available to leverage the representational capacity of pre-trained transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2605_02608
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
Guan, Kevin
Buzaaba, Happy
Fellbaum, Christiane
Computation and Language
Artificial Intelligence
Machine Learning
Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers -- the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT -- across ten typologically diverse languages, with a focus on low-resource African languages. We find that the Biaffine LSTM consistently outperforms transformer models in low-resource regimes, with transformers recovering their advantage as training data increases. The crossover falls within a resource range typical of treebanks for under-resourced languages. Morphological complexity (measured via MATTR) emerges as a significant secondary predictor of transformers' relative disadvantage after controlling for corpus size. These results indicate that the Biaffine LSTM may be better suited for syntactic tool development in low-resource regimes until sufficient annotated data is available to leverage the representational capacity of pre-trained transformers.
title Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.02608