Saved in:
Bibliographic Details
Main Authors: Yang, Changbing, Ma, Franklin, Shi, Freda, Zhu, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.00343
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914128027713536
author Yang, Changbing
Ma, Franklin
Shi, Freda
Zhu, Jian
author_facet Yang, Changbing
Ma, Franklin
Shi, Freda
Zhu, Jian
contents This paper introduces LingGym, a new benchmark that evaluates LLMs' capacity for meta-linguistic reasoning using Interlinear Glossed Text (IGT) and grammatical descriptions extracted from 18 typologically diverse reference grammars. Unlike previous work that focuses on specific downstream tasks, we assess whether LLMs can generalize linguistic inference across low-resource languages and structures not seen during training. We present a controlled evaluation task: Word-Gloss Inference, in which the model must infer a missing word and gloss from context using varying levels of linguistic information (e.g., glosses, grammatical explanations, translations). Our results show that incorporating structured linguistic cues leads to consistent improvements in reasoning performance across all models. This work highlights both the promise and current limitations of using LLMs for typologically informed linguistic analysis and low-resource language documentation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00343
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LingGym: How Far Are LLMs from Thinking Like Field Linguists?
Yang, Changbing
Ma, Franklin
Shi, Freda
Zhu, Jian
Computation and Language
This paper introduces LingGym, a new benchmark that evaluates LLMs' capacity for meta-linguistic reasoning using Interlinear Glossed Text (IGT) and grammatical descriptions extracted from 18 typologically diverse reference grammars. Unlike previous work that focuses on specific downstream tasks, we assess whether LLMs can generalize linguistic inference across low-resource languages and structures not seen during training. We present a controlled evaluation task: Word-Gloss Inference, in which the model must infer a missing word and gloss from context using varying levels of linguistic information (e.g., glosses, grammatical explanations, translations). Our results show that incorporating structured linguistic cues leads to consistent improvements in reasoning performance across all models. This work highlights both the promise and current limitations of using LLMs for typologically informed linguistic analysis and low-resource language documentation.
title LingGym: How Far Are LLMs from Thinking Like Field Linguists?
topic Computation and Language
url https://arxiv.org/abs/2511.00343