Salvato in:
Dettagli Bibliografici
Autori principali: Lian, Da-Chen, Huang, Ri-Sheng, Chen, Pin-Er, Lim, Chunki, Lin, You-Kuan, Tseng, Guan-Yu, Yang, Zi-Cheng, Lin, Zhen-Yu, Chen, Pin-Cheng, Hsieh, Shu-Kai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2507.16809
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908463579267072
author Lian, Da-Chen
Huang, Ri-Sheng
Chen, Pin-Er
Lim, Chunki
Lin, You-Kuan
Tseng, Guan-Yu
Yang, Zi-Cheng
Lin, Zhen-Yu
Chen, Pin-Cheng
Hsieh, Shu-Kai
author_facet Lian, Da-Chen
Huang, Ri-Sheng
Chen, Pin-Er
Lim, Chunki
Lin, You-Kuan
Tseng, Guan-Yu
Yang, Zi-Cheng
Lin, Zhen-Yu
Chen, Pin-Cheng
Hsieh, Shu-Kai
contents We propose LingBench++, a linguistically-informed benchmark and reasoning framework designed to evaluate large language models (LLMs) on complex linguistic tasks inspired by the International Linguistics Olympiad (IOL). Unlike prior benchmarks that focus solely on final answer accuracy, LingBench++ provides structured reasoning traces, stepwise evaluation protocols, and rich typological metadata across over 90 low-resource and cross-cultural languages. We further develop a multi-agent architecture integrating grammatical knowledge retrieval, tool-augmented reasoning, and deliberate hypothesis testing. Through systematic comparisons of baseline and our proposed agentic models, we demonstrate that models equipped with external knowledge sources and iterative reasoning outperform single-pass approaches in both accuracy and interpretability. LingBench++ offers a comprehensive foundation for advancing linguistically grounded, culturally informed, and cognitively plausible reasoning in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16809
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
Lian, Da-Chen
Huang, Ri-Sheng
Chen, Pin-Er
Lim, Chunki
Lin, You-Kuan
Tseng, Guan-Yu
Yang, Zi-Cheng
Lin, Zhen-Yu
Chen, Pin-Cheng
Hsieh, Shu-Kai
Computation and Language
We propose LingBench++, a linguistically-informed benchmark and reasoning framework designed to evaluate large language models (LLMs) on complex linguistic tasks inspired by the International Linguistics Olympiad (IOL). Unlike prior benchmarks that focus solely on final answer accuracy, LingBench++ provides structured reasoning traces, stepwise evaluation protocols, and rich typological metadata across over 90 low-resource and cross-cultural languages. We further develop a multi-agent architecture integrating grammatical knowledge retrieval, tool-augmented reasoning, and deliberate hypothesis testing. Through systematic comparisons of baseline and our proposed agentic models, we demonstrate that models equipped with external knowledge sources and iterative reasoning outperform single-pass approaches in both accuracy and interpretability. LingBench++ offers a comprehensive foundation for advancing linguistically grounded, culturally informed, and cognitively plausible reasoning in LLMs.
title LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
topic Computation and Language
url https://arxiv.org/abs/2507.16809