IOLBENCH: Benchmarking LLMs on Linguistic Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Goyal, Satyam, Dan, Soham |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
by: Dai, Song, et al.
Published: (2025)
by: Dai, Song, et al.
Published: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
by: Parcalabescu, Letitia, et al.
Published: (2021)
by: Parcalabescu, Letitia, et al.
Published: (2021)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents
by: Jehu-Appiah, Rodney
Published: (2026)
by: Jehu-Appiah, Rodney
Published: (2026)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
by: Bayram, M. Ali, et al.
Published: (2025)
by: Bayram, M. Ali, et al.
Published: (2025)
Bridging the Reasoning Gap: Small LLMs Can Plan with Generalised Strategies
by: Borro, Andrey, et al.
Published: (2025)
by: Borro, Andrey, et al.
Published: (2025)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
by: Teixeira, Tiago, et al.
Published: (2026)
by: Teixeira, Tiago, et al.
Published: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
by: Son, Daniel, et al.
Published: (2025)
by: Son, Daniel, et al.
Published: (2025)
On the Compatibility of Generative AI and Generative Linguistics
by: Portelance, Eva, et al.
Published: (2024)
by: Portelance, Eva, et al.
Published: (2024)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
Can LLMs Compute with Reasons?
by: Sandilya, Harshit, et al.
Published: (2024)
by: Sandilya, Harshit, et al.
Published: (2024)
LLMs and the Human Condition
by: Wallis, Peter
Published: (2024)
by: Wallis, Peter
Published: (2024)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
by: Marcuzzo, Matteo, et al.
Published: (2025)
by: Marcuzzo, Matteo, et al.
Published: (2025)
LLMs for Game Theory: Entropy-Guided In-Context Learning and Adaptive CoT Reasoning
by: Banfi, Tommaso Felice, et al.
Published: (2026)
by: Banfi, Tommaso Felice, et al.
Published: (2026)
Natural Language as Policies: Reasoning for Coordinate-Level Embodied Control with LLMs
by: Mikami, Yusuke, et al.
Published: (2024)
by: Mikami, Yusuke, et al.
Published: (2024)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
by: Chourasia, Shivam, et al.
Published: (2026)
by: Chourasia, Shivam, et al.
Published: (2026)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Ontology Learning with LLMs: A Benchmark Study on Axiom Identification
by: Bakker, Roos M., et al.
Published: (2025)
by: Bakker, Roos M., et al.
Published: (2025)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
by: Cui, Jian, et al.
Published: (2026)
by: Cui, Jian, et al.
Published: (2026)
Linguistic Interpretability of Transformer-based Language Models: a systematic review
by: López-Otal, Miguel, et al.
Published: (2025)
by: López-Otal, Miguel, et al.
Published: (2025)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
by: Zhang, Chuyifei, et al.
Published: (2026)
by: Zhang, Chuyifei, et al.
Published: (2026)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
by: Ansell, Rebecca, et al.
Published: (2026)
by: Ansell, Rebecca, et al.
Published: (2026)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
by: Gokdemir, Ozan, et al.
Published: (2025)
by: Gokdemir, Ozan, et al.
Published: (2025)
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
by: Annese, Luca, et al.
Published: (2025)
by: Annese, Luca, et al.
Published: (2025)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
by: Kebe, Gaoussou Youssouf, et al.
Published: (2025)
Improving LLMs with a knowledge from databases
by: Máša, Petr
Published: (2025)
by: Máša, Petr
Published: (2025)
AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models
by: Perez, Miguel Angel Peñaloza, et al.
Published: (2025)
by: Perez, Miguel Angel Peñaloza, et al.
Published: (2025)
LMLPA: Language Model Linguistic Personality Assessment
by: Zheng, Jingyao, et al.
Published: (2024)
by: Zheng, Jingyao, et al.
Published: (2024)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
by: Gwak, Jiho, et al.
Published: (2025)
by: Gwak, Jiho, et al.
Published: (2025)
Similar Items
-
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026) -
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026) -
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026) -
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)