LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917355968266240 |
|---|---|
| author | Ren, Huimin Liang, Yan Su, Baiqiao Sun, Chaobo Lu, Hengtong Zhang, Kaike Wei, Chen |
| author_facet | Ren, Huimin Liang, Yan Su, Baiqiao Sun, Chaobo Lu, Hengtong Zhang, Kaike Wei, Chen |
| contents | The ability of Large Language Models (LLMs) to precisely follow complex and fine-grained lexical instructions is a cornerstone of their utility and controllability. However, evaluating this capability remains a significant challenge. Current methods either rely on subjective and costly human evaluation or on automated LLM-as-a-judge systems, which suffer from inherent biases and unreliability. Existing programmatic benchmarks, while objective, often lack the expressiveness to test intricate, compositional constraints at a granular level. To address these limitations, we introduce LexInstructEval, a new benchmark and evaluation framework for fine-grained lexical instruction following. Our framework is built upon a formal, rule-based grammar that deconstructs complex instructions into a canonical <Procedure, Relation, Value> triplet. This grammar enables the systematic generation of a diverse dataset through a multi-stage, human-in-the-loop pipeline and facilitates objective verification via a transparent, programmatic engine. We release our dataset and open-source evaluation tools to facilitate further research into the controllability and reliability of LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_17561 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models Ren, Huimin Liang, Yan Su, Baiqiao Sun, Chaobo Lu, Hengtong Zhang, Kaike Wei, Chen Computation and Language Artificial Intelligence The ability of Large Language Models (LLMs) to precisely follow complex and fine-grained lexical instructions is a cornerstone of their utility and controllability. However, evaluating this capability remains a significant challenge. Current methods either rely on subjective and costly human evaluation or on automated LLM-as-a-judge systems, which suffer from inherent biases and unreliability. Existing programmatic benchmarks, while objective, often lack the expressiveness to test intricate, compositional constraints at a granular level. To address these limitations, we introduce LexInstructEval, a new benchmark and evaluation framework for fine-grained lexical instruction following. Our framework is built upon a formal, rule-based grammar that deconstructs complex instructions into a canonical <Procedure, Relation, Value> triplet. This grammar enables the systematic generation of a diverse dataset through a multi-stage, human-in-the-loop pipeline and facilitates objective verification via a transparent, programmatic engine. We release our dataset and open-source evaluation tools to facilitate further research into the controllability and reliability of LLMs. |
| title | LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2511.17561 |