LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ren, Huimin, Liang, Yan, Su, Baiqiao, Sun, Chaobo, Lu, Hengtong, Zhang, Kaike, Wei, Chen
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917355968266240
author Ren, Huimin
Liang, Yan
Su, Baiqiao
Sun, Chaobo
Lu, Hengtong
Zhang, Kaike
Wei, Chen
author_facet Ren, Huimin
Liang, Yan
Su, Baiqiao
Sun, Chaobo
Lu, Hengtong
Zhang, Kaike
Wei, Chen
contents The ability of Large Language Models (LLMs) to precisely follow complex and fine-grained lexical instructions is a cornerstone of their utility and controllability. However, evaluating this capability remains a significant challenge. Current methods either rely on subjective and costly human evaluation or on automated LLM-as-a-judge systems, which suffer from inherent biases and unreliability. Existing programmatic benchmarks, while objective, often lack the expressiveness to test intricate, compositional constraints at a granular level. To address these limitations, we introduce LexInstructEval, a new benchmark and evaluation framework for fine-grained lexical instruction following. Our framework is built upon a formal, rule-based grammar that deconstructs complex instructions into a canonical <Procedure, Relation, Value> triplet. This grammar enables the systematic generation of a diverse dataset through a multi-stage, human-in-the-loop pipeline and facilitates objective verification via a transparent, programmatic engine. We release our dataset and open-source evaluation tools to facilitate further research into the controllability and reliability of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17561
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
Ren, Huimin
Liang, Yan
Su, Baiqiao
Sun, Chaobo
Lu, Hengtong
Zhang, Kaike
Wei, Chen
Computation and Language
Artificial Intelligence
The ability of Large Language Models (LLMs) to precisely follow complex and fine-grained lexical instructions is a cornerstone of their utility and controllability. However, evaluating this capability remains a significant challenge. Current methods either rely on subjective and costly human evaluation or on automated LLM-as-a-judge systems, which suffer from inherent biases and unreliability. Existing programmatic benchmarks, while objective, often lack the expressiveness to test intricate, compositional constraints at a granular level. To address these limitations, we introduce LexInstructEval, a new benchmark and evaluation framework for fine-grained lexical instruction following. Our framework is built upon a formal, rule-based grammar that deconstructs complex instructions into a canonical <Procedure, Relation, Value> triplet. This grammar enables the systematic generation of a diverse dataset through a multi-stage, human-in-the-loop pipeline and facilitates objective verification via a transparent, programmatic engine. We release our dataset and open-source evaluation tools to facilitate further research into the controllability and reliability of LLMs.
title LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.17561