DIALEVAL: Automated Type-Theoretic Evaluation of LLM Instruction Following

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Basta, Nardine, Kaafar, Dali
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918368753221632
author Basta, Nardine
Kaafar, Dali
author_facet Basta, Nardine
Kaafar, Dali
contents Evaluating instruction following in Large Language Models requires decomposing instructions into verifiable requirements and assessing satisfaction--tasks currently dependent on manual annotation and uniform criteria that do not align with human judgment patterns. We present DIALEVAL, a type-theoretic framework using dual LLM agents to automate instruction decomposition into typed predicates and implement type-specific satisfaction semantics. The framework enforces formal atomicity and independence constraints during automated extraction, then applies differentiated evaluation criteria--semantic equivalence for content predicates, exact precision for numerical predicates--mirroring empirically observed human assessment patterns. Extended to multi-turn dialogues through history-aware satisfaction functions, DIALEVAL enables evaluation in conversational contexts where single-turn methods fail. Validation demonstrates 90.38% accuracy (26.45% error reduction over baselines) and substantially stronger correlation with human judgment for complex instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DIALEVAL: Automated Type-Theoretic Evaluation of LLM Instruction Following
Basta, Nardine
Kaafar, Dali
Computation and Language
Artificial Intelligence
Evaluating instruction following in Large Language Models requires decomposing instructions into verifiable requirements and assessing satisfaction--tasks currently dependent on manual annotation and uniform criteria that do not align with human judgment patterns. We present DIALEVAL, a type-theoretic framework using dual LLM agents to automate instruction decomposition into typed predicates and implement type-specific satisfaction semantics. The framework enforces formal atomicity and independence constraints during automated extraction, then applies differentiated evaluation criteria--semantic equivalence for content predicates, exact precision for numerical predicates--mirroring empirically observed human assessment patterns. Extended to multi-turn dialogues through history-aware satisfaction functions, DIALEVAL enables evaluation in conversational contexts where single-turn methods fail. Validation demonstrates 90.38% accuracy (26.45% error reduction over baselines) and substantially stronger correlation with human judgment for complex instructions.
title DIALEVAL: Automated Type-Theoretic Evaluation of LLM Instruction Following
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.03321