CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Jingbo, Yao, Guanyu, Hou, Bairu, Yang, Xinghan, Glushnev, Nikolai, Bialynicka-Birula, Iwona, Ding, Duo, Chang, Shiyu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917406256922624
author Yang, Jingbo
Yao, Guanyu
Hou, Bairu
Yang, Xinghan
Glushnev, Nikolai
Bialynicka-Birula, Iwona
Ding, Duo
Chang, Shiyu
author_facet Yang, Jingbo
Yao, Guanyu
Hou, Bairu
Yang, Xinghan
Glushnev, Nikolai
Bialynicka-Birula, Iwona
Ding, Duo
Chang, Shiyu
contents As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex, domain-specific operational guidelines is critical. While utilizing an LLM-as-a-Judge is a promising solution for scalable evaluation, the reliability of these judges in detecting specific policy violations remains largely unexplored. This gap is primarily due to the lack of a systematic data generation method, which has been hindered by the extensive cost of fine-grained human annotation and the difficulty of synthesizing realistic agent violations. In this paper, we introduce CompliBench, a novel benchmark designed to evaluate the ability of LLM judges to detect and localize guideline violations in multi-turn dialogues. To overcome data scarcity, we develop a scalable, automated data generation pipeline that simulates user-agent interactions. Our controllable flaw injection process automatically yields precise ground-truth labels for the violated guideline and the exact conversation turn, while an adversarial search method ensures these introduced perturbations are highly challenging. Our comprehensive evaluation reveals that current state-of-the-art proprietary LLMs struggle significantly with this task. In addition, we demonstrate that a small-scale judge model fine-tuned on our synthesized data outperforms leading LLMs and generalizes well to unseen business domains, highlighting our pipeline as an effective foundation for training robust generative reward models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12312
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
Yang, Jingbo
Yao, Guanyu
Hou, Bairu
Yang, Xinghan
Glushnev, Nikolai
Bialynicka-Birula, Iwona
Ding, Duo
Chang, Shiyu
Computation and Language
As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex, domain-specific operational guidelines is critical. While utilizing an LLM-as-a-Judge is a promising solution for scalable evaluation, the reliability of these judges in detecting specific policy violations remains largely unexplored. This gap is primarily due to the lack of a systematic data generation method, which has been hindered by the extensive cost of fine-grained human annotation and the difficulty of synthesizing realistic agent violations. In this paper, we introduce CompliBench, a novel benchmark designed to evaluate the ability of LLM judges to detect and localize guideline violations in multi-turn dialogues. To overcome data scarcity, we develop a scalable, automated data generation pipeline that simulates user-agent interactions. Our controllable flaw injection process automatically yields precise ground-truth labels for the violated guideline and the exact conversation turn, while an adversarial search method ensures these introduced perturbations are highly challenging. Our comprehensive evaluation reveals that current state-of-the-art proprietary LLMs struggle significantly with this task. In addition, we demonstrate that a small-scale judge model fine-tuned on our synthesized data outperforms leading LLMs and generalizes well to unseen business domains, highlighting our pipeline as an effective foundation for training robust generative reward models.
title CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
topic Computation and Language
url https://arxiv.org/abs/2604.12312