Rethinking Human Preference Evaluation of LLM Rationales

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ziang, Ganti, Manasi, Ma, Zixian, Vasconcelos, Helena, He, Qijia, Krishna, Ranjay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912585636380672
author Li, Ziang
Ganti, Manasi
Ma, Zixian
Vasconcelos, Helena
He, Qijia
Krishna, Ranjay
author_facet Li, Ziang
Ganti, Manasi
Ma, Zixian
Vasconcelos, Helena
He, Qijia
Krishna, Ranjay
contents Large language models (LLMs) often generate natural language rationales -- free-form explanations that help improve performance on complex reasoning tasks and enhance interpretability for human users. However, evaluating these rationales remains challenging. While recent work has relied on binary preference judgments from humans or LLM judges, such evaluations are often opaque and coarse-grained, offering limited insight into what makes one rationale better than another. In this work, we rethink preference evaluation for LLM-generated rationales by asking: (1) What attributes define good rationales? (2) Can human preferences be explained by these attributes? (3) Can attribute-based evaluation overcome the limitations of binary comparisons? We identify a set of key rationale attributes from prior literature and assess them using automatic metrics, LLM judgments, and human annotations. We then analyze two standard human preference datasets MT Bench and Chatbot Arena using SHAP to identify which attributes best explain human preference outcomes. Finally, we re-evaluate model-generated rationales using attribute-specific ELO scores, revealing more nuanced model comparisons and insights. Our findings suggest that fine-grained attribute evaluations can better characterize rationale quality and guide future research toward more interpretable and reliable evaluation practices.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11026
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Human Preference Evaluation of LLM Rationales
Li, Ziang
Ganti, Manasi
Ma, Zixian
Vasconcelos, Helena
He, Qijia
Krishna, Ranjay
Artificial Intelligence
Computation and Language
Large language models (LLMs) often generate natural language rationales -- free-form explanations that help improve performance on complex reasoning tasks and enhance interpretability for human users. However, evaluating these rationales remains challenging. While recent work has relied on binary preference judgments from humans or LLM judges, such evaluations are often opaque and coarse-grained, offering limited insight into what makes one rationale better than another. In this work, we rethink preference evaluation for LLM-generated rationales by asking: (1) What attributes define good rationales? (2) Can human preferences be explained by these attributes? (3) Can attribute-based evaluation overcome the limitations of binary comparisons? We identify a set of key rationale attributes from prior literature and assess them using automatic metrics, LLM judgments, and human annotations. We then analyze two standard human preference datasets MT Bench and Chatbot Arena using SHAP to identify which attributes best explain human preference outcomes. Finally, we re-evaluate model-generated rationales using attribute-specific ELO scores, revealing more nuanced model comparisons and insights. Our findings suggest that fine-grained attribute evaluations can better characterize rationale quality and guide future research toward more interpretable and reliable evaluation practices.
title Rethinking Human Preference Evaluation of LLM Rationales
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.11026