On the Reliability of User-Centric Evaluation of Conversational Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Müller, Michael, Mohammadi, Amir Reza, Peintner, Andreas, Gstrein, Beatriz Barroso, Specht, Günther, Zangerle, Eva
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912913003905024
author Müller, Michael
Mohammadi, Amir Reza
Peintner, Andreas
Gstrein, Beatriz Barroso
Specht, Günther
Zangerle, Eva
author_facet Müller, Michael
Mohammadi, Amir Reza
Peintner, Andreas
Gstrein, Beatriz Barroso
Specht, Günther
Zangerle, Eva
contents User-centric evaluation has become a key paradigm for assessing Conversational Recommender Systems (CRS), aiming to capture subjective qualities such as satisfaction, trust, and rapport. To enable scalable evaluation, recent work increasingly relies on third-party annotations of static dialogue logs by crowd workers or large language models. However, the reliability of this practice remains largely unexamined. In this paper, we present a large-scale empirical study investigating the reliability and structure of user-centric CRS evaluation on static dialogue transcripts. We collected 1,053 annotations from 124 crowd workers on 200 ReDial dialogues using the 18-dimensional CRS-Que framework. Using random-effects reliability models and correlation analysis, we quantify the stability of individual dimensions and their interdependencies. Our results show that utilitarian and outcome-oriented dimensions such as accuracy, usefulness, and satisfaction achieve moderate reliability under aggregation, whereas socially grounded constructs such as humanness and rapport are substantially less reliable. Furthermore, many dimensions collapse into a single global quality signal, revealing a strong halo effect in third-party judgments. These findings challenge the validity of single-annotator and LLM-based evaluation protocols and motivate the need for multi-rater aggregation and dimension reduction in offline CRS evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17264
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Reliability of User-Centric Evaluation of Conversational Recommender Systems
Müller, Michael
Mohammadi, Amir Reza
Peintner, Andreas
Gstrein, Beatriz Barroso
Specht, Günther
Zangerle, Eva
Information Retrieval
User-centric evaluation has become a key paradigm for assessing Conversational Recommender Systems (CRS), aiming to capture subjective qualities such as satisfaction, trust, and rapport. To enable scalable evaluation, recent work increasingly relies on third-party annotations of static dialogue logs by crowd workers or large language models. However, the reliability of this practice remains largely unexamined. In this paper, we present a large-scale empirical study investigating the reliability and structure of user-centric CRS evaluation on static dialogue transcripts. We collected 1,053 annotations from 124 crowd workers on 200 ReDial dialogues using the 18-dimensional CRS-Que framework. Using random-effects reliability models and correlation analysis, we quantify the stability of individual dimensions and their interdependencies. Our results show that utilitarian and outcome-oriented dimensions such as accuracy, usefulness, and satisfaction achieve moderate reliability under aggregation, whereas socially grounded constructs such as humanness and rapport are substantially less reliable. Furthermore, many dimensions collapse into a single global quality signal, revealing a strong halo effect in third-party judgments. These findings challenge the validity of single-annotator and LLM-based evaluation protocols and motivate the need for multi-rater aggregation and dimension reduction in offline CRS evaluation.
title On the Reliability of User-Centric Evaluation of Conversational Recommender Systems
topic Information Retrieval
url https://arxiv.org/abs/2602.17264