CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bernard, Nolwenn, Joko, Hideaki, Hasibi, Faegheh, Balog, Krisztian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916523343347712
author Bernard, Nolwenn
Joko, Hideaki
Hasibi, Faegheh
Balog, Krisztian
author_facet Bernard, Nolwenn
Joko, Hideaki
Hasibi, Faegheh
Balog, Krisztian
contents We introduce CRS Arena, a research platform for scalable benchmarking of Conversational Recommender Systems (CRS) based on human feedback. The platform displays pairwise battles between anonymous conversational recommender systems, where users interact with the systems one after the other before declaring either a winner or a draw. CRS Arena collects conversations and user feedback, providing a foundation for reliable evaluation and ranking of CRSs. We conduct experiments with CRS Arena on both open and closed crowdsourcing platforms, confirming that both setups produce highly correlated rankings of CRSs and conversations with similar characteristics. We release CRSArena-Dial, a dataset of 474 conversations and their corresponding user feedback, along with a preliminary ranking of the systems based on the Elo rating system. The platform is accessible at https://iai-group-crsarena.hf.space/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10514
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems
Bernard, Nolwenn
Joko, Hideaki
Hasibi, Faegheh
Balog, Krisztian
Information Retrieval
We introduce CRS Arena, a research platform for scalable benchmarking of Conversational Recommender Systems (CRS) based on human feedback. The platform displays pairwise battles between anonymous conversational recommender systems, where users interact with the systems one after the other before declaring either a winner or a draw. CRS Arena collects conversations and user feedback, providing a foundation for reliable evaluation and ranking of CRSs. We conduct experiments with CRS Arena on both open and closed crowdsourcing platforms, confirming that both setups produce highly correlated rankings of CRSs and conversations with similar characteristics. We release CRSArena-Dial, a dataset of 474 conversations and their corresponding user feedback, along with a preliminary ranking of the systems based on the Elo rating system. The platform is accessible at https://iai-group-crsarena.hf.space/.
title CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems
topic Information Retrieval
url https://arxiv.org/abs/2412.10514