REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Su, Kun, Sayana, Krishna, Pham, Hubert, Pine, James, Vasilevski, Yuri, Vasudeva, Raghavendra, Kyriakidi, Marialena, Hebert, Liam, Jash, Ambarish, Subbiah, Anushya, Sodhi, Sukhdeep
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918088989999104
author Su, Kun
Sayana, Krishna
Pham, Hubert
Pine, James
Vasilevski, Yuri
Vasudeva, Raghavendra
Kyriakidi, Marialena
Hebert, Liam
Jash, Ambarish
Subbiah, Anushya
Sodhi, Sukhdeep
author_facet Su, Kun
Sayana, Krishna
Pham, Hubert
Pine, James
Vasilevski, Yuri
Vasudeva, Raghavendra
Kyriakidi, Marialena
Hebert, Liam
Jash, Ambarish
Subbiah, Anushya
Sodhi, Sukhdeep
contents This paper introduces a novel dataset REGEN (Reviews Enhanced with GEnerative Narratives), designed to benchmark the conversational capabilities of recommender Large Language Models (LLMs), addressing the limitations of existing datasets that primarily focus on sequential item prediction. REGEN extends the Amazon Product Reviews dataset by inpainting two key natural language features: (1) user critiques, representing user "steering" queries that lead to the selection of a subsequent item, and (2) narratives, rich textual outputs associated with each recommended item taking into account prior context. The narratives include product endorsements, purchase explanations, and summaries of user preferences. Further, we establish an end-to-end modeling benchmark for the task of conversational recommendation, where models are trained to generate both recommendations and corresponding narratives conditioned on user history (items and critiques). For this joint task, we introduce a modeling framework LUMEN (LLM-based Unified Multi-task Model with Critiques, Recommendations, and Narratives) which uses an LLM as a backbone for critiquing, retrieval and generation. We also evaluate the dataset's quality using standard auto-rating techniques and benchmark it by training both traditional and LLM-based recommender models. Our results demonstrate that incorporating critiques enhances recommendation quality by enabling the recommender to learn language understanding and integrate it with recommendation signals. Furthermore, LLMs trained on our dataset effectively generate both recommendations and contextual narratives, achieving performance comparable to state-of-the-art recommenders and language models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives
Su, Kun
Sayana, Krishna
Pham, Hubert
Pine, James
Vasilevski, Yuri
Vasudeva, Raghavendra
Kyriakidi, Marialena
Hebert, Liam
Jash, Ambarish
Subbiah, Anushya
Sodhi, Sukhdeep
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
This paper introduces a novel dataset REGEN (Reviews Enhanced with GEnerative Narratives), designed to benchmark the conversational capabilities of recommender Large Language Models (LLMs), addressing the limitations of existing datasets that primarily focus on sequential item prediction. REGEN extends the Amazon Product Reviews dataset by inpainting two key natural language features: (1) user critiques, representing user "steering" queries that lead to the selection of a subsequent item, and (2) narratives, rich textual outputs associated with each recommended item taking into account prior context. The narratives include product endorsements, purchase explanations, and summaries of user preferences. Further, we establish an end-to-end modeling benchmark for the task of conversational recommendation, where models are trained to generate both recommendations and corresponding narratives conditioned on user history (items and critiques). For this joint task, we introduce a modeling framework LUMEN (LLM-based Unified Multi-task Model with Critiques, Recommendations, and Narratives) which uses an LLM as a backbone for critiquing, retrieval and generation. We also evaluate the dataset's quality using standard auto-rating techniques and benchmark it by training both traditional and LLM-based recommender models. Our results demonstrate that incorporating critiques enhances recommendation quality by enabling the recommender to learn language understanding and integrate it with recommendation signals. Furthermore, LLMs trained on our dataset effectively generate both recommendations and contextual narratives, achieving performance comparable to state-of-the-art recommenders and language models.
title REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2503.11924