UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chao, Wu, Neo, Ning, Lin, Wu, Jiaxing, Liu, Luyang, Xie, Jun, O'Banion, Shawn, Green, Bradley
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917770129571840
author Wang, Chao
Wu, Neo
Ning, Lin
Wu, Jiaxing
Liu, Luyang
Xie, Jun
O'Banion, Shawn
Green, Bradley
author_facet Wang, Chao
Wu, Neo
Ning, Lin
Wu, Jiaxing
Liu, Luyang
Xie, Jun
O'Banion, Shawn
Green, Bradley
contents Large language models (LLMs) have shown remarkable capabilities in generating user summaries from a long list of raw user activity data. These summaries capture essential user information such as preferences and interests, and therefore are invaluable for LLM-based personalization applications, such as explainable recommender systems. However, the development of new summarization techniques is hindered by the lack of ground-truth labels, the inherent subjectivity of user summaries, and human evaluation which is often costly and time-consuming. To address these challenges, we introduce \UserSumBench, a benchmark framework designed to facilitate iterative development of LLM-based summarization approaches. This framework offers two key components: (1) A reference-free summary quality metric. We show that this metric is effective and aligned with human preferences across three diverse datasets (MovieLens, Yelp and Amazon Review). (2) A novel robust summarization method that leverages time-hierarchical summarizer and self-critique verifier to produce high-quality summaries while eliminating hallucination. This method serves as a strong baseline for further innovation in summarization techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16966
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
Wang, Chao
Wu, Neo
Ning, Lin
Wu, Jiaxing
Liu, Luyang
Xie, Jun
O'Banion, Shawn
Green, Bradley
Machine Learning
Artificial Intelligence
Computation and Language
Large language models (LLMs) have shown remarkable capabilities in generating user summaries from a long list of raw user activity data. These summaries capture essential user information such as preferences and interests, and therefore are invaluable for LLM-based personalization applications, such as explainable recommender systems. However, the development of new summarization techniques is hindered by the lack of ground-truth labels, the inherent subjectivity of user summaries, and human evaluation which is often costly and time-consuming. To address these challenges, we introduce \UserSumBench, a benchmark framework designed to facilitate iterative development of LLM-based summarization approaches. This framework offers two key components: (1) A reference-free summary quality metric. We show that this metric is effective and aligned with human preferences across three diverse datasets (MovieLens, Yelp and Amazon Review). (2) A novel robust summarization method that leverages time-hierarchical summarizer and self-critique verifier to produce high-quality summaries while eliminating hallucination. This method serves as a strong baseline for further innovation in summarization techniques.
title UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2408.16966