OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nayeem, Mir Tafseer, Rafiei, Davood
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917054178656256
author Nayeem, Mir Tafseer
Rafiei, Davood
author_facet Nayeem, Mir Tafseer
Rafiei, Davood
contents We study the problem of opinion highlights generation from large volumes of user reviews, often exceeding thousands per entity, where existing methods either fail to scale or produce generic, one-size-fits-all summaries that overlook personalized needs. To tackle this, we introduce OpinioRAG, a scalable, training-free framework that combines RAG-based evidence retrieval with LLMs to efficiently produce tailored summaries. Additionally, we propose novel reference-free verification metrics designed for sentiment-rich domains, where accurately capturing opinions and sentiment alignment is essential. These metrics offer a fine-grained, context-sensitive assessment of factual consistency. To facilitate evaluation, we contribute the first large-scale dataset of long-form user reviews, comprising entities with over a thousand reviews each, paired with unbiased expert summaries and manually annotated queries. Through extensive experiments, we identify key challenges, provide actionable insights into improving systems, pave the way for future research, and position OpinioRAG as a robust framework for generating accurate, relevant, and structured summaries at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
Nayeem, Mir Tafseer
Rafiei, Davood
Computation and Language
Artificial Intelligence
Information Retrieval
We study the problem of opinion highlights generation from large volumes of user reviews, often exceeding thousands per entity, where existing methods either fail to scale or produce generic, one-size-fits-all summaries that overlook personalized needs. To tackle this, we introduce OpinioRAG, a scalable, training-free framework that combines RAG-based evidence retrieval with LLMs to efficiently produce tailored summaries. Additionally, we propose novel reference-free verification metrics designed for sentiment-rich domains, where accurately capturing opinions and sentiment alignment is essential. These metrics offer a fine-grained, context-sensitive assessment of factual consistency. To facilitate evaluation, we contribute the first large-scale dataset of long-form user reviews, comprising entities with over a thousand reviews each, paired with unbiased expert summaries and manually annotated queries. Through extensive experiments, we identify key challenges, provide actionable insights into improving systems, pave the way for future research, and position OpinioRAG as a robust framework for generating accurate, relevant, and structured summaries at scale.
title OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2509.00285