ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nkhata, Gibson, Oyshi, Uttamasha Anjally, Mai, Quan, Gauch, Susan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911695889235968
author Nkhata, Gibson
Oyshi, Uttamasha Anjally
Mai, Quan
Gauch, Susan
author_facet Nkhata, Gibson
Oyshi, Uttamasha Anjally
Mai, Quan
Gauch, Susan
contents Personalized Retrieval-Augmented Generation (RAG) relies on accurately selecting user-relevant documents. In practice, existing RAG approaches often suffer from high retrieval costs and overlook that collaborative signals from similar users can enhance personalized generation for the current user. We propose ClusterRAG, a Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation. ClusterRAG represents users through their profile documents, organizes users into semantically coherent clusters using density-based clustering, and performs retrieval at both the cluster and document levels via cluster-level similarity and fine-grained ranking. Extensive experiments on the LaMP benchmark demonstrate that jointly leveraging the target user's profile and profiles from top similar users consistently yields the best performance across diverse tasks. Further analysis shows that ClusterRAG integrates seamlessly with different dense retrievers and rankers, and remains effective when paired with both fine-tuned and zero-shot language models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18769
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation
Nkhata, Gibson
Oyshi, Uttamasha Anjally
Mai, Quan
Gauch, Susan
Information Retrieval
Artificial Intelligence
Computation and Language
Personalized Retrieval-Augmented Generation (RAG) relies on accurately selecting user-relevant documents. In practice, existing RAG approaches often suffer from high retrieval costs and overlook that collaborative signals from similar users can enhance personalized generation for the current user. We propose ClusterRAG, a Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation. ClusterRAG represents users through their profile documents, organizes users into semantically coherent clusters using density-based clustering, and performs retrieval at both the cluster and document levels via cluster-level similarity and fine-grained ranking. Extensive experiments on the LaMP benchmark demonstrate that jointly leveraging the target user's profile and profiles from top similar users consistently yields the best performance across diverse tasks. Further analysis shows that ClusterRAG integrates seamlessly with different dense retrievers and rankers, and remains effective when paired with both fine-tuned and zero-shot language models.
title ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.18769