Datasets for Navigating Sensitive Topics in Recommendation Systems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kovacs, Amelia, Chee, Jerry, Kazemian, Kimia, Dean, Sarah
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918138095861760
author Kovacs, Amelia
Chee, Jerry
Kazemian, Kimia
Dean, Sarah
author_facet Kovacs, Amelia
Chee, Jerry
Kazemian, Kimia
Dean, Sarah
contents Personalized AI systems, from recommendation systems to chatbots, are a prevalent method for distributing content to users based on their learned preferences. However, there is growing concern about the adverse effects of these systems, including their potential tendency to expose users to sensitive or harmful material, negatively impacting overall well-being. To address this concern quantitatively, it is necessary to create datasets with relevant sensitivity labels for content, enabling researchers to evaluate personalized systems beyond mere engagement metrics. To this end, we introduce two novel datasets that include a taxonomy of sensitivity labels alongside user-content ratings: one that integrates MovieLens rating data with content warnings from the Does the Dog Die? community ratings website, and another that combines fan-fiction interaction data and user-generated warnings from Archive of Our Own.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Datasets for Navigating Sensitive Topics in Recommendation Systems
Kovacs, Amelia
Chee, Jerry
Kazemian, Kimia
Dean, Sarah
Information Retrieval
Artificial Intelligence
Personalized AI systems, from recommendation systems to chatbots, are a prevalent method for distributing content to users based on their learned preferences. However, there is growing concern about the adverse effects of these systems, including their potential tendency to expose users to sensitive or harmful material, negatively impacting overall well-being. To address this concern quantitatively, it is necessary to create datasets with relevant sensitivity labels for content, enabling researchers to evaluate personalized systems beyond mere engagement metrics. To this end, we introduce two novel datasets that include a taxonomy of sensitivity labels alongside user-content ratings: one that integrates MovieLens rating data with content warnings from the Does the Dog Die? community ratings website, and another that combines fan-fiction interaction data and user-generated warnings from Archive of Our Own.
title Datasets for Navigating Sensitive Topics in Recommendation Systems
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2509.07269