Can We Trust Recommender System Fairness Evaluation? The Role of Fairness and Relevance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rampisela, Theresia Veronika, Ruotsalo, Tuukka, Maistro, Maria, Lioma, Christina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911891245236224
author Rampisela, Theresia Veronika
Ruotsalo, Tuukka
Maistro, Maria
Lioma, Christina
author_facet Rampisela, Theresia Veronika
Ruotsalo, Tuukka
Maistro, Maria
Lioma, Christina
contents Relevance and fairness are two major objectives of recommender systems (RSs). Recent work proposes measures of RS fairness that are either independent from relevance (fairness-only) or conditioned on relevance (joint measures). While fairness-only measures have been studied extensively, we look into whether joint measures can be trusted. We collect all joint evaluation measures of RS relevance and fairness, and ask: How much do they agree with each other? To what extent do they agree with relevance/fairness measures? How sensitive are they to changes in rank position, or to increasingly fair and relevant recommendations? We empirically study for the first time the behaviour of these measures across 4 real-world datasets and 4 recommenders. We find that most of these measures: i) correlate weakly with one another and even contradict each other at times; ii) are less sensitive to rank position changes than relevance- and fairness-only measures, meaning that they are less granular than traditional RS measures; and iii) tend to compress scores at the low end of their range, meaning that they are not very expressive. We counter the above limitations with a set of guidelines on the appropriate usage of such measures, i.e., they should be used with caution due to their tendency to contradict each other and of having a very small empirical range.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18276
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can We Trust Recommender System Fairness Evaluation? The Role of Fairness and Relevance
Rampisela, Theresia Veronika
Ruotsalo, Tuukka
Maistro, Maria
Lioma, Christina
Information Retrieval
Relevance and fairness are two major objectives of recommender systems (RSs). Recent work proposes measures of RS fairness that are either independent from relevance (fairness-only) or conditioned on relevance (joint measures). While fairness-only measures have been studied extensively, we look into whether joint measures can be trusted. We collect all joint evaluation measures of RS relevance and fairness, and ask: How much do they agree with each other? To what extent do they agree with relevance/fairness measures? How sensitive are they to changes in rank position, or to increasingly fair and relevant recommendations? We empirically study for the first time the behaviour of these measures across 4 real-world datasets and 4 recommenders. We find that most of these measures: i) correlate weakly with one another and even contradict each other at times; ii) are less sensitive to rank position changes than relevance- and fairness-only measures, meaning that they are less granular than traditional RS measures; and iii) tend to compress scores at the low end of their range, meaning that they are not very expressive. We counter the above limitations with a set of guidelines on the appropriate usage of such measures, i.e., they should be used with caution due to their tendency to contradict each other and of having a very small empirical range.
title Can We Trust Recommender System Fairness Evaluation? The Role of Fairness and Relevance
topic Information Retrieval
url https://arxiv.org/abs/2405.18276