Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gupta, Shashank, Oosterhuis, Harrie, de Rijke, Maarten
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909317229182976
author Gupta, Shashank
Oosterhuis, Harrie
de Rijke, Maarten
author_facet Gupta, Shashank
Oosterhuis, Harrie
de Rijke, Maarten
contents Counterfactual learning to rank (CLTR) can be risky and, in various circumstances, can produce sub-optimal models that hurt performance when deployed. Safe CLTR was introduced to mitigate these risks when using inverse propensity scoring to correct for position bias. However, the existing safety measure for CLTR is not applicable to state-of-the-art CLTR methods, cannot handle trust bias, and relies on specific assumptions about user behavior. We propose a novel approach, proximal ranking policy optimization (PRPO), that provides safety in deployment without assumptions about user behavior. PRPO removes incentives for learning ranking behavior that is too dissimilar to a safe ranking model. Thereby, PRPO imposes a limit on how much learned models can degrade performance metrics, without relying on any specific user assumptions. Our experiments show that PRPO provides higher performance than the existing safe inverse propensity scoring approach. PRPO always maintains safety, even in maximally adversarial situations. By avoiding assumptions, PRPO is the first method with unconditional safety in deployment that translates to robust safety for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09881
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank
Gupta, Shashank
Oosterhuis, Harrie
de Rijke, Maarten
Machine Learning
Information Retrieval
Counterfactual learning to rank (CLTR) can be risky and, in various circumstances, can produce sub-optimal models that hurt performance when deployed. Safe CLTR was introduced to mitigate these risks when using inverse propensity scoring to correct for position bias. However, the existing safety measure for CLTR is not applicable to state-of-the-art CLTR methods, cannot handle trust bias, and relies on specific assumptions about user behavior. We propose a novel approach, proximal ranking policy optimization (PRPO), that provides safety in deployment without assumptions about user behavior. PRPO removes incentives for learning ranking behavior that is too dissimilar to a safe ranking model. Thereby, PRPO imposes a limit on how much learned models can degrade performance metrics, without relying on any specific user assumptions. Our experiments show that PRPO provides higher performance than the existing safe inverse propensity scoring approach. PRPO always maintains safety, even in maximally adversarial situations. By avoiding assumptions, PRPO is the first method with unconditional safety in deployment that translates to robust safety for real-world applications.
title Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank
topic Machine Learning
Information Retrieval
url https://arxiv.org/abs/2409.09881