Saved in:
Bibliographic Details
Main Author: Nagpal, Chirag
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.13189
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Approaches for estimating preferences from human annotated data typically involves inducing a distribution over a ranked list of choices such as the Plackett-Luce model. Indeed, modern AI alignment tools such as Reward Modelling and Direct Preference Optimization are based on the statistical assumptions posed by the Plackett-Luce model. In this paper, I will connect the Plackett-Luce model to another classical and well known statistical model, the Cox Proportional Hazards model and attempt to shed some light on the implications of the connection therein.