Debiasing Reward Models by Representation Learning with Guarantees

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ng, Ignavier, Blöbaum, Patrick, Bhandari, Siddharth, Zhang, Kun, Kasiviswanathan, Shiva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!