Inference-Time Reward Hacking in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khalaf, Hadi, Verdun, Claudio Mayrink, Oesterling, Alex, Lakkaraju, Himabindu, Calmon, Flavio du Pin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!