Kwa, T., Thomas, D., & Garriga-Alonso, A. (2024). Catastrophic Goodhart: Regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification.
Style de citation Chicago (17e éd.)Kwa, Thomas, Drake Thomas, et Adrià Garriga-Alonso. Catastrophic Goodhart: Regularizing RLHF with KL Divergence Does Not Mitigate Heavy-tailed Reward Misspecification. 2024.
Style de citation MLA (9e éd.)Kwa, Thomas, et al. Catastrophic Goodhart: Regularizing RLHF with KL Divergence Does Not Mitigate Heavy-tailed Reward Misspecification. 2024.
Attention : ces citations peuvent ne pas être correctes à 100%.