Nguyen, D., Prasad, A., Stengel-Eskin, E., & Bansal, M. (2024). LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits.
Chicago-Zitierstil (17. Ausg.)Nguyen, Duy, Archiki Prasad, Elias Stengel-Eskin, und Mohit Bansal. LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits. 2024.
MLA-Zitierstil (9. Ausg.)Nguyen, Duy, et al. LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits. 2024.
Achtung: Diese Zitate sind unter Umständen nicht zu 100% korrekt.