Galatolo, A., Dai, Z., Winkle, K., & Beloucif, M. (2025). Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models.
Chicago-Zitierstil (17. Ausg.)Galatolo, Alessio, Zhenbang Dai, Katie Winkle, und Meriem Beloucif. Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models. 2025.
MLA-Zitierstil (9. Ausg.)Galatolo, Alessio, et al. Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models. 2025.
Achtung: Diese Zitate sind unter Umständen nicht zu 100% korrekt.