Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., & Rakhlin, A. (2024). Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.
Chicago-Zitierstil (17. Ausg.)Xie, Tengyang, Dylan J. Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah, und Alexander Rakhlin. Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. 2024.
MLA-Zitierstil (9. Ausg.)Xie, Tengyang, et al. Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF. 2024.
Achtung: Diese Zitate sind unter Umständen nicht zu 100% korrekt.