Tang, Y., Cohen, T., Zhang, D. W., Valko, M., & Munos, R. (2025). RL-finetuning LLMs from on- and off-policy data with a single algorithm.
Citazione stile Chigago Style (17a edizione)Tang, Yunhao, Taco Cohen, David W. Zhang, Michal Valko, e Rémi Munos. RL-finetuning LLMs from on- and Off-policy Data with a Single Algorithm. 2025.
Citatione MLA (9a ed.)Tang, Yunhao, et al. RL-finetuning LLMs from on- and Off-policy Data with a Single Algorithm. 2025.
Attenzione: Queste citazioni potrebbero non essere precise al 100%.