Wu, Y., Han, S., & Cai, H. (2026). Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.
Style de citation Chicago (17e éd.)Wu, Yecheng, Song Han, et Hai Cai. Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation. 2026.
Style de citation MLA (9e éd.)Wu, Yecheng, et al. Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation. 2026.
Attention : ces citations peuvent ne pas être correctes à 100%.