Pavlenko, K., Golubev, A., Karasik, S., & Yangel, B. (2026). Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards.
Style de citation Chicago (17e éd.)Pavlenko, Kirill, Alexander Golubev, Simon Karasik, et Boris Yangel. Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards. 2026.
Style de citation MLA (9e éd.)Pavlenko, Kirill, et al. Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards. 2026.
Attention : ces citations peuvent ne pas être correctes à 100%.