He, S., Xia, T., Zhou, X., & Wei, H. (2025). Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective.
Chicago Style (17th ed.) CitationHe, Shenghua, Tian Xia, Xuan Zhou, and Hui Wei. Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective. 2025.
MLA (9th ed.) CitationHe, Shenghua, et al. Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective. 2025.
Warning: These citations may not always be 100% accurate.