Li, Y., Ma, L., Zhang, J., Tang, L., Zhang, W., & Luo, G. (2025). Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model.
Chicago Style (17th ed.) CitationLi, Yanhao, Lu Ma, Jiaran Zhang, Lexiang Tang, Wentao Zhang, and Guibo Luo. Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model. 2025.
MLA (9th ed.) CitationLi, Yanhao, et al. Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model. 2025.
Warning: These citations may not always be 100% accurate.