Pang, L., Luo, J., & Jin, R. (2025). TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback.
Chicago Style (17th ed.) CitationPang, Lei, Jun Luo, and Ruinan Jin. TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback. 2025.
MLA (9th ed.) CitationPang, Lei, et al. TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback. 2025.
Warning: These citations may not always be 100% accurate.