Zhang, Z., Dong, H., Pei, K., & Mao, C. (2026). R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning.
Chicago Style (17th ed.) CitationZhang, Zirui, Haoyu Dong, Kexin Pei, and Chengzhi Mao. R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning. 2026.
MLA (9th ed.) CitationZhang, Zirui, et al. R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning. 2026.
Warning: These citations may not always be 100% accurate.