Li, X., Li, M., Men, R., Zhang, Y., Bao, K., Wang, W., . . . Lin, J. (2025). HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning.
Chicago Style (17th ed.) CitationLi, Xiaoyuan, Moxin Li, Rui Men, Yichang Zhang, Keqin Bao, Wenjie Wang, Fuli Feng, Dayiheng Liu, and Junyang Lin. HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning. 2025.
MLA (9th ed.) CitationLi, Xiaoyuan, et al. HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning. 2025.
Warning: These citations may not always be 100% accurate.