FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911285900214272 |
|---|---|
| author | Qin, Bowen Yue, Chen Yin, Fang Wang, Hui Yao, JG Liu, Jiakang Zheng, Jing-Shu Chen, Miguel Hu Xuan, Richeng Meng, Shibei Zhou, Shiqi Dai, Teng Ren, Tong-Shuai Cui, Wei Yang, Xi Du, Xialin Xu, Xiaojing Sun, Xue Li, Xuejing Liu, Yaming Liu, Yesheng Liu, Ying Lin, Yonghua Zhao, Yu Zhang, Yunduo Luo, Yuwen He, Zheqi He, Zhiyuan Wang, Zhongyuan |
| author_facet | Qin, Bowen Yue, Chen Yin, Fang Wang, Hui Yao, JG Liu, Jiakang Zheng, Jing-Shu Chen, Miguel Hu Xuan, Richeng Meng, Shibei Zhou, Shiqi Dai, Teng Ren, Tong-Shuai Cui, Wei Yang, Xi Du, Xialin Xu, Xiaojing Sun, Xue Li, Xuejing Liu, Yaming Liu, Yesheng Liu, Ying Lin, Yonghua Zhao, Yu Zhang, Yunduo Luo, Yuwen He, Zheqi He, Zhiyuan Wang, Zhongyuan |
| contents | We conduct a moderate-scale contamination-free (to some extent) evaluation of current large reasoning models (LRMs) with some preliminary findings. We also release ROME, our evaluation benchmark for vision language models intended to test reasoning from visual clues. We attach links to the benchmark, evaluation data, and other updates on this website: https://flageval-baai.github.io/LRM-Eval/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17177 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions Qin, Bowen Yue, Chen Yin, Fang Wang, Hui Yao, JG Liu, Jiakang Zheng, Jing-Shu Chen, Miguel Hu Xuan, Richeng Meng, Shibei Zhou, Shiqi Dai, Teng Ren, Tong-Shuai Cui, Wei Yang, Xi Du, Xialin Xu, Xiaojing Sun, Xue Li, Xuejing Liu, Yaming Liu, Yesheng Liu, Ying Lin, Yonghua Zhao, Yu Zhang, Yunduo Luo, Yuwen He, Zheqi He, Zhiyuan Wang, Zhongyuan Computation and Language Computer Vision and Pattern Recognition Machine Learning We conduct a moderate-scale contamination-free (to some extent) evaluation of current large reasoning models (LRMs) with some preliminary findings. We also release ROME, our evaluation benchmark for vision language models intended to test reasoning from visual clues. We attach links to the benchmark, evaluation data, and other updates on this website: https://flageval-baai.github.io/LRM-Eval/ |
| title | FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions |
| topic | Computation and Language Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2509.17177 |