Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915832123097088 |
|---|---|
| author | Gao, Chongyang Yang, Diji Zhou, Shuyan Yan, Xichen Song, Luchuan Li, Shuo Chen, Kezhen |
| author_facet | Gao, Chongyang Yang, Diji Zhou, Shuyan Yan, Xichen Song, Luchuan Li, Shuo Chen, Kezhen |
| contents | We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university homework and exam problems, paired with reference solutions provided by course instructors. CFE-Bench remains challenging for frontier models: the newly released Gemini-3.1-pro-preview achieves 59.69% overall accuracy, while the second-best model, Gemini-3-flash-preview, reaches 55.46%, leaving substantial room for improvement. Beyond aggregate scores, we conduct a diagnostic analysis by decomposing instructor reference solutions into structured reasoning flows. We find that while frontier models often answer intermediate sub-questions correctly, they struggle to reliably derive and maintain correct intermediate states throughout multi-step solutions. We further observe that model-generated solutions typically contain more reasoning steps than instructor solutions, indicating lower step efficiency and a higher risk of error accumulation. Data and code are available at https://github.com/Analogy-AI/CFE_Bench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_19517 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Classroom Final Exam: An Instructor-Tested Reasoning Benchmark Gao, Chongyang Yang, Diji Zhou, Shuyan Yan, Xichen Song, Luchuan Li, Shuo Chen, Kezhen Artificial Intelligence Computational Engineering, Finance, and Science Computation and Language Computer Vision and Pattern Recognition We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university homework and exam problems, paired with reference solutions provided by course instructors. CFE-Bench remains challenging for frontier models: the newly released Gemini-3.1-pro-preview achieves 59.69% overall accuracy, while the second-best model, Gemini-3-flash-preview, reaches 55.46%, leaving substantial room for improvement. Beyond aggregate scores, we conduct a diagnostic analysis by decomposing instructor reference solutions into structured reasoning flows. We find that while frontier models often answer intermediate sub-questions correctly, they struggle to reliably derive and maintain correct intermediate states throughout multi-step solutions. We further observe that model-generated solutions typically contain more reasoning steps than instructor solutions, indicating lower step efficiency and a higher risk of error accumulation. Data and code are available at https://github.com/Analogy-AI/CFE_Bench. |
| title | Classroom Final Exam: An Instructor-Tested Reasoning Benchmark |
| topic | Artificial Intelligence Computational Engineering, Finance, and Science Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.19517 |