Classroom Final Exam: An Instructor-Tested Reasoning Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Chongyang, Yang, Diji, Zhou, Shuyan, Yan, Xichen, Song, Luchuan, Li, Shuo, Chen, Kezhen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915832123097088
author Gao, Chongyang
Yang, Diji
Zhou, Shuyan
Yan, Xichen
Song, Luchuan
Li, Shuo
Chen, Kezhen
author_facet Gao, Chongyang
Yang, Diji
Zhou, Shuyan
Yan, Xichen
Song, Luchuan
Li, Shuo
Chen, Kezhen
contents We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university homework and exam problems, paired with reference solutions provided by course instructors. CFE-Bench remains challenging for frontier models: the newly released Gemini-3.1-pro-preview achieves 59.69% overall accuracy, while the second-best model, Gemini-3-flash-preview, reaches 55.46%, leaving substantial room for improvement. Beyond aggregate scores, we conduct a diagnostic analysis by decomposing instructor reference solutions into structured reasoning flows. We find that while frontier models often answer intermediate sub-questions correctly, they struggle to reliably derive and maintain correct intermediate states throughout multi-step solutions. We further observe that model-generated solutions typically contain more reasoning steps than instructor solutions, indicating lower step efficiency and a higher risk of error accumulation. Data and code are available at https://github.com/Analogy-AI/CFE_Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19517
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
Gao, Chongyang
Yang, Diji
Zhou, Shuyan
Yan, Xichen
Song, Luchuan
Li, Shuo
Chen, Kezhen
Artificial Intelligence
Computational Engineering, Finance, and Science
Computation and Language
Computer Vision and Pattern Recognition
We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university homework and exam problems, paired with reference solutions provided by course instructors. CFE-Bench remains challenging for frontier models: the newly released Gemini-3.1-pro-preview achieves 59.69% overall accuracy, while the second-best model, Gemini-3-flash-preview, reaches 55.46%, leaving substantial room for improvement. Beyond aggregate scores, we conduct a diagnostic analysis by decomposing instructor reference solutions into structured reasoning flows. We find that while frontier models often answer intermediate sub-questions correctly, they struggle to reliably derive and maintain correct intermediate states throughout multi-step solutions. We further observe that model-generated solutions typically contain more reasoning steps than instructor solutions, indicating lower step efficiency and a higher risk of error accumulation. Data and code are available at https://github.com/Analogy-AI/CFE_Bench.
title Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
topic Artificial Intelligence
Computational Engineering, Finance, and Science
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19517