Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Ming, Kyng, Rasmus, Solda, Federico, Yuan, Weixuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918026524229632
author Ding, Ming
Kyng, Rasmus
Solda, Federico
Yuan, Weixuan
author_facet Ding, Ming
Kyng, Rasmus
Solda, Federico
Yuan, Weixuan
contents As large language models (LLMs) advance, their role in higher education, particularly in free-response problem-solving, requires careful examination. This study assesses the performance of GPT-4o and o1-preview under realistic educational conditions in an undergraduate algorithms course. Anonymous GPT-generated solutions to take-home exams were graded by teaching assistants unaware of their origin. Our analysis examines both coarse-grained performance (scores) and fine-grained reasoning quality (error patterns). Results show that GPT-4o consistently struggles, failing to reach the passing threshold, while o1-preview performs significantly better, surpassing the passing score and even exceeding the student median in certain exercises. However, both models exhibit issues with unjustified claims and misleading arguments. These findings highlight the need for robust assessment strategies and AI-aware grading policies in education.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13664
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading
Ding, Ming
Kyng, Rasmus
Solda, Federico
Yuan, Weixuan
Computers and Society
Computation and Language
As large language models (LLMs) advance, their role in higher education, particularly in free-response problem-solving, requires careful examination. This study assesses the performance of GPT-4o and o1-preview under realistic educational conditions in an undergraduate algorithms course. Anonymous GPT-generated solutions to take-home exams were graded by teaching assistants unaware of their origin. Our analysis examines both coarse-grained performance (scores) and fine-grained reasoning quality (error patterns). Results show that GPT-4o consistently struggles, failing to reach the passing threshold, while o1-preview performs significantly better, surpassing the passing score and even exceeding the student median in certain exercises. However, both models exhibit issues with unjustified claims and misleading arguments. These findings highlight the need for robust assessment strategies and AI-aware grading policies in education.
title Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading
topic Computers and Society
Computation and Language
url https://arxiv.org/abs/2505.13664