Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Ming, Kyng, Rasmus, Solda, Federico, Yuan, Weixuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Artificial Intelligence Driven Course Generation: A Case Study Using ChatGPT
by: Rouabhia, Djaber
Published: (2024)
by: Rouabhia, Djaber
Published: (2024)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
by: Caraeni, Adriana, et al.
Published: (2024)
by: Caraeni, Adriana, et al.
Published: (2024)
Analyzing the Performance of ChatGPT in Cardiology and Vascular Pathologies
by: Hariri, Walid
Published: (2023)
by: Hariri, Walid
Published: (2023)
A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
by: Yeadon, Will, et al.
Published: (2024)
by: Yeadon, Will, et al.
Published: (2024)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
by: Torrielli, Federico
Published: (2024)
by: Torrielli, Federico
Published: (2024)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Generative AI in Higher Education: Seeing ChatGPT Through Universities' Policies, Resources, and Guidelines
by: Wang, Hui, et al.
Published: (2023)
by: Wang, Hui, et al.
Published: (2023)
Improving Graduate Outcomes by Identifying Skills Gaps and Recommending Courses Based on Career Interests
by: Soni, Rahul, et al.
Published: (2025)
by: Soni, Rahul, et al.
Published: (2025)
An Exploration of Higher Education Course Evaluation by Large Language Models
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
SteLLA: A Structured Grading System Using LLMs with RAG
by: Qiu, Hefei, et al.
Published: (2025)
by: Qiu, Hefei, et al.
Published: (2025)
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
ChatGPT for President! Presupposed content in politicians versus GPT-generated texts
by: Garassino, Davide, et al.
Published: (2025)
by: Garassino, Davide, et al.
Published: (2025)
Application of GPT Language Models for Innovation in Activities in University Teaching
by: de Buenaga, Manuel, et al.
Published: (2024)
by: de Buenaga, Manuel, et al.
Published: (2024)
Assessing Engineering Student Perceptions of Introductory CS Courses in an Indian Context
by: Nareti, Utsav Kumar, et al.
Published: (2025)
by: Nareti, Utsav Kumar, et al.
Published: (2025)
DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4
by: Liu, Zhengliang, et al.
Published: (2023)
by: Liu, Zhengliang, et al.
Published: (2023)
LLMs left, right, and center: Assessing GPT's capabilities to label political bias from web domains
by: Hernandes, Raphael, et al.
Published: (2024)
by: Hernandes, Raphael, et al.
Published: (2024)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
by: Kim, Seungju, et al.
Published: (2024)
by: Kim, Seungju, et al.
Published: (2024)
Enhancing Programming Education with ChatGPT: A Case Study on Student Perceptions and Interactions in a Python Course
by: Ma, Boxaun, et al.
Published: (2024)
by: Ma, Boxaun, et al.
Published: (2024)
Math anxiety and associative knowledge structure are entwined in psychology students but not in Large Language Models like GPT-3.5 and GPT-4o
by: Ciringione, Luciana, et al.
Published: (2025)
by: Ciringione, Luciana, et al.
Published: (2025)
Leveraging Large Language Models for Actionable Course Evaluation Student Feedback to Lecturers
by: Zhang, Mike, et al.
Published: (2024)
by: Zhang, Mike, et al.
Published: (2024)
Exploring the change in scientific readability following the release of ChatGPT
by: Alsudais, Abdulkareem
Published: (2025)
by: Alsudais, Abdulkareem
Published: (2025)
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring
by: Armstrong, Lena, et al.
Published: (2024)
by: Armstrong, Lena, et al.
Published: (2024)
Handling Students Dropouts in an LLM-driven Interactive Online Course Using Language Models
by: Wang, Yuanchun, et al.
Published: (2025)
by: Wang, Yuanchun, et al.
Published: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Evaluating the Performance of ChatGPT for Spam Email Detection
by: Si, Shijing, et al.
Published: (2024)
by: Si, Shijing, et al.
Published: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
Better Call GPT, Comparing Large Language Models Against Lawyers
by: Martin, Lauren, et al.
Published: (2024)
by: Martin, Lauren, et al.
Published: (2024)
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant
by: Ballestero-Ribó, Marc, et al.
Published: (2025)
by: Ballestero-Ribó, Marc, et al.
Published: (2025)
RogueGPT: dis-ethical tuning transforms ChatGPT4 into a Rogue AI in 158 Words
by: Buscemi, Alessio, et al.
Published: (2024)
by: Buscemi, Alessio, et al.
Published: (2024)
Answering Students' Questions on Course Forums Using Multiple Chain-of-Thought Reasoning and Finetuning RAG-Enabled LLM
by: Wang, Neo, et al.
Published: (2025)
by: Wang, Neo, et al.
Published: (2025)
Politicians vs ChatGPT. A study of presuppositions in French and Italian political communication
by: Garassino, Davide, et al.
Published: (2024)
by: Garassino, Davide, et al.
Published: (2024)
GPT as ghostwriter at the White House
by: Savoy, Jacques
Published: (2024)
by: Savoy, Jacques
Published: (2024)
ChatEd: A Chatbot Leveraging ChatGPT for an Enhanced Learning Experience in Higher Education
by: Wang, Kevin, et al.
Published: (2023)
by: Wang, Kevin, et al.
Published: (2023)
GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?
by: Jin, Yiping, et al.
Published: (2024)
by: Jin, Yiping, et al.
Published: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice
by: Divya, V Sai, et al.
Published: (2026)
by: Divya, V Sai, et al.
Published: (2026)
Assessing the Impact of Conspiracy Theories Using Large Language Models
by: Jiang, Bohan, et al.
Published: (2024)
by: Jiang, Bohan, et al.
Published: (2024)
Towards Hybrid Intelligence in Journalism: Findings and Lessons Learnt from a Collaborative Analysis of Greek Political Rhetoric by ChatGPT and Humans
by: Troboukis, Thanasis, et al.
Published: (2024)
by: Troboukis, Thanasis, et al.
Published: (2024)
Phare: A Safety Probe for Large Language Models
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Similar Items
-
Artificial Intelligence Driven Course Generation: A Case Study Using ChatGPT
by: Rouabhia, Djaber
Published: (2024) -
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
by: Caraeni, Adriana, et al.
Published: (2024) -
Analyzing the Performance of ChatGPT in Cardiology and Vascular Pathologies
by: Hariri, Walid
Published: (2023) -
A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
by: Yeadon, Will, et al.
Published: (2024) -
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
by: Dong, Wenhan, et al.
Published: (2025)