Can LLMs Identify Gaps and Misconceptions in Students' Code Explanations?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oli, Priti, Banjade, Rabin, Olney, Andrew M., Rus, Vasile
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909460302135296
author Oli, Priti
Banjade, Rabin
Olney, Andrew M.
Rus, Vasile
author_facet Oli, Priti
Banjade, Rabin
Olney, Andrew M.
Rus, Vasile
contents This paper investigates various approaches using Large Language Models (LLMs) to identify gaps and misconceptions in students' self-explanations of specific instructional material, in our case explanations of code examples. This research is a part of our larger effort to automate the assessment of students' freely generated responses, focusing specifically on their self-explanations of code examples during activities related to code comprehension. In this work, we experiment with zero-shot prompting, Supervised Fine-Tuning (SFT), and preference alignment of LLMs to identify gaps in students' self-explanation. With simple prompting, GPT-4 consistently outperformed LLaMA3 and Mistral in identifying gaps and misconceptions, as confirmed by human evaluations. Additionally, our results suggest that fine-tuned large language models are more effective at identifying gaps in students' explanations compared to zero-shot and few-shot prompting techniques. Furthermore, our findings show that the preference optimization approach using Odds Ratio Preference Optimization (ORPO) outperforms SFT in identifying gaps and misconceptions in students' code explanations.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can LLMs Identify Gaps and Misconceptions in Students' Code Explanations?
Oli, Priti
Banjade, Rabin
Olney, Andrew M.
Rus, Vasile
Computers and Society
Artificial Intelligence
Software Engineering
This paper investigates various approaches using Large Language Models (LLMs) to identify gaps and misconceptions in students' self-explanations of specific instructional material, in our case explanations of code examples. This research is a part of our larger effort to automate the assessment of students' freely generated responses, focusing specifically on their self-explanations of code examples during activities related to code comprehension. In this work, we experiment with zero-shot prompting, Supervised Fine-Tuning (SFT), and preference alignment of LLMs to identify gaps in students' self-explanation. With simple prompting, GPT-4 consistently outperformed LLaMA3 and Mistral in identifying gaps and misconceptions, as confirmed by human evaluations. Additionally, our results suggest that fine-tuned large language models are more effective at identifying gaps in students' explanations compared to zero-shot and few-shot prompting techniques. Furthermore, our findings show that the preference optimization approach using Odds Ratio Preference Optimization (ORPO) outperforms SFT in identifying gaps and misconceptions in students' code explanations.
title Can LLMs Identify Gaps and Misconceptions in Students' Code Explanations?
topic Computers and Society
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2501.10365