Open, Small, Rigmarole -- Evaluating Llama 3.2 3B's Feedback for Programming Exercises

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Azaiz, Imen, Kiesler, Natalie, Strickroth, Sven, Zhang, Anni
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916670955585536
author Azaiz, Imen
Kiesler, Natalie
Strickroth, Sven
Zhang, Anni
author_facet Azaiz, Imen
Kiesler, Natalie
Strickroth, Sven
Zhang, Anni
contents Large Language Models (LLMs) have been subject to extensive research in the past few years. This is particularly true for the potential of LLMs to generate formative programming feedback for novice learners at university. In contrast to Generative AI (GenAI) tools based on LLMs, such as GPT, smaller and open models have received much less attention. Yet, they offer several benefits, as educators can let them run on a virtual machine or personal computer. This can help circumvent some major concerns applicable to other GenAI tools and LLMs (e. g., data protection, lack of control over changes, privacy). Therefore, this study explores the feedback characteristics of the open, lightweight LLM Llama 3.2 (3B). In particular, we investigate the models' responses to authentic student solutions to introductory programming exercises written in Java. The generated output is qualitatively analyzed to help evaluate the feedback's quality, content, structure, and other features. The results provide a comprehensive overview of the feedback capabilities and serious shortcomings of this open, small LLM. We further discuss the findings in the context of previous research on LLMs and contribute to benchmarking recently available GenAI tools and their feedback for novice learners of programming. Thereby, this work has implications for educators, learners, and tool developers attempting to utilize all variants of LLMs (including open, and small models) to generate formative feedback and support learning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_01054
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Open, Small, Rigmarole -- Evaluating Llama 3.2 3B's Feedback for Programming Exercises
Azaiz, Imen
Kiesler, Natalie
Strickroth, Sven
Zhang, Anni
Computers and Society
Software Engineering
Large Language Models (LLMs) have been subject to extensive research in the past few years. This is particularly true for the potential of LLMs to generate formative programming feedback for novice learners at university. In contrast to Generative AI (GenAI) tools based on LLMs, such as GPT, smaller and open models have received much less attention. Yet, they offer several benefits, as educators can let them run on a virtual machine or personal computer. This can help circumvent some major concerns applicable to other GenAI tools and LLMs (e. g., data protection, lack of control over changes, privacy). Therefore, this study explores the feedback characteristics of the open, lightweight LLM Llama 3.2 (3B). In particular, we investigate the models' responses to authentic student solutions to introductory programming exercises written in Java. The generated output is qualitatively analyzed to help evaluate the feedback's quality, content, structure, and other features. The results provide a comprehensive overview of the feedback capabilities and serious shortcomings of this open, small LLM. We further discuss the findings in the context of previous research on LLMs and contribute to benchmarking recently available GenAI tools and their feedback for novice learners of programming. Thereby, this work has implications for educators, learners, and tool developers attempting to utilize all variants of LLMs (including open, and small models) to generate formative feedback and support learning.
title Open, Small, Rigmarole -- Evaluating Llama 3.2 3B's Feedback for Programming Exercises
topic Computers and Society
Software Engineering
url https://arxiv.org/abs/2504.01054