A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohn, Clayton, Hutchins, Nicole, Le, Tuan, Biswas, Gautam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910429889953792
author Cohn, Clayton
Hutchins, Nicole
Le, Tuan
Biswas, Gautam
author_facet Cohn, Clayton
Hutchins, Nicole
Le, Tuan
Biswas, Gautam
contents This paper explores the use of large language models (LLMs) to score and explain short-answer assessments in K-12 science. While existing methods can score more structured math and computer science assessments, they often do not provide explanations for the scores. Our study focuses on employing GPT-4 for automated assessment in middle school Earth Science, combining few-shot and active learning with chain-of-thought reasoning. Using a human-in-the-loop approach, we successfully score and provide meaningful explanations for formative assessment responses. A systematic analysis of our method's pros and cons sheds light on the potential for human-in-the-loop techniques to enhance automated grading for open-ended science assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14565
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science
Cohn, Clayton
Hutchins, Nicole
Le, Tuan
Biswas, Gautam
Computation and Language
This paper explores the use of large language models (LLMs) to score and explain short-answer assessments in K-12 science. While existing methods can score more structured math and computer science assessments, they often do not provide explanations for the scores. Our study focuses on employing GPT-4 for automated assessment in middle school Earth Science, combining few-shot and active learning with chain-of-thought reasoning. Using a human-in-the-loop approach, we successfully score and provide meaningful explanations for formative assessment responses. A systematic analysis of our method's pros and cons sheds light on the potential for human-in-the-loop techniques to enhance automated grading for open-ended science assessments.
title A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science
topic Computation and Language
url https://arxiv.org/abs/2403.14565