Analyzing Large Language Models for Classroom Discussion Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tran, Nhat, Pierce, Benjamin, Litman, Diane, Correnti, Richard, Matsumura, Lindsay Clare
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929384661712896
author Tran, Nhat
Pierce, Benjamin
Litman, Diane
Correnti, Richard
Matsumura, Lindsay Clare
author_facet Tran, Nhat
Pierce, Benjamin
Litman, Diane
Correnti, Richard
Matsumura, Lindsay Clare
contents Automatically assessing classroom discussion quality is becoming increasingly feasible with the help of new NLP advancements such as large language models (LLMs). In this work, we examine how the assessment performance of 2 LLMs interacts with 3 factors that may affect performance: task formulation, context length, and few-shot examples. We also explore the computational efficiency and predictive consistency of the 2 LLMs. Our results suggest that the 3 aforementioned factors do affect the performance of the tested LLMs and there is a relation between consistency and performance. We recommend a LLM-based assessment approach that has a good balance in terms of predictive performance, computational efficiency, and consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08680
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Analyzing Large Language Models for Classroom Discussion Assessment
Tran, Nhat
Pierce, Benjamin
Litman, Diane
Correnti, Richard
Matsumura, Lindsay Clare
Computation and Language
Automatically assessing classroom discussion quality is becoming increasingly feasible with the help of new NLP advancements such as large language models (LLMs). In this work, we examine how the assessment performance of 2 LLMs interacts with 3 factors that may affect performance: task formulation, context length, and few-shot examples. We also explore the computational efficiency and predictive consistency of the 2 LLMs. Our results suggest that the 3 aforementioned factors do affect the performance of the tested LLMs and there is a relation between consistency and performance. We recommend a LLM-based assessment approach that has a good balance in terms of predictive performance, computational efficiency, and consistency.
title Analyzing Large Language Models for Classroom Discussion Assessment
topic Computation and Language
url https://arxiv.org/abs/2406.08680