Analyzing Large Language Models for Classroom Discussion Assessment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929384661712896 |
|---|---|
| author | Tran, Nhat Pierce, Benjamin Litman, Diane Correnti, Richard Matsumura, Lindsay Clare |
| author_facet | Tran, Nhat Pierce, Benjamin Litman, Diane Correnti, Richard Matsumura, Lindsay Clare |
| contents | Automatically assessing classroom discussion quality is becoming increasingly feasible with the help of new NLP advancements such as large language models (LLMs). In this work, we examine how the assessment performance of 2 LLMs interacts with 3 factors that may affect performance: task formulation, context length, and few-shot examples. We also explore the computational efficiency and predictive consistency of the 2 LLMs. Our results suggest that the 3 aforementioned factors do affect the performance of the tested LLMs and there is a relation between consistency and performance. We recommend a LLM-based assessment approach that has a good balance in terms of predictive performance, computational efficiency, and consistency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_08680 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Analyzing Large Language Models for Classroom Discussion Assessment Tran, Nhat Pierce, Benjamin Litman, Diane Correnti, Richard Matsumura, Lindsay Clare Computation and Language Automatically assessing classroom discussion quality is becoming increasingly feasible with the help of new NLP advancements such as large language models (LLMs). In this work, we examine how the assessment performance of 2 LLMs interacts with 3 factors that may affect performance: task formulation, context length, and few-shot examples. We also explore the computational efficiency and predictive consistency of the 2 LLMs. Our results suggest that the 3 aforementioned factors do affect the performance of the tested LLMs and there is a relation between consistency and performance. We recommend a LLM-based assessment approach that has a good balance in terms of predictive performance, computational efficiency, and consistency. |
| title | Analyzing Large Language Models for Classroom Discussion Assessment |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2406.08680 |