Evaluating Large Language Models for automatic analysis of teacher simulations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de-Fitero-Dominguez, David, Albaladejo-González, Mariano, Garcia-Cabot, Antonio, Garcia-Lopez, Eva, Moreno-Cediel, Antonio, Barno, Erin, Reich, Justin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929441467269120
author de-Fitero-Dominguez, David
Albaladejo-González, Mariano
Garcia-Cabot, Antonio
Garcia-Lopez, Eva
Moreno-Cediel, Antonio
Barno, Erin
Reich, Justin
author_facet de-Fitero-Dominguez, David
Albaladejo-González, Mariano
Garcia-Cabot, Antonio
Garcia-Lopez, Eva
Moreno-Cediel, Antonio
Barno, Erin
Reich, Justin
contents Digital Simulations (DS) provide safe environments where users interact with an agent through conversational prompts, providing engaging learning experiences that can be used to train teacher candidates in realistic classroom scenarios. These simulations usually include open-ended questions, allowing teacher candidates to express their thoughts but complicating an automatic response analysis. To address this issue, we have evaluated Large Language Models (LLMs) to identify characteristics (user behaviors) in the responses of DS for teacher education. We evaluated the performance of DeBERTaV3 and Llama 3, combined with zero-shot, few-shot, and fine-tuning. Our experiments discovered a significant variation in the LLMs' performance depending on the characteristic to identify. Additionally, we noted that DeBERTaV3 significantly reduced its performance when it had to identify new characteristics. In contrast, Llama 3 performed better than DeBERTaV3 in detecting new characteristics and showing more stable performance. Therefore, in DS where teacher educators need to introduce new characteristics because they change depending on the simulation or the educational objectives, it is more recommended to use Llama 3. These results can guide other researchers in introducing LLMs to provide the highly demanded automatic evaluations in DS.
format Preprint
id arxiv_https___arxiv_org_abs_2407_20360
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Large Language Models for automatic analysis of teacher simulations
de-Fitero-Dominguez, David
Albaladejo-González, Mariano
Garcia-Cabot, Antonio
Garcia-Lopez, Eva
Moreno-Cediel, Antonio
Barno, Erin
Reich, Justin
Artificial Intelligence
Digital Simulations (DS) provide safe environments where users interact with an agent through conversational prompts, providing engaging learning experiences that can be used to train teacher candidates in realistic classroom scenarios. These simulations usually include open-ended questions, allowing teacher candidates to express their thoughts but complicating an automatic response analysis. To address this issue, we have evaluated Large Language Models (LLMs) to identify characteristics (user behaviors) in the responses of DS for teacher education. We evaluated the performance of DeBERTaV3 and Llama 3, combined with zero-shot, few-shot, and fine-tuning. Our experiments discovered a significant variation in the LLMs' performance depending on the characteristic to identify. Additionally, we noted that DeBERTaV3 significantly reduced its performance when it had to identify new characteristics. In contrast, Llama 3 performed better than DeBERTaV3 in detecting new characteristics and showing more stable performance. Therefore, in DS where teacher educators need to introduce new characteristics because they change depending on the simulation or the educational objectives, it is more recommended to use Llama 3. These results can guide other researchers in introducing LLMs to provide the highly demanded automatic evaluations in DS.
title Evaluating Large Language Models for automatic analysis of teacher simulations
topic Artificial Intelligence
url https://arxiv.org/abs/2407.20360