Saved in:
Bibliographic Details
Main Authors: Molina, Ivo Lodovico, Švábenský, Valdemar, Minematsu, Tsubasa, Chen, Li, Okubo, Fumiya, Shimada, Atsushi
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.20578
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912028137881600
author Molina, Ivo Lodovico
Švábenský, Valdemar
Minematsu, Tsubasa
Chen, Li
Okubo, Fumiya
Shimada, Atsushi
author_facet Molina, Ivo Lodovico
Švábenský, Valdemar
Minematsu, Tsubasa
Chen, Li
Okubo, Fumiya
Shimada, Atsushi
contents This study explores the effectiveness of Large Language Models (LLMs) for Automatic Question Generation in educational settings. Three LLMs are compared in their ability to create questions from university slide text without fine-tuning. Questions were obtained in a two-step pipeline: first, answer phrases were extracted from slides using Llama 2-Chat 13B; then, the three models generated questions for each answer. To analyze whether the questions would be suitable in educational applications for students, a survey was conducted with 46 students who evaluated a total of 246 questions across five metrics: clarity, relevance, difficulty, slide relation, and question-answer alignment. Results indicate that GPT-3.5 and Llama 2-Chat 13B outperform Flan T5 XXL by a small margin, particularly in terms of clarity and question-answer alignment. GPT-3.5 especially excels at tailoring questions to match the input answers. The contribution of this research is the analysis of the capacity of LLMs for Automatic Question Generation in education.
format Preprint
id arxiv_https___arxiv_org_abs_2407_20578
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comparison of Large Language Models for Generating Contextually Relevant Questions
Molina, Ivo Lodovico
Švábenský, Valdemar
Minematsu, Tsubasa
Chen, Li
Okubo, Fumiya
Shimada, Atsushi
Computation and Language
Artificial Intelligence
Computers and Society
K.3
This study explores the effectiveness of Large Language Models (LLMs) for Automatic Question Generation in educational settings. Three LLMs are compared in their ability to create questions from university slide text without fine-tuning. Questions were obtained in a two-step pipeline: first, answer phrases were extracted from slides using Llama 2-Chat 13B; then, the three models generated questions for each answer. To analyze whether the questions would be suitable in educational applications for students, a survey was conducted with 46 students who evaluated a total of 246 questions across five metrics: clarity, relevance, difficulty, slide relation, and question-answer alignment. Results indicate that GPT-3.5 and Llama 2-Chat 13B outperform Flan T5 XXL by a small margin, particularly in terms of clarity and question-answer alignment. GPT-3.5 especially excels at tailoring questions to match the input answers. The contribution of this research is the analysis of the capacity of LLMs for Automatic Question Generation in education.
title Comparison of Large Language Models for Generating Contextually Relevant Questions
topic Computation and Language
Artificial Intelligence
Computers and Society
K.3
url https://arxiv.org/abs/2407.20578