Use of Fine-Tuned LLMs in Engineering Education: Evaluating Answer Quality using Metric-Based Scoring

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Siddhesh More, Saniya Nande
Format: Recurso digital
Published: Zenodo 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901209833537536
author Siddhesh More
Saniya Nande
author_facet Siddhesh More
Saniya Nande
contents Using a variety of large language models (LLMs)—such as Llama 3.1, Llama 3, Llama 2, Mistral and Phi 3.5—this study aims to provide chapter-by-chapter academic content for engineering courses. The objective is to assess these models by providing answers to exam questions from the past. Several quantitative metrics, such as ROUGE, Cosine Similarity, METEOR, Coherence, and Accuracy scores, are used to compare the generated replies with the original answers. Teaching academics evaluate the responses provided by the top three performing models manually, assigning a grade based on the academic quality and relevance of each response. The goal of this comparative study is to determine which model performs the best when it comes to algorithmic and human-based evaluation.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18137132
institution Zenodo
language
publishDate 2024
publisher Zenodo
record_format zenodo
spellingShingle Use of Fine-Tuned LLMs in Engineering Education: Evaluating Answer Quality using Metric-Based Scoring
Siddhesh More
Saniya Nande
Large language models
Llama 3.1
Llama 3
Llama 2
Mistral
Phi 3.5
engineering courses
chapter-wise content
past exam questions
ROUGE
Cosine Similarity
METEOR
Coherence
Accuracy score
teaching academics
human evaluation
academic quality
relevance
algorithmic performance
educational content delivery.
Using a variety of large language models (LLMs)—such as Llama 3.1, Llama 3, Llama 2, Mistral and Phi 3.5—this study aims to provide chapter-by-chapter academic content for engineering courses. The objective is to assess these models by providing answers to exam questions from the past. Several quantitative metrics, such as ROUGE, Cosine Similarity, METEOR, Coherence, and Accuracy scores, are used to compare the generated replies with the original answers. Teaching academics evaluate the responses provided by the top three performing models manually, assigning a grade based on the academic quality and relevance of each response. The goal of this comparative study is to determine which model performs the best when it comes to algorithmic and human-based evaluation.
title Use of Fine-Tuned LLMs in Engineering Education: Evaluating Answer Quality using Metric-Based Scoring
topic Large language models
Llama 3.1
Llama 3
Llama 2
Mistral
Phi 3.5
engineering courses
chapter-wise content
past exam questions
ROUGE
Cosine Similarity
METEOR
Coherence
Accuracy score
teaching academics
human evaluation
academic quality
relevance
algorithmic performance
educational content delivery.
url https://doi.org/10.5281/zenodo.18137132