Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large
Fuente:
arXiv
Saved in:
| Main Authors: | Jauhiainen, Jussi S., Guerra, Agustín Garagorry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPT-3.5 for Grammatical Error Correction
by: Katinskaia, Anisia, et al.
Published: (2024)
by: Katinskaia, Anisia, et al.
Published: (2024)
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2024)
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2024)
How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses
by: Lin, Jionghao, et al.
Published: (2024)
by: Lin, Jionghao, et al.
Published: (2024)
ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models
by: Wen, Zehao, et al.
Published: (2024)
by: Wen, Zehao, et al.
Published: (2024)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo
by: Wood, Michael C., et al.
Published: (2024)
by: Wood, Michael C., et al.
Published: (2024)
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2023)
Exploring Combinatorial Problem Solving with Large Language Models: A Case Study on the Travelling Salesman Problem Using GPT-3.5 Turbo
by: Masoud, Mahmoud, et al.
Published: (2024)
by: Masoud, Mahmoud, et al.
Published: (2024)
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
by: Shakil, Hassan, et al.
Published: (2024)
by: Shakil, Hassan, et al.
Published: (2024)
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT
by: Krauss, Patrick, et al.
Published: (2024)
by: Krauss, Patrick, et al.
Published: (2024)
OpenAI GPT-5 System Card
by: Singh, Aaditya, et al.
Published: (2025)
by: Singh, Aaditya, et al.
Published: (2025)
Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
by: Hackl, Veronika, et al.
Published: (2023)
by: Hackl, Veronika, et al.
Published: (2023)
Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics
by: Greco, Candida M., et al.
Published: (2024)
by: Greco, Candida M., et al.
Published: (2024)
Enhancing Next-Generation Language Models with Knowledge Graphs: Extending Claude, Mistral IA, and GPT-4 via KG-BERT
by: Chaabene, Nour El Houda Ben, et al.
Published: (2025)
by: Chaabene, Nour El Houda Ben, et al.
Published: (2025)
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions
by: Rouzegar, Hamdireza, et al.
Published: (2024)
by: Rouzegar, Hamdireza, et al.
Published: (2024)
Is GPT-4 Less Politically Biased than GPT-3.5? A Renewed Investigation of ChatGPT's Political Biases
by: Weber, Erik, et al.
Published: (2024)
by: Weber, Erik, et al.
Published: (2024)
Whose LLM is it Anyway? Linguistic Comparison and LLM Attribution for GPT-3.5, GPT-4 and Bard
by: Rosenfeld, Ariel, et al.
Published: (2024)
by: Rosenfeld, Ariel, et al.
Published: (2024)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
DaVinci at SemEval-2024 Task 9: Few-shot prompting GPT-3.5 for Unconventional Reasoning
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
by: Lachenmaier, Clara, et al.
Published: (2026)
by: Lachenmaier, Clara, et al.
Published: (2026)
Generative Artificial Intelligence and Agents in Research and Teaching
by: Jauhiainen, Jussi S., et al.
Published: (2025)
by: Jauhiainen, Jussi S., et al.
Published: (2025)
HateGPT: Unleashing GPT-3.5 Turbo to Combat Hate Speech on X
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
by: Plevris, Vagelis, et al.
Published: (2023)
by: Plevris, Vagelis, et al.
Published: (2023)
A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course
by: Yeadon, Will, et al.
Published: (2024)
by: Yeadon, Will, et al.
Published: (2024)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
by: Mao, Rui, et al.
Published: (2023)
by: Mao, Rui, et al.
Published: (2023)
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
by: Sassella, Andrea, et al.
Published: (2026)
by: Sassella, Andrea, et al.
Published: (2026)
Multilevel Analysis of Cryptocurrency News using RAG Approach with Fine-Tuned Mistral Large Language Model
by: Pavlyshenko, Bohdan M.
Published: (2025)
by: Pavlyshenko, Bohdan M.
Published: (2025)
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
by: Bayani, David
Published: (2023)
by: Bayani, David
Published: (2023)
An Evaluation of GPT-4V for Transcribing the Urban Renewal Hand-Written Collection
by: Lee, Myeong, et al.
Published: (2024)
by: Lee, Myeong, et al.
Published: (2024)
GPT-4 Technical Report
by: OpenAI, et al.
Published: (2023)
by: OpenAI, et al.
Published: (2023)
An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
by: Sands, Brendan, et al.
Published: (2025)
by: Sands, Brendan, et al.
Published: (2025)
Can GPT-3.5 Generate and Code Discharge Summaries?
by: Falis, Matúš, et al.
Published: (2024)
by: Falis, Matúš, et al.
Published: (2024)
ArabianGPT: Native Arabic GPT-based Large Language Model
by: Koubaa, Anis, et al.
Published: (2024)
by: Koubaa, Anis, et al.
Published: (2024)
Math anxiety and associative knowledge structure are entwined in psychology students but not in Large Language Models like GPT-3.5 and GPT-4o
by: Ciringione, Luciana, et al.
Published: (2025)
by: Ciringione, Luciana, et al.
Published: (2025)
EngGPT2: Sovereign, Efficient and Open Intelligence
by: Ciarfaglia, G., et al.
Published: (2026)
by: Ciarfaglia, G., et al.
Published: (2026)
Comparative Analysis of ChatGPT, GPT-4, and Microsoft Bing Chatbots for GRE Test
by: Abu-Haifa, Mohammad, et al.
Published: (2023)
by: Abu-Haifa, Mohammad, et al.
Published: (2023)
HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis
by: Vasilatos, Christoforos, et al.
Published: (2023)
by: Vasilatos, Christoforos, et al.
Published: (2023)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
by: Wang, Ming, et al.
Published: (2024)
by: Wang, Ming, et al.
Published: (2024)
Using Hallucinations to Bypass GPT4's Filter
by: Lemkin, Benjamin
Published: (2024)
by: Lemkin, Benjamin
Published: (2024)
Similar Items
-
GPT-3.5 for Grammatical Error Correction
by: Katinskaia, Anisia, et al.
Published: (2024) -
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
by: Safavi-Naini, Seyed Amir Ahmad, et al.
Published: (2024) -
How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses
by: Lin, Jionghao, et al.
Published: (2024) -
ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models
by: Wen, Zehao, et al.
Published: (2024) -
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
by: Koutcheme, Charles, et al.
Published: (2024)