EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Naeem, Numaan, Mekki, Abdellah El, Abdul-Mageed, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025)
by: Ahmad, Sarfraz, et al.
Published: (2025)
Graph Guided Question Answer Generation for Procedural Question-Answering
by: Pham, Hai X., et al.
Published: (2024)
by: Pham, Hai X., et al.
Published: (2024)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
by: Yuan, Weikang, et al.
Published: (2025)
by: Yuan, Weikang, et al.
Published: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
by: Palit, Sayon, et al.
Published: (2025)
by: Palit, Sayon, et al.
Published: (2025)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
by: Muller, Sacha, et al.
Published: (2024)
by: Muller, Sacha, et al.
Published: (2024)
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
by: Er, Yakup Abrek, et al.
Published: (2025)
by: Er, Yakup Abrek, et al.
Published: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
by: Basit, Abdul, et al.
Published: (2024)
by: Basit, Abdul, et al.
Published: (2024)
Subjective Question Generation and Answer Evaluation using NLP
by: Islam, G. M. Refatul, et al.
Published: (2025)
by: Islam, G. M. Refatul, et al.
Published: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
by: Mehenni, Gaya, et al.
Published: (2025)
by: Mehenni, Gaya, et al.
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
An Adaptive Framework for Generating Systematic Explanatory Answer in Online Q&A Platforms
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
by: Eyasir, Abrar, et al.
Published: (2026)
by: Eyasir, Abrar, et al.
Published: (2026)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
by: Simione II, Robert L
Published: (2024)
by: Simione II, Robert L
Published: (2024)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
by: Iqbal, Hasan, et al.
Published: (2024)
by: Iqbal, Hasan, et al.
Published: (2024)
PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
by: Mukhopadhyay, Srija, et al.
Published: (2025)
by: Mukhopadhyay, Srija, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Constructing Cloze Questions Generatively
by: Sun, Yicheng, et al.
Published: (2024)
by: Sun, Yicheng, et al.
Published: (2024)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
by: Dayarathne, Ranul, et al.
Published: (2025)
by: Dayarathne, Ranul, et al.
Published: (2025)
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer
by: Sun, Huashan, et al.
Published: (2023)
by: Sun, Huashan, et al.
Published: (2023)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
DQA: Diagnostic Question Answering for IT Support
by: Kapoor, Vishaal, et al.
Published: (2026)
by: Kapoor, Vishaal, et al.
Published: (2026)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
by: Fiaz, Layba, et al.
Published: (2025)
by: Fiaz, Layba, et al.
Published: (2025)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
by: Vieira, Inês, et al.
Published: (2026)
by: Vieira, Inês, et al.
Published: (2026)
A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
by: Park, Chiwan, et al.
Published: (2025)
by: Park, Chiwan, et al.
Published: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Similar Items
-
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
by: Naeem, Numaan, et al.
Published: (2025) -
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025) -
Graph Guided Question Answer Generation for Procedural Question-Answering
by: Pham, Hai X., et al.
Published: (2024) -
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025) -
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)