Small, Private Language Models as Teammates for Educational Assessment Design

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jaldi, Chris Davis, Saini, Anmol, Zhang, Shan, Schroeder, Noah, Shimizu, Cogan, Ilkou, Eleni
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911685872189440
author Jaldi, Chris Davis
Saini, Anmol
Zhang, Shan
Schroeder, Noah
Shimizu, Cogan
Ilkou, Eleni
author_facet Jaldi, Chris Davis
Saini, Anmol
Zhang, Shan
Schroeder, Noah
Shimizu, Cogan
Ilkou, Eleni
contents Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bloom's taxonomy). However, they often rely on subjective or limited evaluation methods; focus primarily on proprietary models; or rarely systematically examine generation, evaluation, or deployment constraints in real educational settings. Meanwhile, Small Language Models (SLMs) have emerged as local alternatives that better address privacy and resource limitations; yet their effectiveness for assessment tasks remains underexplored. To address this gap, we systematically compare LLMs and SLMs for assessment question design; evaluate generation quality across Bloom's taxonomy levels using reproducible, pedagogically grounded metrics; and further assess model-based judging against expert-informed evaluation by analyzing reliability and agreement patterns. Results show that SLMs achieve competitive performance across key pedagogically motivated quality dimensions while enabling local, privacy-sensitive deployment. However, model-based evaluations also exhibit systematic inconsistencies and bias relative to expert ratings. These findings provide evidence to posit language models as bounded assistants in assessment workflows; underscore the necessity of Human-in-the-Loop; and advance the automated educational question generation field by examining quality, reliability, and deployment-aware trade-offs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15015
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Small, Private Language Models as Teammates for Educational Assessment Design
Jaldi, Chris Davis
Saini, Anmol
Zhang, Shan
Schroeder, Noah
Shimizu, Cogan
Ilkou, Eleni
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Generative AI increasingly supports educational design tasks, e.g., through Large Language Models (LLMs), demonstrating the capability to design assessment questions that are aligned with pedagogical frameworks (e.g., Bloom's taxonomy). However, they often rely on subjective or limited evaluation methods; focus primarily on proprietary models; or rarely systematically examine generation, evaluation, or deployment constraints in real educational settings. Meanwhile, Small Language Models (SLMs) have emerged as local alternatives that better address privacy and resource limitations; yet their effectiveness for assessment tasks remains underexplored. To address this gap, we systematically compare LLMs and SLMs for assessment question design; evaluate generation quality across Bloom's taxonomy levels using reproducible, pedagogically grounded metrics; and further assess model-based judging against expert-informed evaluation by analyzing reliability and agreement patterns. Results show that SLMs achieve competitive performance across key pedagogically motivated quality dimensions while enabling local, privacy-sensitive deployment. However, model-based evaluations also exhibit systematic inconsistencies and bias relative to expert ratings. These findings provide evidence to posit language models as bounded assistants in assessment workflows; underscore the necessity of Human-in-the-Loop; and advance the automated educational question generation field by examining quality, reliability, and deployment-aware trade-offs.
title Small, Private Language Models as Teammates for Educational Assessment Design
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2605.15015