Just Read the Question: Enabling Generalization to New Assessment Items with Text Awareness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khan, Arisha, Li, Nathaniel, Shen, Tori, Rafferty, Anna N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915382735929344
author Khan, Arisha
Li, Nathaniel
Shen, Tori
Rafferty, Anna N.
author_facet Khan, Arisha
Li, Nathaniel
Shen, Tori
Rafferty, Anna N.
contents Machine learning has been proposed as a way to improve educational assessment by making fine-grained predictions about student performance and learning relationships between items. One challenge with many machine learning approaches is incorporating new items, as these approaches rely heavily on historical data. We develop Text-LENS by extending the LENS partial variational auto-encoder for educational assessment to leverage item text embeddings, and explore the impact on predictive performance and generalization to previously unseen items. We examine performance on two datasets: Eedi, a publicly available dataset that includes item content, and LLM-Sim, a novel dataset with test items produced by an LLM. We find that Text-LENS matches LENS' performance on seen items and improves upon it in a variety of conditions involving unseen items; it effectively learns student proficiency from and makes predictions about student performance on new items.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08154
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Just Read the Question: Enabling Generalization to New Assessment Items with Text Awareness
Khan, Arisha
Li, Nathaniel
Shen, Tori
Rafferty, Anna N.
Machine Learning
Machine learning has been proposed as a way to improve educational assessment by making fine-grained predictions about student performance and learning relationships between items. One challenge with many machine learning approaches is incorporating new items, as these approaches rely heavily on historical data. We develop Text-LENS by extending the LENS partial variational auto-encoder for educational assessment to leverage item text embeddings, and explore the impact on predictive performance and generalization to previously unseen items. We examine performance on two datasets: Eedi, a publicly available dataset that includes item content, and LLM-Sim, a novel dataset with test items produced by an LLM. We find that Text-LENS matches LENS' performance on seen items and improves upon it in a variety of conditions involving unseen items; it effectively learns student proficiency from and makes predictions about student performance on new items.
title Just Read the Question: Enabling Generalization to New Assessment Items with Text Awareness
topic Machine Learning
url https://arxiv.org/abs/2507.08154