Saved in:
Bibliographic Details
Main Authors: Maeda, Hotaka, Lu, Yikai
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.07017
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912681882025984
author Maeda, Hotaka
Lu, Yikai
author_facet Maeda, Hotaka
Lu, Yikai
contents We fine-tuned and compared several encoder-based Transformer large language models (LLM) to predict differential item functioning (DIF) from the item text. We then applied explainable artificial intelligence (XAI) methods to these models to identify specific words associated with DIF. The data included 42,180 items designed for English language arts and mathematics summative state assessments among students in grades 3 to 11. Prediction $R^2$ ranged from .04 to .32 among eight focal and reference group pairs. Our findings suggest that many words associated with DIF reflect minor sub-domains included in the test blueprint by design, rather than construct-irrelevant item content that should be removed from assessments. This may explain why qualitative reviews of DIF items often yield confusing or inconclusive results. Our approach can be used to screen words associated with DIF during the item-writing process for immediate revision, or help review traditional DIF analysis results by highlighting key words in the text. Extensions of this research can enhance the fairness of assessment programs, especially those that lack resources to build high-quality items, and among smaller subpopulations where we do not have sufficient sample sizes for traditional DIF analyses.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07017
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI
Maeda, Hotaka
Lu, Yikai
Computation and Language
Artificial Intelligence
We fine-tuned and compared several encoder-based Transformer large language models (LLM) to predict differential item functioning (DIF) from the item text. We then applied explainable artificial intelligence (XAI) methods to these models to identify specific words associated with DIF. The data included 42,180 items designed for English language arts and mathematics summative state assessments among students in grades 3 to 11. Prediction $R^2$ ranged from .04 to .32 among eight focal and reference group pairs. Our findings suggest that many words associated with DIF reflect minor sub-domains included in the test blueprint by design, rather than construct-irrelevant item content that should be removed from assessments. This may explain why qualitative reviews of DIF items often yield confusing or inconclusive results. Our approach can be used to screen words associated with DIF during the item-writing process for immediate revision, or help review traditional DIF analysis results by highlighting key words in the text. Extensions of this research can enhance the fairness of assessment programs, especially those that lack resources to build high-quality items, and among smaller subpopulations where we do not have sufficient sample sizes for traditional DIF analyses.
title Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.07017