Long Context Automated Essay Scoring with Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ormerod, Christopher, Kehat, Gitit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912585007235072
author Ormerod, Christopher
Kehat, Gitit
author_facet Ormerod, Christopher
Kehat, Gitit
contents Transformer-based language models are architecturally constrained to process text of a fixed maximum length. Essays written by higher-grade students frequently exceed the maximum allowed length for many popular open-source models. A common approach to addressing this issue when using these models for Automated Essay Scoring is to truncate the input text. This raises serious validity concerns as it undermines the model's ability to fully capture and evaluate organizational elements of the scoring rubric, which requires long contexts to assess. In this study, we evaluate several models that incorporate architectural modifications of the standard transformer architecture to overcome these length limitations using the Kaggle ASAP 2.0 dataset. The models considered in this study include fine-tuned versions of XLNet, Longformer, ModernBERT, Mamba, and Llama models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long Context Automated Essay Scoring with Language Models
Ormerod, Christopher
Kehat, Gitit
Computation and Language
Transformer-based language models are architecturally constrained to process text of a fixed maximum length. Essays written by higher-grade students frequently exceed the maximum allowed length for many popular open-source models. A common approach to addressing this issue when using these models for Automated Essay Scoring is to truncate the input text. This raises serious validity concerns as it undermines the model's ability to fully capture and evaluate organizational elements of the scoring rubric, which requires long contexts to assess. In this study, we evaluate several models that incorporate architectural modifications of the standard transformer architecture to overcome these length limitations using the Kaggle ASAP 2.0 dataset. The models considered in this study include fine-tuned versions of XLNet, Longformer, ModernBERT, Mamba, and Llama models.
title Long Context Automated Essay Scoring with Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.10417