Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yun, Xia, Xin, Wu, Xuansheng, Zhai, Xiaoming, Liu, Ninghao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911726327300096
author Wang, Yun
Xia, Xin
Wu, Xuansheng
Zhai, Xiaoming
Liu, Ninghao
author_facet Wang, Yun
Xia, Xin
Wu, Xuansheng
Zhai, Xiaoming
Liu, Ninghao
contents LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric construction. Human experts bypass this bottleneck through evaluation heuristics developed over extensive practice. We ask whether LLMs can learn similar heuristics directly from scoring experience, and formalize this as the concept of assessment skills: item-independent natural-language procedural knowledge that guides LLMs through specific stages of the scoring workflow. Focusing on rubric construction as a first instantiation, we propose an iterative framework that decomposes a skill into a fixed scaffold and learnable item-agnostic rules, refining the rules through LLM-driven diagnosis of scoring errors and validation-gated selection. The framework requires no expert-written rubric. On all ten ASAP-SAS items, optimized skills substantially improve LLM-based scoring and frequently surpass the dataset-provided expert rubric. Cross-item transfer experiments further reveal that learned skills capture both generalizable and item-specific patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29274
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
Wang, Yun
Xia, Xin
Wu, Xuansheng
Zhai, Xiaoming
Liu, Ninghao
Computation and Language
LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric construction. Human experts bypass this bottleneck through evaluation heuristics developed over extensive practice. We ask whether LLMs can learn similar heuristics directly from scoring experience, and formalize this as the concept of assessment skills: item-independent natural-language procedural knowledge that guides LLMs through specific stages of the scoring workflow. Focusing on rubric construction as a first instantiation, we propose an iterative framework that decomposes a skill into a fixed scaffold and learnable item-agnostic rules, refining the rules through LLM-driven diagnosis of scoring errors and validation-gated selection. The framework requires no expert-written rubric. On all ten ASAP-SAS items, optimized skills substantially improve LLM-based scoring and frequently surpass the dataset-provided expert rubric. Cross-item transfer experiments further reveal that learned skills capture both generalizable and item-specific patterns.
title Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
topic Computation and Language
url https://arxiv.org/abs/2605.29274