KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Zhangqi, Fernandez, Nigel, Lan, Andrew
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917504029294592
author Duan, Zhangqi
Fernandez, Nigel
Lan, Andrew
author_facet Duan, Zhangqi
Fernandez, Nigel
Lan, Andrew
contents Open-ended tasks, such as coding problems that are common in computer science education, provide detailed insights into student knowledge. However, training large language models (LLMs) to simulate and predict possible student errors in their responses to these problems can be challenging: they often suffer from mode collapse and fail to fully capture the diversity in syntax, style, and solution approach in student responses. In this work, we present KASER (Knowledge-Aligned Student Error Simulator), a novel approach that aligns errors with student knowledge. We propose a training method based on reinforcement learning using a hybrid reward that reflects three aspects of student code prediction: i) code similarity to the ground-truth, ii) error matching, and iii) code prediction diversity. On two real-world datasets, we perform two levels of evaluation and show that: At the per-student-problem pair level, our method outperforms baselines on code and error prediction; at the per-problem level, our method outperforms baselines on error coverage and simulated code diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06633
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
Duan, Zhangqi
Fernandez, Nigel
Lan, Andrew
Machine Learning
Artificial Intelligence
Computation and Language
Computers and Society
Open-ended tasks, such as coding problems that are common in computer science education, provide detailed insights into student knowledge. However, training large language models (LLMs) to simulate and predict possible student errors in their responses to these problems can be challenging: they often suffer from mode collapse and fail to fully capture the diversity in syntax, style, and solution approach in student responses. In this work, we present KASER (Knowledge-Aligned Student Error Simulator), a novel approach that aligns errors with student knowledge. We propose a training method based on reinforcement learning using a hybrid reward that reflects three aspects of student code prediction: i) code similarity to the ground-truth, ii) error matching, and iii) code prediction diversity. On two real-world datasets, we perform two levels of evaluation and show that: At the per-student-problem pair level, our method outperforms baselines on code and error prediction; at the per-problem level, our method outperforms baselines on error coverage and simulated code diversity.
title KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
topic Machine Learning
Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2601.06633