DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tianyou, Chen, Xinglu, Zhang, Jingshen, Qiu, Xinying, Niu, Ruiying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916845387251712
author Huang, Tianyou
Chen, Xinglu
Zhang, Jingshen
Qiu, Xinying
Niu, Ruiying
author_facet Huang, Tianyou
Chen, Xinglu
Zhang, Jingshen
Qiu, Xinying
Niu, Ruiying
contents This paper introduces DualReward, a novel reinforcement learning framework for automatic distractor generation in cloze tests. Unlike conventional approaches that rely primarily on supervised learning or static generative models, our method employs a dual reward structure with adaptive scaling that differentiates between human-created gold standard distractors and model-generated candidates. The framework dynamically adjusts reward signal intensity based on model performance and confidence. We evaluate our approach on both passage-level (CLOTH-F) and sentence-level (MCQ) cloze test datasets, demonstrating consistent improvements over state-of-the-art baselines. Experimental results show that our adaptive reward scaling mechanism provides modest but consistent benefits on homogeneous datasets (CLOTH-F) and more substantial improvements (3.48-3.86% in P@1) on diverse, cross-domain data (MCQ), suggesting its particular effectiveness for handling varied question types and domains. Our work offers a flexible framework that effectively balances learning from reliable human examples while exploring novel, high-quality distractors for automated test generation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation
Huang, Tianyou
Chen, Xinglu
Zhang, Jingshen
Qiu, Xinying
Niu, Ruiying
Computation and Language
This paper introduces DualReward, a novel reinforcement learning framework for automatic distractor generation in cloze tests. Unlike conventional approaches that rely primarily on supervised learning or static generative models, our method employs a dual reward structure with adaptive scaling that differentiates between human-created gold standard distractors and model-generated candidates. The framework dynamically adjusts reward signal intensity based on model performance and confidence. We evaluate our approach on both passage-level (CLOTH-F) and sentence-level (MCQ) cloze test datasets, demonstrating consistent improvements over state-of-the-art baselines. Experimental results show that our adaptive reward scaling mechanism provides modest but consistent benefits on homogeneous datasets (CLOTH-F) and more substantial improvements (3.48-3.86% in P@1) on diverse, cross-domain data (MCQ), suggesting its particular effectiveness for handling varied question types and domains. Our work offers a flexible framework that effectively balances learning from reliable human examples while exploring novel, high-quality distractors for automated test generation.
title DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation
topic Computation and Language
url https://arxiv.org/abs/2507.11875