PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Qikai, Zhang, Zhenrong, Chen, Linbo, Hu, Pengfei, Zhang, Jianshu, Guo, Youhui, Du, Jun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910269191487488
author Chang, Qikai
Zhang, Zhenrong
Chen, Linbo
Hu, Pengfei
Zhang, Jianshu
Guo, Youhui
Du, Jun
author_facet Chang, Qikai
Zhang, Zhenrong
Chen, Linbo
Hu, Pengfei
Zhang, Jianshu
Guo, Youhui
Du, Jun
contents Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide progressive Socratic guidance and balance multiple pedagogical objectives across multi-turn interactions. However, training such tutors remains challenging due to limited-fidelity and weakly controllable student simulation, under-specified pedagogical reward modeling, and unstable multi-objective optimization. To overcome these limitations, we propose PEARL, a pedagogically aligned reinforcement learning framework for training Socratic tutoring agents, consisting of three key components. First, we introduce a controllable student simulator that decouples latent cognitive states from response generation to model diverse abilities and misconceptions. Second, we develop a generative reward model that jointly evaluates pedagogical quality and objective correctness for policy optimization. Finally, we propose a stable multi-objective RL scheme that discretizes rewards within each dimension and aggregates normalized advantages across dimensions, preventing high-variance objectives from dominating updates. Experiments on multiple benchmarks show that PEARL achieves the best performance among open-source models and remains competitive with leading proprietary LLMs, despite using only a 30B policy model.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29582
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
Chang, Qikai
Zhang, Zhenrong
Chen, Linbo
Hu, Pengfei
Zhang, Jianshu
Guo, Youhui
Du, Jun
Machine Learning
Computation and Language
Large Language Models (LLMs) have shown promise as educational tutors, yet effective tutoring requires more than solving problems: it must provide progressive Socratic guidance and balance multiple pedagogical objectives across multi-turn interactions. However, training such tutors remains challenging due to limited-fidelity and weakly controllable student simulation, under-specified pedagogical reward modeling, and unstable multi-objective optimization. To overcome these limitations, we propose PEARL, a pedagogically aligned reinforcement learning framework for training Socratic tutoring agents, consisting of three key components. First, we introduce a controllable student simulator that decouples latent cognitive states from response generation to model diverse abilities and misconceptions. Second, we develop a generative reward model that jointly evaluates pedagogical quality and objective correctness for policy optimization. Finally, we propose a stable multi-objective RL scheme that discretizes rewards within each dimension and aggregates normalized advantages across dimensions, preventing high-variance objectives from dominating updates. Experiments on multiple benchmarks show that PEARL achieves the best performance among open-source models and remains competitive with leading proprietary LLMs, despite using only a 30B policy model.
title PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.29582