Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zimmer, Matthieu, Ji, Xiaotong, Nguyen, Tu, Ammar, Haitham Bou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915517243064320
author Zimmer, Matthieu
Ji, Xiaotong
Nguyen, Tu
Ammar, Haitham Bou
author_facet Zimmer, Matthieu
Ji, Xiaotong
Nguyen, Tu
Ammar, Haitham Bou
contents We introduce a novel approach to large language model (LLM) distillation by formulating it as a constrained reinforcement learning problem. While recent work has begun exploring the integration of task-specific rewards into distillation processes, existing methods typically rely on ad-hoc reward weighting. We propose a principled optimization framework that maximizes task-specific rewards while constraining the divergence from the teacher model to remain below a specified threshold. Our approach adapts constrained state augmented reinforcement learning to the distillation setting, introducing a modified reward function that maintains theoretical guarantees of constraint satisfaction without requiring state augmentation or teacher model access during deployment and without the computational overhead of the dual Lagrangian methods. Through extensive experiments on mathematical reasoning tasks, we demonstrate that our method achieves better constraint satisfaction rates and better reasoning compared to the soft Lagrangian relaxation baselines while maintaining competitive task performance. Our framework provides a theoretically grounded and practically efficient solution for reward-aware distillation in resource-constrained settings.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
Zimmer, Matthieu
Ji, Xiaotong
Nguyen, Tu
Ammar, Haitham Bou
Machine Learning
Artificial Intelligence
We introduce a novel approach to large language model (LLM) distillation by formulating it as a constrained reinforcement learning problem. While recent work has begun exploring the integration of task-specific rewards into distillation processes, existing methods typically rely on ad-hoc reward weighting. We propose a principled optimization framework that maximizes task-specific rewards while constraining the divergence from the teacher model to remain below a specified threshold. Our approach adapts constrained state augmented reinforcement learning to the distillation setting, introducing a modified reward function that maintains theoretical guarantees of constraint satisfaction without requiring state augmentation or teacher model access during deployment and without the computational overhead of the dual Lagrangian methods. Through extensive experiments on mathematical reasoning tasks, we demonstrate that our method achieves better constraint satisfaction rates and better reasoning compared to the soft Lagrangian relaxation baselines while maintaining competitive task performance. Our framework provides a theoretically grounded and practically efficient solution for reward-aware distillation in resource-constrained settings.
title Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.22921