AdapThink: Adaptive Thinking Preferences for Reasoning Language Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wan, Xu, Wang, Wei, Xu, Wenyue, Yin, Wotao, Song, Jie, Sun, Mingyang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911018935910400
author Wan, Xu
Wang, Wei
Xu, Wenyue
Yin, Wotao
Song, Jie
Sun, Mingyang
author_facet Wan, Xu
Wang, Wei
Xu, Wenyue
Yin, Wotao
Song, Jie
Sun, Mingyang
contents Reinforcement Learning (RL)-based post-training has significantly advanced the complex reasoning capabilities of language models, fostering sophisticated self-reflection processes. However, this ``slow thinking'' paradigm presents a critical challenge to reasoning efficiency: models may expend excessive computation on simple questions and shift reasoning prematurely for complex ones. Previous mechanisms typically rely on static length budgets or predefined rules, lacking the adaptability for varying question complexities and models' evolving capabilities. To this end, we propose AdapThink, an adaptive post-training framework designed to induce more efficient thinking while maintaining the performance of reasoning language models. Specifically, AdapThink incorporates two key mechanisms: 1) A group-relative reward function that leverages model confidence and response's characteristic to dynamically adjust the preference of reflection-related transition words without resorting to a fixed length preference. 2) A diversity-aware sampling mechanism that balances the training group's solution accuracy with reasoning diversity via an entropy-guided score. Experiments on several mathematical reasoning datasets with DeepSeek-distilled models demonstrate AdapThink's advantages in enabling adaptive reasoning patterns and mitigating the inefficiencies.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18237
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
Wan, Xu
Wang, Wei
Xu, Wenyue
Yin, Wotao
Song, Jie
Sun, Mingyang
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement Learning (RL)-based post-training has significantly advanced the complex reasoning capabilities of language models, fostering sophisticated self-reflection processes. However, this ``slow thinking'' paradigm presents a critical challenge to reasoning efficiency: models may expend excessive computation on simple questions and shift reasoning prematurely for complex ones. Previous mechanisms typically rely on static length budgets or predefined rules, lacking the adaptability for varying question complexities and models' evolving capabilities. To this end, we propose AdapThink, an adaptive post-training framework designed to induce more efficient thinking while maintaining the performance of reasoning language models. Specifically, AdapThink incorporates two key mechanisms: 1) A group-relative reward function that leverages model confidence and response's characteristic to dynamically adjust the preference of reflection-related transition words without resorting to a fixed length preference. 2) A diversity-aware sampling mechanism that balances the training group's solution accuracy with reasoning diversity via an entropy-guided score. Experiments on several mathematical reasoning datasets with DeepSeek-distilled models demonstrate AdapThink's advantages in enabling adaptive reasoning patterns and mitigating the inefficiencies.
title AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.18237