Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917480368177152 |
|---|---|
| author | Tan, Wenhui Parascandolo, Fiorenzo Sangineto, Enver Ju, Jianzhong Luo, Zhenbo Cao, Qian Cucchiara, Rita Song, Ruihua Luan, Jian |
| author_facet | Tan, Wenhui Parascandolo, Fiorenzo Sangineto, Enver Ju, Jianzhong Luo, Zhenbo Cao, Qian Cucchiara, Rita Song, Ruihua Luan, Jian |
| contents | Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases pass@$n$ accuracy. Empirically, the final-layer posterior of post-trained LRMs exhibit sharply reduced entropy, while the entropy of intermediate layers remains relatively high. Motivated by this entropy asymmetry, we propose Latent Exploration Decoding (LED), a depth-conditioned decoding strategy. LED aggregates intermediate posteriors via cumulative sum and selects depth configurations with maximal entropy as exploration candidates. Without additional training or parameters, LED consistently improves pass@1 and pass@16 accuracy by 0.61 and 1.03 percentage points across multiple reasoning benchmarks and models. Furthermore, integrating LED into reinforcement learning, e.g., using GRPO as the rollout strategy, yields faster reward improvement and higher final performance, due to the efficient exploration capability of LED. Project page: https://github.com/AlbertTan404/LED. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_01698 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models Tan, Wenhui Parascandolo, Fiorenzo Sangineto, Enver Ju, Jianzhong Luo, Zhenbo Cao, Qian Cucchiara, Rita Song, Ruihua Luan, Jian Computation and Language Machine Learning Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases pass@$n$ accuracy. Empirically, the final-layer posterior of post-trained LRMs exhibit sharply reduced entropy, while the entropy of intermediate layers remains relatively high. Motivated by this entropy asymmetry, we propose Latent Exploration Decoding (LED), a depth-conditioned decoding strategy. LED aggregates intermediate posteriors via cumulative sum and selects depth configurations with maximal entropy as exploration candidates. Without additional training or parameters, LED consistently improves pass@1 and pass@16 accuracy by 0.61 and 1.03 percentage points across multiple reasoning benchmarks and models. Furthermore, integrating LED into reinforcement learning, e.g., using GRPO as the rollout strategy, yields faster reward improvement and higher final performance, due to the efficient exploration capability of LED. Project page: https://github.com/AlbertTan404/LED. |
| title | Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2602.01698 |