FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Kehan, Dong, Haonan, Kang, Zhaolu, Zhu, Zhengzhou, Song, Guojie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908934852313088
author Jiang, Kehan
Dong, Haonan
Kang, Zhaolu
Zhu, Zhengzhou
Song, Guojie
author_facet Jiang, Kehan
Dong, Haonan
Kang, Zhaolu
Zhu, Zhengzhou
Song, Guojie
contents Recent Large Reasoning Models (LRMs) like DeepSeek-R1 have demonstrated remarkable success in complex reasoning tasks, exhibiting human-like patterns in exploring multiple alternative solutions. Upon closer inspection, however, we uncover a surprising phenomenon: The First is The Best, where alternative solutions are not merely suboptimal but potentially detrimental. This observation challenges widely accepted test-time scaling laws, leading us to hypothesize that errors within the reasoning path scale concurrently with test time. Through comprehensive empirical analysis, we characterize errors as a forest-structured Forest of Errors (FoE) and conclude that FoE makes the First the Best, which is underpinned by rigorous theoretical analysis. Leveraging these insights, we propose RED, a self-guided efficient reasoning framework comprising two components: I) Refining First, which suppresses FoE growth in the first solution; and II) Discarding Subs, which prunes subsequent FoE via dual-consistency. Extensive experiments across five benchmarks and six backbone models demonstrate that RED outperforms eight competitive baselines, achieving performance gains of up to 19.0% while reducing token consumption by 37.7% ~ 70.4%. Moreover, comparative experiments on FoE metrics shed light on how RED achieves effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02967
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models
Jiang, Kehan
Dong, Haonan
Kang, Zhaolu
Zhu, Zhengzhou
Song, Guojie
Artificial Intelligence
Computation and Language
Recent Large Reasoning Models (LRMs) like DeepSeek-R1 have demonstrated remarkable success in complex reasoning tasks, exhibiting human-like patterns in exploring multiple alternative solutions. Upon closer inspection, however, we uncover a surprising phenomenon: The First is The Best, where alternative solutions are not merely suboptimal but potentially detrimental. This observation challenges widely accepted test-time scaling laws, leading us to hypothesize that errors within the reasoning path scale concurrently with test time. Through comprehensive empirical analysis, we characterize errors as a forest-structured Forest of Errors (FoE) and conclude that FoE makes the First the Best, which is underpinned by rigorous theoretical analysis. Leveraging these insights, we propose RED, a self-guided efficient reasoning framework comprising two components: I) Refining First, which suppresses FoE growth in the first solution; and II) Discarding Subs, which prunes subsequent FoE via dual-consistency. Extensive experiments across five benchmarks and six backbone models demonstrate that RED outperforms eight competitive baselines, achieving performance gains of up to 19.0% while reducing token consumption by 37.7% ~ 70.4%. Moreover, comparative experiments on FoE metrics shed light on how RED achieves effectiveness.
title FoE: Forest of Errors Makes the First Solution the Best in Large Reasoning Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.02967