Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | |
| Veröffentlicht: |
Zenodo
2026
|
| Online-Zugang: | https://doi.org/10.5281/zenodo.19720960 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Inhaltsangabe:
- <p class="15"><span>Current performance improvements in Large Inference Models (LRMs) heavily rely on long inference chains, but this mechanism is essentially an unconstrained sequence generation process, leading to redundant inference, computational waste, and decreased generalization ability. Existing methods (such as Self-Braking Tuning) mainly alleviate over-inference through "termination mechanisms," but they still fail to address the core problem: the inference process itself lacks structural optimality constraints.</span></p> <p class="15"> </p> <p class="15"><span>This paper proposes a novel training paradigm: Combinatorial Trajectory Optimization (CTO), which for the first time systematically introduces combinatorial optimization theory into the modeling of LRM inference processes. We formalize the inference process as a discrete structural optimization problem, treating each step of inference as node selection in a graph, and the complete inference path as a combinatorial object. By constructing a global optimality objective function, we perform structural-level optimization of the inference path.</span></p> <p class="15"> </p> <p class="15"><span>The core innovations of this method include:</span></p> <p class="15"> </p> <p class="15"><span>(1) proposing a unified modeling framework of "inference path = combinatorial structure";</span></p> <p class="15"> </p> <p class="15"><span>(2) defining inference redundancy as the dominance relationship and redundant substructures in the combinatorial structure;</span></p> <p class="15"> </p> <p class="15"><span>(3) introducing "combinatorial pruning operators" for training signal generation;</span></p> <p class="15"> </p> <p class="15"><span>(4) designing "structural consistency loss" to achieve end-to-end training;</span></p> <p class="15"> </p> <p class="15"><span>(5) proposing "differentiable combinatorial relaxation" to achieve gradient optimization.</span></p> <p class="15"> </p> <p class="15"><span>This paradigm fundamentally transforms LRM training from a "sequence fitting problem" to a "combinatorial structure optimization problem," providing a theoretical foundation for efficient inference.</span></p>