PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Haoyu, Zhu, Yun, Yuan, Yuqian, Yuan, Bo, Zhang, Wenqiao, Tang, Siliang, Xiao, Jun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914470448594944
author Zheng, Haoyu
Zhu, Yun
Yuan, Yuqian
Yuan, Bo
Zhang, Wenqiao
Tang, Siliang
Xiao, Jun
author_facet Zheng, Haoyu
Zhu, Yun
Yuan, Yuqian
Yuan, Bo
Zhang, Wenqiao
Tang, Siliang
Xiao, Jun
contents Strategic planning is critical for multi-step reasoning, yet compact Large Language Models (LLMs) often lack the capacity to formulate global strategies, leading to error propagation in long-horizon tasks. Our analysis reveals that LLMs possess latent reasoning capabilities that can be unlocked when conditioned on explicit plans from a teacher model; however, runtime reliance on external guidance is often impractical due to latency and availability constraints. To bridge this gap, we propose PILOT (Planning via Internalized Latent Optimization Trajectories), a non-invasive framework designed to internalize the strategic oversight of large models into intrinsic Latent Guidance. Instead of altering backbone weights, PILOT employs a lightweight Hyper-Network to synthesize a query-conditioned Latent Guidance vector. This vector acts as an internal steering mechanism, guiding the model's representations toward optimal reasoning paths. Extensive experiments on mathematical and coding benchmarks demonstrate that PILOT effectively stabilizes reasoning trajectories, consistently outperforming strong baselines (e.g., +8.9% on MATH500) with negligible inference latency.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19917
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
Zheng, Haoyu
Zhu, Yun
Yuan, Yuqian
Yuan, Bo
Zhang, Wenqiao
Tang, Siliang
Xiao, Jun
Computation and Language
Strategic planning is critical for multi-step reasoning, yet compact Large Language Models (LLMs) often lack the capacity to formulate global strategies, leading to error propagation in long-horizon tasks. Our analysis reveals that LLMs possess latent reasoning capabilities that can be unlocked when conditioned on explicit plans from a teacher model; however, runtime reliance on external guidance is often impractical due to latency and availability constraints. To bridge this gap, we propose PILOT (Planning via Internalized Latent Optimization Trajectories), a non-invasive framework designed to internalize the strategic oversight of large models into intrinsic Latent Guidance. Instead of altering backbone weights, PILOT employs a lightweight Hyper-Network to synthesize a query-conditioned Latent Guidance vector. This vector acts as an internal steering mechanism, guiding the model's representations toward optimal reasoning paths. Extensive experiments on mathematical and coding benchmarks demonstrate that PILOT effectively stabilizes reasoning trajectories, consistently outperforming strong baselines (e.g., +8.9% on MATH500) with negligible inference latency.
title PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2601.19917