Policy of Thoughts: Scaling LLM Reasoning via Test-time Policy Evolution
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Zhengbo, Xian, Hongyu, Wang, Qinglong, Ma, Yunpu, Wang, Zhebo, Zhang, Zifan, Kong, Dezhang, Han, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
by: Wang, Zhebo, et al.
Published: (2026)
by: Wang, Zhebo, et al.
Published: (2026)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
by: Bi, Zhenni, et al.
Published: (2024)
by: Bi, Zhenni, et al.
Published: (2024)
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning
by: Xu, Zifan, et al.
Published: (2023)
by: Xu, Zifan, et al.
Published: (2023)
Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
Agentic Proposing: Enhancing Large Language Model Reasoning via Compositional Skill Synthesis
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
by: Ye, Wengao, et al.
Published: (2025)
by: Ye, Wengao, et al.
Published: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
by: Yang, Ling, et al.
Published: (2025)
by: Yang, Ling, et al.
Published: (2025)
LLM Reasoning Is Latent, Not the Chain of Thought
by: Wang, Wenshuo
Published: (2026)
by: Wang, Wenshuo
Published: (2026)
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
by: Han, Xinchen, et al.
Published: (2026)
by: Han, Xinchen, et al.
Published: (2026)
The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
by: Wei, Zihao, et al.
Published: (2025)
by: Wei, Zihao, et al.
Published: (2025)
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
by: Chen, Yihang, et al.
Published: (2025)
by: Chen, Yihang, et al.
Published: (2025)
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
by: Li, Chenyi, et al.
Published: (2026)
by: Li, Chenyi, et al.
Published: (2026)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
Agentic Policy Optimization via Instruction-Policy Co-Evolution
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
by: Wang, Guangzhi, et al.
Published: (2025)
by: Wang, Guangzhi, et al.
Published: (2025)
ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models
by: Xu, Zhenhua, et al.
Published: (2026)
by: Xu, Zhenhua, et al.
Published: (2026)
Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
Scaling Test-time Compute for LLM Agents
by: Zhu, King, et al.
Published: (2025)
by: Zhu, King, et al.
Published: (2025)
Code Evolution for Control: Synthesizing Policies via LLM-Driven Evolutionary Search
by: Guo, Ping, et al.
Published: (2026)
by: Guo, Ping, et al.
Published: (2026)
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
by: Yin, Bo, et al.
Published: (2026)
by: Yin, Bo, et al.
Published: (2026)
Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
by: Kong, Dezhang, et al.
Published: (2025)
by: Kong, Dezhang, et al.
Published: (2025)
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
by: Hua, Xingyuan, et al.
Published: (2026)
by: Hua, Xingyuan, et al.
Published: (2026)
DRT: Deep Reasoning Translation via Long Chain-of-Thought
by: Wang, Jiaan, et al.
Published: (2024)
by: Wang, Jiaan, et al.
Published: (2024)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
by: Jiao, Rui, et al.
Published: (2025)
by: Jiao, Rui, et al.
Published: (2025)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
Scaling Graph Chain-of-Thought Reasoning: A Multi-Agent Framework with Efficient LLM Serving
by: Huan, Chengying, et al.
Published: (2025)
by: Huan, Chengying, et al.
Published: (2025)
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
by: Hao, Qianyue, et al.
Published: (2025)
by: Hao, Qianyue, et al.
Published: (2025)
Atom of Thoughts for Markov LLM Test-Time Scaling
by: Teng, Fengwei, et al.
Published: (2025)
by: Teng, Fengwei, et al.
Published: (2025)
Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning
by: Baldwin, Wilder, et al.
Published: (2026)
by: Baldwin, Wilder, et al.
Published: (2026)
CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
by: Zhang, Shijie, et al.
Published: (2025)
by: Zhang, Shijie, et al.
Published: (2025)
Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
by: Chun, Inkook, et al.
Published: (2025)
by: Chun, Inkook, et al.
Published: (2025)
Reasoning Compression with Mixed-Policy Distillation
by: Yang, Han, et al.
Published: (2026)
by: Yang, Han, et al.
Published: (2026)
Similar Items
-
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
by: Wang, Zhebo, et al.
Published: (2026) -
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
by: Bi, Zhenni, et al.
Published: (2024) -
Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning
by: Jiao, Zhengbo, et al.
Published: (2026) -
LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning
by: Xu, Zifan, et al.
Published: (2023) -
Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction
by: Jiao, Zhengbo, et al.
Published: (2026)