From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jiaxiang, Wang, Zhuo, Zou, Mingxi, Li, Zhucong, Zhou, Zhijian, Wang, Song, Xu, Zenglin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Guideline Forest: Retrieval-Augmented Reasoning with Branching Experience-Induced Guidelines
by: Chen, Jiaxiang, et al.
Published: (2025)
by: Chen, Jiaxiang, et al.
Published: (2025)
FinHEAR: Human Expertise and Adaptive Risk-Aware Temporal Reasoning for Financial Decision-Making
by: Chen, Jiaxiang, et al.
Published: (2025)
by: Chen, Jiaxiang, et al.
Published: (2025)
DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding
by: Zou, Mingxi, et al.
Published: (2026)
by: Zou, Mingxi, et al.
Published: (2026)
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
by: Sheshanarayana, Disha, et al.
Published: (2026)
by: Sheshanarayana, Disha, et al.
Published: (2026)
TimeCNN: Refining Cross-Variable Interaction on Time Point for Time Series Forecasting
by: Hu, Ao, et al.
Published: (2024)
by: Hu, Ao, et al.
Published: (2024)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
by: Wang, Qibin, et al.
Published: (2025)
by: Wang, Qibin, et al.
Published: (2025)
Can we only use guideline instead of shot in prompt?
by: Chen, Jiaxiang, et al.
Published: (2024)
by: Chen, Jiaxiang, et al.
Published: (2024)
From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs
by: Wang, Xiaoxuan, et al.
Published: (2025)
by: Wang, Xiaoxuan, et al.
Published: (2025)
ChemAmp: Amplified Chemistry Tools via Composable Agents
by: Li, Zhucong, et al.
Published: (2025)
by: Li, Zhucong, et al.
Published: (2025)
Nested Spatio-Temporal Time Series Forecasting
by: Ai, Yinghao, et al.
Published: (2026)
by: Ai, Yinghao, et al.
Published: (2026)
CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs
by: Zeng, Yongcheng, et al.
Published: (2025)
by: Zeng, Yongcheng, et al.
Published: (2025)
Safety Reasoning with Guidelines
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
by: Xie, Lipeng, et al.
Published: (2025)
by: Xie, Lipeng, et al.
Published: (2025)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Partial Differential Equations is All You Need for Generating Neural Architectures -- A Theory for Physical Artificial Intelligence Systems
by: Guo, Ping, et al.
Published: (2021)
by: Guo, Ping, et al.
Published: (2021)
GNNavigator: Towards Adaptive Training of Graph Neural Networks via Automatic Guideline Exploration
by: Qiao, Tong, et al.
Published: (2024)
by: Qiao, Tong, et al.
Published: (2024)
Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion
by: Tan, Zhen, et al.
Published: (2026)
by: Tan, Zhen, et al.
Published: (2026)
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Can LLMs Learn to Reason Robustly under Noisy Supervision?
by: Yang, Shenzhi, et al.
Published: (2026)
by: Yang, Shenzhi, et al.
Published: (2026)
GeoPro-Net: Learning Interpretable Spatiotemporal Prediction Models through Statistically-Guided Geo-Prototyping
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs
by: Zhou, Yitong, et al.
Published: (2025)
by: Zhou, Yitong, et al.
Published: (2025)
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
by: Chen, Xinzhu, et al.
Published: (2025)
by: Chen, Xinzhu, et al.
Published: (2025)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
by: Wu, Jianfei, et al.
Published: (2026)
by: Wu, Jianfei, et al.
Published: (2026)
Reinforcing Numerical Reasoning in LLMs for Tabular Prediction via Structural Priors
by: Cai, Pengxiang, et al.
Published: (2025)
by: Cai, Pengxiang, et al.
Published: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
FA-INR: Adaptive Implicit Neural Representations for Interpretable Exploration of Simulation Ensembles
by: Li, Ziwei, et al.
Published: (2025)
by: Li, Ziwei, et al.
Published: (2025)
Revisiting Long-term Time Series Forecasting: An Investigation on Linear Mapping
by: Li, Zhe, et al.
Published: (2023)
by: Li, Zhe, et al.
Published: (2023)
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
Can LLMs Effectively Leverage Graph Structural Information through Prompts, and Why?
by: Huang, Jin, et al.
Published: (2023)
by: Huang, Jin, et al.
Published: (2023)
Incomplete Depression Feature Selection with Missing EEG Channels
by: Gong, Zhijian, et al.
Published: (2025)
by: Gong, Zhijian, et al.
Published: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models
by: Zhuo, Zhijian, et al.
Published: (2024)
by: Zhuo, Zhijian, et al.
Published: (2024)
Divide-Verify-Refine: Can LLMs Self-Align with Complex Instructions?
by: Zhang, Xianren, et al.
Published: (2024)
by: Zhang, Xianren, et al.
Published: (2024)
Similar Items
-
Guideline Forest: Retrieval-Augmented Reasoning with Branching Experience-Induced Guidelines
by: Chen, Jiaxiang, et al.
Published: (2025) -
FinHEAR: Human Expertise and Adaptive Risk-Aware Temporal Reasoning for Financial Decision-Making
by: Chen, Jiaxiang, et al.
Published: (2025) -
DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding
by: Zou, Mingxi, et al.
Published: (2026) -
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
by: Sheshanarayana, Disha, et al.
Published: (2026) -
TimeCNN: Refining Cross-Variable Interaction on Time Point for Time Series Forecasting
by: Hu, Ao, et al.
Published: (2024)