Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Jiahe, Paladugu, Abhijay, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025)
by: Wang, Tevin, et al.
Published: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
Can Coding Agents Be General Agents?
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
by: Li, Xiaochuan, et al.
Published: (2024)
by: Li, Xiaochuan, et al.
Published: (2024)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
by: Tan, Zelin, et al.
Published: (2025)
by: Tan, Zelin, et al.
Published: (2025)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
by: Hwang, Dongyoon, et al.
Published: (2025)
by: Hwang, Dongyoon, et al.
Published: (2025)
The Impact of Post-training on Data Contamination
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2026)
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
by: Liu, Tong, et al.
Published: (2026)
by: Liu, Tong, et al.
Published: (2026)
Diagnosis of Fuel Cell Health Status with Deep Sparse Auto-Encoder Neural Network
by: Fei, Chenyan, et al.
Published: (2025)
by: Fei, Chenyan, et al.
Published: (2025)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
by: Sun, Liwen, et al.
Published: (2024)
by: Sun, Liwen, et al.
Published: (2024)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
Efficient Agent Training for Computer Use
by: He, Yanheng, et al.
Published: (2025)
by: He, Yanheng, et al.
Published: (2025)
The Context Gathering Decision Process: A POMDP Framework for Agentic Search
by: Kausik, Chinmaya, et al.
Published: (2026)
by: Kausik, Chinmaya, et al.
Published: (2026)
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
by: Yao, Yihang, et al.
Published: (2026)
by: Yao, Yihang, et al.
Published: (2026)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025)
by: Wen, Kaiyue, et al.
Published: (2025)
Low Rank Gradients and Where to Find Them
by: Sonthalia, Rishi, et al.
Published: (2025)
by: Sonthalia, Rishi, et al.
Published: (2025)
BoA: Attention-aware Post-training Quantization without Backpropagation
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
by: Guan, Xinyan, et al.
Published: (2024)
by: Guan, Xinyan, et al.
Published: (2024)
Subgoal Search For Complex Reasoning Tasks
by: Czechowski, Konrad, et al.
Published: (2021)
by: Czechowski, Konrad, et al.
Published: (2021)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Prioritized Replay for RL Post-training
by: Fatemi, Mehdi
Published: (2026)
by: Fatemi, Mehdi
Published: (2026)
GNN Explanations that do not Explain and How to find Them
by: Azzolin, Steve, et al.
Published: (2026)
by: Azzolin, Steve, et al.
Published: (2026)
What Cohort INRs Encode and Where to Freeze Them
by: Sideri-Lampretsa, Vasiliki, et al.
Published: (2026)
by: Sideri-Lampretsa, Vasiliki, et al.
Published: (2026)
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
by: Zhang, Huishuai, et al.
Published: (2025)
by: Zhang, Huishuai, et al.
Published: (2025)
Multimodal Fine-grained Reasoning for Post Quality Evaluation
by: Guo, Xiaoxu, et al.
Published: (2025)
by: Guo, Xiaoxu, et al.
Published: (2025)
Obtaining Example-Based Explanations from Deep Neural Networks
by: Dong, Genghua, et al.
Published: (2025)
by: Dong, Genghua, et al.
Published: (2025)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
by: Han, Dongge, et al.
Published: (2025)
by: Han, Dongge, et al.
Published: (2025)
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
by: Agarwal, Shivam, et al.
Published: (2025)
by: Agarwal, Shivam, et al.
Published: (2025)
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
by: Joshi, Siddharth, et al.
Published: (2023)
by: Joshi, Siddharth, et al.
Published: (2023)
The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning
by: Zhuang, Ren, et al.
Published: (2026)
by: Zhuang, Ren, et al.
Published: (2026)
CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
by: Mroueh, Youssef, et al.
Published: (2026)
by: Mroueh, Youssef, et al.
Published: (2026)
How to Square Tensor Networks and Circuits Without Squaring Them
by: Loconte, Lorenzo, et al.
Published: (2025)
by: Loconte, Lorenzo, et al.
Published: (2025)
Calibrated Language Models and How to Find Them with Label Smoothing
by: Huang, Jerry, et al.
Published: (2025)
by: Huang, Jerry, et al.
Published: (2025)
From Simulation to Enaction: Post-trained language models recognize and react to their own generations
by: G., Asvin, et al.
Published: (2026)
by: G., Asvin, et al.
Published: (2026)
Transcendence: Generative Models Can Outperform The Experts That Train Them
by: Zhang, Edwin, et al.
Published: (2024)
by: Zhang, Edwin, et al.
Published: (2024)
Similar Items
-
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025) -
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026) -
Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
by: Zhang, Zhenyu, et al.
Published: (2025) -
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026) -
Can Coding Agents Be General Agents?
by: Ivanov, Maksim, et al.
Published: (2026)