DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Fang, Xuan, Weihao, Qi, Heli, Lu, Ximing, Tu, Aaron, Li, Li Erran, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LightSearcher: Efficient DeepSearch via Experiential Memory
by: Lan, Hengzhi, et al.
Published: (2025)
by: Lan, Hengzhi, et al.
Published: (2025)
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
by: Wu, Fang, et al.
Published: (2025)
by: Wu, Fang, et al.
Published: (2025)
In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided Search
by: Li, Huihan, et al.
Published: (2023)
by: Li, Huihan, et al.
Published: (2023)
TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs
by: Xie, Yutao, et al.
Published: (2026)
by: Xie, Yutao, et al.
Published: (2026)
Annotation-Free Reinforcement Learning Query Rewriting via Verifiable Search Reward
by: Cha, Sungguk, et al.
Published: (2025)
by: Cha, Sungguk, et al.
Published: (2025)
Improve Value Estimation of Q Function and Reshape Reward with Monte Carlo Tree Search
by: Li, Jiamian
Published: (2024)
by: Li, Jiamian
Published: (2024)
Deep Reinforcement Learning Xiangqi Player with Monte Carlo Tree Search
by: Yilmaz, Berk, et al.
Published: (2025)
by: Yilmaz, Berk, et al.
Published: (2025)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
by: Seo, Wooseok, et al.
Published: (2025)
by: Seo, Wooseok, et al.
Published: (2025)
Tailoring Self-Rationalizers with Multi-Reward Distillation
by: Ramnath, Sahana, et al.
Published: (2023)
by: Ramnath, Sahana, et al.
Published: (2023)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
by: Liu, Jiacheng, et al.
Published: (2023)
by: Liu, Jiacheng, et al.
Published: (2023)
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
by: Xin, Huajian, et al.
Published: (2024)
by: Xin, Huajian, et al.
Published: (2024)
Monte Carlo Search Algorithms Discovering Monte Carlo Tree Search Exploration Terms
by: Cazenave, Tristan
Published: (2024)
by: Cazenave, Tristan
Published: (2024)
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning
by: Liu, Huanyu, et al.
Published: (2025)
by: Liu, Huanyu, et al.
Published: (2025)
Monte Carlo Permutation Search
by: Cazenave, Tristan
Published: (2025)
by: Cazenave, Tristan
Published: (2025)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
Empirical-MCTS: Continuous Agent Evolution via Dual-Experience Monte Carlo Tree Search
by: Lu, Hao, et al.
Published: (2026)
by: Lu, Hao, et al.
Published: (2026)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
by: Lee, Jaeyoung, et al.
Published: (2024)
by: Lee, Jaeyoung, et al.
Published: (2024)
Vital: Vulnerability-Oriented Symbolic Execution via Type-Unsafe Pointer-Guided Monte Carlo Tree Search
by: Tu, Haoxin, et al.
Published: (2024)
by: Tu, Haoxin, et al.
Published: (2024)
Discovering Mathematical Formulas from Data via GPT-guided Monte Carlo Tree Search
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
Searching Efficient Deep Architectures for Radar Target Detection using Monte-Carlo Tree Search
by: Lallouet, Noé, et al.
Published: (2025)
by: Lallouet, Noé, et al.
Published: (2025)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)
by: Lu, Ximing, et al.
Published: (2025)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
by: Zhao, Qingfei, et al.
Published: (2025)
by: Zhao, Qingfei, et al.
Published: (2025)
Continuous Monte Carlo Graph Search
by: Kujanpää, Kalle, et al.
Published: (2022)
by: Kujanpää, Kalle, et al.
Published: (2022)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
Verifiable, Efficient and Confidentiality-Preserving Graph Search with Transparency
by: Wang, Qiuhao, et al.
Published: (2025)
by: Wang, Qiuhao, et al.
Published: (2025)
Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards
by: Zhang, Jiajie, et al.
Published: (2026)
by: Zhang, Jiajie, et al.
Published: (2026)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
by: Wang, Peisong, et al.
Published: (2025)
by: Wang, Peisong, et al.
Published: (2025)
Segment Anything with Multiple Modalities
by: Xiao, Aoran, et al.
Published: (2024)
by: Xiao, Aoran, et al.
Published: (2024)
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
UNSAT Solver Synthesis via Monte Carlo Forest Search
by: Cameron, Chris, et al.
Published: (2022)
by: Cameron, Chris, et al.
Published: (2022)
Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control
by: Alzorgan, Hazim, et al.
Published: (2025)
by: Alzorgan, Hazim, et al.
Published: (2025)
Explainable Graph Neural Architecture Search via Monte-Carlo Tree Search (Full version)
by: Sasaki, Yuya
Published: (2023)
by: Sasaki, Yuya
Published: (2023)
IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
by: Guo, Xu, et al.
Published: (2025)
by: Guo, Xu, et al.
Published: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
by: Li, Jian, et al.
Published: (2026)
by: Li, Jian, et al.
Published: (2026)
Toward Template-Free Explainability for Monte Carlo Tree Search
by: Lu, Siqi, et al.
Published: (2026)
by: Lu, Siqi, et al.
Published: (2026)
Monte Carlo Spin Simulations of Magnetic Noise -- The Search for Pivoting
by: Mickelsen, D. L., et al.
Published: (2024)
by: Mickelsen, D. L., et al.
Published: (2024)
RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code Generation
by: Lin, Yuanyuan, et al.
Published: (2025)
by: Lin, Yuanyuan, et al.
Published: (2025)
Similar Items
-
LightSearcher: Efficient DeepSearch via Experiential Memory
by: Lan, Hengzhi, et al.
Published: (2025) -
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
by: Yu, Tao, et al.
Published: (2026) -
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
by: Wu, Fang, et al.
Published: (2025) -
In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided Search
by: Li, Huihan, et al.
Published: (2023) -
TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs
by: Xie, Yutao, et al.
Published: (2026)