Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Chuzhan, Feng, Wenfeng, Jiang, Guochao, Quan, Guofeng, Liu, Guohua, Zhang, Yuewei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RASD: Retrieval-Augmented Speculative Decoding
by: Quan, Guofeng, et al.
Published: (2025)
by: Quan, Guofeng, et al.
Published: (2025)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward Reinforcement Learning
by: Hao, Chuzhan, et al.
Published: (2025)
by: Hao, Chuzhan, et al.
Published: (2025)
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
by: Jiang, Guochao, et al.
Published: (2025)
by: Jiang, Guochao, et al.
Published: (2025)
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
by: Jiang, Guochao, et al.
Published: (2026)
by: Jiang, Guochao, et al.
Published: (2026)
Mixture-of-LoRAs: An Efficient Multitask Tuning for Large Language Models
by: Feng, Wenfeng, et al.
Published: (2024)
by: Feng, Wenfeng, et al.
Published: (2024)
FAQ: Mitigating Quantization Error via Regenerating Calibration Data with Family-Aware Quantization
by: Xiao, Haiyang, et al.
Published: (2026)
by: Xiao, Haiyang, et al.
Published: (2026)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
by: Li, Weiqing, et al.
Published: (2025)
by: Li, Weiqing, et al.
Published: (2025)
FlashThink: An Early Exit Method For Efficient Reasoning
by: Jiang, Guochao, et al.
Published: (2025)
by: Jiang, Guochao, et al.
Published: (2025)
Why the Valuable Capabilities of LLMs Are Precisely the Unexplainable Ones
by: Cheng, Quan
Published: (2026)
by: Cheng, Quan
Published: (2026)
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling
by: Liu, Yuchen, et al.
Published: (2026)
by: Liu, Yuchen, et al.
Published: (2026)
Beyond Monolithic Architectures: A Multi-Agent Search and Knowledge Optimization Framework for Agentic Search
by: Chen, Yiqun, et al.
Published: (2026)
by: Chen, Yiqun, et al.
Published: (2026)
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
by: Li, Zhuofeng, et al.
Published: (2026)
by: Li, Zhuofeng, et al.
Published: (2026)
OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search
by: Zhang, Erhan, et al.
Published: (2026)
by: Zhang, Erhan, et al.
Published: (2026)
Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation
by: Jivrajani, Madhav, et al.
Published: (2026)
by: Jivrajani, Madhav, et al.
Published: (2026)
LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
by: He, Hang, et al.
Published: (2025)
by: He, Hang, et al.
Published: (2025)
Conformal Segmentation in Industrial Surface Defect Detection with Statistical Guarantees
by: Shen, Cheng, et al.
Published: (2025)
by: Shen, Cheng, et al.
Published: (2025)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
GraphScout: Empowering Large Language Models with Intrinsic Exploration Ability for Agentic Graph Reasoning
by: Ying, Yuchen, et al.
Published: (2026)
by: Ying, Yuchen, et al.
Published: (2026)
SAGE: Steerable Agentic Data Generation for Deep Search with Execution Feedback
by: Xu, Fangyuan, et al.
Published: (2026)
by: Xu, Fangyuan, et al.
Published: (2026)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
by: Wei, Qianshan, et al.
Published: (2026)
by: Wei, Qianshan, et al.
Published: (2026)
ActionCodec: What Makes for Good Action Tokenizers
by: Dong, Zibin, et al.
Published: (2026)
by: Dong, Zibin, et al.
Published: (2026)
Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
by: Liu, Zibin, et al.
Published: (2025)
by: Liu, Zibin, et al.
Published: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
by: Liu, Wei, et al.
Published: (2023)
by: Liu, Wei, et al.
Published: (2023)
BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search
by: Liu, Shiyu, et al.
Published: (2026)
by: Liu, Shiyu, et al.
Published: (2026)
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
by: Shen, Jianghan, et al.
Published: (2026)
by: Shen, Jianghan, et al.
Published: (2026)
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
by: Chu, Meng, et al.
Published: (2026)
by: Chu, Meng, et al.
Published: (2026)
AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning
by: Sun, Jingbo, et al.
Published: (2026)
by: Sun, Jingbo, et al.
Published: (2026)
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning
by: Wang, Haixin, et al.
Published: (2026)
by: Wang, Haixin, et al.
Published: (2026)
Towards Agentic Self-Learning LLMs in Search Environment
by: Sun, Wangtao, et al.
Published: (2025)
by: Sun, Wangtao, et al.
Published: (2025)
Agentic Architect: An Agentic AI Framework for Architecture Design Exploration and Optimization
by: Blasberg, Alexander, et al.
Published: (2026)
by: Blasberg, Alexander, et al.
Published: (2026)
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
by: Wu, Peilin, et al.
Published: (2025)
by: Wu, Peilin, et al.
Published: (2025)
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
by: Tran, Dao, et al.
Published: (2026)
by: Tran, Dao, et al.
Published: (2026)
Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training
by: Yang, Suorong, et al.
Published: (2026)
by: Yang, Suorong, et al.
Published: (2026)
Agentic Critical Training
by: Liu, Weize, et al.
Published: (2026)
by: Liu, Weize, et al.
Published: (2026)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
by: Liu, Xianyang, et al.
Published: (2025)
by: Liu, Xianyang, et al.
Published: (2025)
Dr. Zero: Self-Evolving Search Agents without Training Data
by: Yue, Zhenrui, et al.
Published: (2026)
by: Yue, Zhenrui, et al.
Published: (2026)
Similar Items
-
RASD: Retrieval-Augmented Speculative Decoding
by: Quan, Guofeng, et al.
Published: (2025) -
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
by: Feng, Wenfeng, et al.
Published: (2025) -
AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
by: Feng, Wenfeng, et al.
Published: (2025) -
DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward Reinforcement Learning
by: Hao, Chuzhan, et al.
Published: (2025) -
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
by: Jiang, Guochao, et al.
Published: (2025)