Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Kotoge, Rikuto, Nishimura, Mai, Ma, Jiaxin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
by: Kotoge, Rikuto, et al.
Published: (2025)
by: Kotoge, Rikuto, et al.
Published: (2025)
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
by: Cheng, Xueqi, et al.
Published: (2026)
by: Cheng, Xueqi, et al.
Published: (2026)
DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
by: Kang, Minki, et al.
Published: (2026)
by: Kang, Minki, et al.
Published: (2026)
CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models
by: Gong, Peiyuan, et al.
Published: (2024)
by: Gong, Peiyuan, et al.
Published: (2024)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles
by: Shah, Kulin, et al.
Published: (2024)
by: Shah, Kulin, et al.
Published: (2024)
MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability
by: Wu, Weiqi, et al.
Published: (2025)
by: Wu, Weiqi, et al.
Published: (2025)
AI-VaxGuide: An Agentic RAG-Based LLM for Vaccination Decisions
by: Zeggai, Abdellah, et al.
Published: (2025)
by: Zeggai, Abdellah, et al.
Published: (2025)
On-Policy Context Distillation for Language Models
by: Ye, Tianzhu, et al.
Published: (2026)
by: Ye, Tianzhu, et al.
Published: (2026)
Combining On-Policy Optimization and Distillation for Long-Context Reasoning in Large Language Models
by: Ramos, Miguel Moura, et al.
Published: (2026)
by: Ramos, Miguel Moura, et al.
Published: (2026)
Self-Policy Distillation via Capability-Selective Subspace Projection
by: Hao, Guangya, et al.
Published: (2026)
by: Hao, Guangya, et al.
Published: (2026)
From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents
by: Sayeed, Mohammad Amaan, et al.
Published: (2025)
by: Sayeed, Mohammad Amaan, et al.
Published: (2025)
Is Agentic RAG worth it? An experimental comparison of RAG approaches
by: Ferrazzi, Pietro, et al.
Published: (2026)
by: Ferrazzi, Pietro, et al.
Published: (2026)
Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
AT$^2$PO: Agentic Turn-based Policy Optimization via Tree Search
by: Zong, Zefang, et al.
Published: (2026)
by: Zong, Zefang, et al.
Published: (2026)
Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
by: He, Bowei, et al.
Published: (2025)
by: He, Bowei, et al.
Published: (2025)
Causal Distillation: Transferring Structured Explanations from Large to Compact Language Models
by: Muhebwa, Aggrey, et al.
Published: (2025)
by: Muhebwa, Aggrey, et al.
Published: (2025)
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
by: Kubis, Marek, et al.
Published: (2025)
by: Kubis, Marek, et al.
Published: (2025)
Distilling LLMs' Decomposition Abilities into Compact Language Models
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
Compact Language Models via Pruning and Knowledge Distillation
by: Muralidharan, Saurav, et al.
Published: (2024)
by: Muralidharan, Saurav, et al.
Published: (2024)
Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
by: Do, Cong-Thanh, et al.
Published: (2025)
by: Do, Cong-Thanh, et al.
Published: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Mixture-of-Agents Enhances Large Language Model Capabilities
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
by: Li, Weizhen, et al.
Published: (2025)
by: Li, Weizhen, et al.
Published: (2025)
A Novel Paradigm Boosting Translation Capabilities of Large Language Models
by: Guo, Jiaxin, et al.
Published: (2024)
by: Guo, Jiaxin, et al.
Published: (2024)
Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided
by: Zhan, Hongli, et al.
Published: (2024)
by: Zhan, Hongli, et al.
Published: (2024)
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
by: Ma, Yufei, et al.
Published: (2026)
by: Ma, Yufei, et al.
Published: (2026)
Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
by: Deng, Wenlong, et al.
Published: (2025)
by: Deng, Wenlong, et al.
Published: (2025)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
by: Dai, Hankun, et al.
Published: (2025)
by: Dai, Hankun, et al.
Published: (2025)
Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG
by: Hazoom, Moshe, et al.
Published: (2026)
by: Hazoom, Moshe, et al.
Published: (2026)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
by: Zhang, Yuanchi, et al.
Published: (2024)
by: Zhang, Yuanchi, et al.
Published: (2024)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Probing-RAG: Self-Probing to Guide Language Models in Selective Document Retrieval
by: Baek, Ingeol, et al.
Published: (2024)
by: Baek, Ingeol, et al.
Published: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Similar Items
-
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
by: Kotoge, Rikuto, et al.
Published: (2025) -
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
by: Cheng, Xueqi, et al.
Published: (2026) -
DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
by: Gu, Chenyang, et al.
Published: (2025) -
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
by: Ling Team, et al.
Published: (2025) -
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models
by: Wang, Rui, et al.
Published: (2024)