Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Singh, Joykirat, Magazine, Raghav, Pandya, Yash, Nambi, Akshay |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspaces
por: Gupta, Karan, et al.
Publicado: (2026)
por: Gupta, Karan, et al.
Publicado: (2026)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
por: Singh, Joykirat, et al.
Publicado: (2024)
por: Singh, Joykirat, et al.
Publicado: (2024)
PromptWizard: Task-Aware Prompt Optimization Framework
por: Agarwal, Eshaan, et al.
Publicado: (2024)
por: Agarwal, Eshaan, et al.
Publicado: (2024)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
por: Singh, Joykirat, et al.
Publicado: (2025)
por: Singh, Joykirat, et al.
Publicado: (2025)
Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
por: Agarwal, Aradhye, et al.
Publicado: (2026)
por: Agarwal, Aradhye, et al.
Publicado: (2026)
Fara-7B: An Efficient Agentic Model for Computer Use
por: Awadallah, Ahmed, et al.
Publicado: (2025)
por: Awadallah, Ahmed, et al.
Publicado: (2025)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
por: Singh, Joykirat, et al.
Publicado: (2025)
por: Singh, Joykirat, et al.
Publicado: (2025)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
por: Kumar, Somnath, et al.
Publicado: (2024)
por: Kumar, Somnath, et al.
Publicado: (2024)
Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty
por: Singh, Joykirat, et al.
Publicado: (2026)
por: Singh, Joykirat, et al.
Publicado: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
por: Xu, Ran, et al.
Publicado: (2025)
por: Xu, Ran, et al.
Publicado: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
por: Lu, Meng, et al.
Publicado: (2025)
por: Lu, Meng, et al.
Publicado: (2025)
Exposing Weak Links in Multi-Agent Systems under Adversarial Prompting
por: Arora, Nirmit, et al.
Publicado: (2025)
por: Arora, Nirmit, et al.
Publicado: (2025)
Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
por: Chavan, Arnav, et al.
Publicado: (2024)
por: Chavan, Arnav, et al.
Publicado: (2024)
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
por: Wang, Jianing, et al.
Publicado: (2026)
por: Wang, Jianing, et al.
Publicado: (2026)
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning
por: He, Yang, et al.
Publicado: (2026)
por: He, Yang, et al.
Publicado: (2026)
Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
por: Zhang, Hang, et al.
Publicado: (2026)
por: Zhang, Hang, et al.
Publicado: (2026)
Mechanistic Behavior Editing of Language Models
por: Singh, Joykirat, et al.
Publicado: (2024)
por: Singh, Joykirat, et al.
Publicado: (2024)
ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
por: Le, Quy Minh, et al.
Publicado: (2025)
por: Le, Quy Minh, et al.
Publicado: (2025)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
por: Kumar, Somnath, et al.
Publicado: (2023)
por: Kumar, Somnath, et al.
Publicado: (2023)
Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning
por: Chen, Liuji, et al.
Publicado: (2026)
por: Chen, Liuji, et al.
Publicado: (2026)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL
por: Xu, Zekun, et al.
Publicado: (2025)
por: Xu, Zekun, et al.
Publicado: (2025)
Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
por: Xu, Siyuan, et al.
Publicado: (2026)
por: Xu, Siyuan, et al.
Publicado: (2026)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
por: Kumar, Somnath, et al.
Publicado: (2024)
por: Kumar, Somnath, et al.
Publicado: (2024)
Tools as Continuous Flow for Evolving Agentic Reasoning
por: Huang, Tairan, et al.
Publicado: (2026)
por: Huang, Tairan, et al.
Publicado: (2026)
Graph-Memoized Reasoning: Foundations Structured Workflow Reuse in Intelligent Systems
por: Singh, Yash Raj
Publicado: (2025)
por: Singh, Yash Raj
Publicado: (2025)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
por: Wu, Junde, et al.
Publicado: (2025)
por: Wu, Junde, et al.
Publicado: (2025)
EROS: Entity-Driven Controlled Policy Document Summarization
por: Singh, Joykirat, et al.
Publicado: (2024)
por: Singh, Joykirat, et al.
Publicado: (2024)
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
por: Zhang, Guibin, et al.
Publicado: (2025)
por: Zhang, Guibin, et al.
Publicado: (2025)
Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
por: Chen, Yifei, et al.
Publicado: (2025)
por: Chen, Yifei, et al.
Publicado: (2025)
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
por: Feng, Jiazhan, et al.
Publicado: (2025)
por: Feng, Jiazhan, et al.
Publicado: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
por: Li, Chenglin, et al.
Publicado: (2026)
por: Li, Chenglin, et al.
Publicado: (2026)
How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
por: Lin, Minhua, et al.
Publicado: (2026)
por: Lin, Minhua, et al.
Publicado: (2026)
Dynamic System Instructions and Tool Exposure for Efficient Agentic LLMs
por: Franko, Uria
Publicado: (2025)
por: Franko, Uria
Publicado: (2025)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
por: Chen, Mingyang, et al.
Publicado: (2025)
por: Chen, Mingyang, et al.
Publicado: (2025)
MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism
por: Liu, Shulin, et al.
Publicado: (2025)
por: Liu, Shulin, et al.
Publicado: (2025)
Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
por: Wu, Wenxun, et al.
Publicado: (2025)
por: Wu, Wenxun, et al.
Publicado: (2025)
Synapse Compendium Aware Federated Knowledge Exchange for Tool Routed LLMs
por: Chakraborty, Abhijit, et al.
Publicado: (2026)
por: Chakraborty, Abhijit, et al.
Publicado: (2026)
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
por: He, Jiashu, et al.
Publicado: (2025)
por: He, Jiashu, et al.
Publicado: (2025)
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
por: Zheng, Haizhong, et al.
Publicado: (2026)
por: Zheng, Haizhong, et al.
Publicado: (2026)
Ejemplares similares
-
Scaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspaces
por: Gupta, Karan, et al.
Publicado: (2026) -
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
por: Singh, Joykirat, et al.
Publicado: (2024) -
PromptWizard: Task-Aware Prompt Optimization Framework
por: Agarwal, Eshaan, et al.
Publicado: (2024) -
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
por: Singh, Joykirat, et al.
Publicado: (2025) -
Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
por: Agarwal, Aradhye, et al.
Publicado: (2026)