Reasoning through Exploration: A Reinforcement Learning Framework for Robust Function Calling
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Bingguang, Xu, Zengzhuang, Wang, Maolin, Wen, Yuntao, Chen, Yicheng, Peng, Cunyin, Chen, Long, Wang, Dong, Zhao, Xiangyu, Gu, Jinjie, Zhuang, Chenyi, Zhang, Ji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement
by: Hao, Bingguang, et al.
Published: (2025)
by: Hao, Bingguang, et al.
Published: (2025)
FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
by: Xu, Zengzhuang, et al.
Published: (2025)
by: Xu, Zengzhuang, et al.
Published: (2025)
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
by: Hao, Bingguang, et al.
Published: (2026)
by: Hao, Bingguang, et al.
Published: (2026)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026)
by: Bai, Haoyue, et al.
Published: (2026)
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
by: Chen, Jikai, et al.
Published: (2026)
by: Chen, Jikai, et al.
Published: (2026)
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
by: Chen, Jikai, et al.
Published: (2025)
by: Chen, Jikai, et al.
Published: (2025)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
by: Wang, Maolin, et al.
Published: (2023)
by: Wang, Maolin, et al.
Published: (2023)
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
by: He, Kaiwen, et al.
Published: (2025)
by: He, Kaiwen, et al.
Published: (2025)
DNS-Rec: Data-aware Neural Architecture Search for Recommender Systems
by: Zhang, Sheng, et al.
Published: (2024)
by: Zhang, Sheng, et al.
Published: (2024)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
by: Tan, Zhiwen, et al.
Published: (2025)
by: Tan, Zhiwen, et al.
Published: (2025)
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
by: Xie, Zhitian, et al.
Published: (2025)
by: Xie, Zhitian, et al.
Published: (2025)
Explainable Behavior Cloning: Teaching Large Language Model Agents through Learning by Demonstration
by: Guan, Yanchu, et al.
Published: (2024)
by: Guan, Yanchu, et al.
Published: (2024)
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy
by: Zhao, Yao, et al.
Published: (2023)
by: Zhao, Yao, et al.
Published: (2023)
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM
by: Yu, Chengyue, et al.
Published: (2024)
by: Yu, Chengyue, et al.
Published: (2024)
LiveAgentBench: Comprehensive Benchmarking of Agentic Systems Across 104 Real-World Challenges
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation
by: Zhang, Hongxuan, et al.
Published: (2024)
by: Zhang, Hongxuan, et al.
Published: (2024)
Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
by: Gao, Jingtong, et al.
Published: (2025)
by: Gao, Jingtong, et al.
Published: (2025)
Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster
by: Zhang, Hongxuan, et al.
Published: (2023)
by: Zhang, Hongxuan, et al.
Published: (2023)
MetaLoRA: Tensor-Enhanced Adaptive Low-Rank Fine-tuning
by: Wang, Maolin, et al.
Published: (2025)
by: Wang, Maolin, et al.
Published: (2025)
Mitigate Position Bias with Coupled Ranking Bias on CTR Prediction
by: Zhao, Yao, et al.
Published: (2024)
by: Zhao, Yao, et al.
Published: (2024)
Tensorized Hypergraph Neural Networks
by: Wang, Maolin, et al.
Published: (2023)
by: Wang, Maolin, et al.
Published: (2023)
Reinforced Preference Optimization for Reasoning-Augmented Recommendations
by: Gao, Jingtong, et al.
Published: (2026)
by: Gao, Jingtong, et al.
Published: (2026)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
by: Peng, Jingyu, et al.
Published: (2025)
by: Peng, Jingyu, et al.
Published: (2025)
Hammer: Robust Function-Calling for On-Device Language Models via Function Masking
by: Lin, Qiqiang, et al.
Published: (2024)
by: Lin, Qiqiang, et al.
Published: (2024)
PowLU: An Activation Function for Stable Pre-Training of LLMs
by: Jiang, Peijie, et al.
Published: (2026)
by: Jiang, Peijie, et al.
Published: (2026)
Don't Just Fine-tune the Agent, Tune the Environment
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts
by: Xie, Zhitian, et al.
Published: (2024)
by: Xie, Zhitian, et al.
Published: (2024)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
by: Chai, Huacan, et al.
Published: (2025)
by: Chai, Huacan, et al.
Published: (2025)
Learning Agile and Robust Omnidirectional Aerial Motion on Overactuated Tiltable-Quadrotors
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
by: Ye, Zhiling, et al.
Published: (2025)
by: Ye, Zhiling, et al.
Published: (2025)
Calling in crisis: How intolerance of uncertainty shaped occupational calling before and during the pandemic
by: Qing Yang, et al.
Published: (2025)
by: Qing Yang, et al.
Published: (2025)
SQLCritic: Correcting Text-to-SQL Generation via Clause-wise Critic
by: Chen, Jikai, et al.
Published: (2025)
by: Chen, Jikai, et al.
Published: (2025)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
by: Wei, Lei, et al.
Published: (2026)
by: Wei, Lei, et al.
Published: (2026)
Stepwise Reasoning Error Disruption Attack of LLMs
by: Peng, Jingyu, et al.
Published: (2024)
by: Peng, Jingyu, et al.
Published: (2024)
Know Your Needs Better: Towards Structured Understanding of Marketer Demands with Analogical Reasoning Augmented LLMs
by: Wang, Junjie, et al.
Published: (2024)
by: Wang, Junjie, et al.
Published: (2024)
BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
by: Lu, Bingguang, et al.
Published: (2025)
by: Lu, Bingguang, et al.
Published: (2025)
Assessing Surface Water Hydrological Connectivity and Spatiotemporal Evolution in Xinjiang (2000–2020)
by: Yue Ding, et al.
Published: (2025)
by: Yue Ding, et al.
Published: (2025)
GIFS: Neural Implicit Function for General Shape Representation
by: Ye, Jianglong, et al.
Published: (2022)
by: Ye, Jianglong, et al.
Published: (2022)
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning
by: Liu, Huanyu, et al.
Published: (2025)
by: Liu, Huanyu, et al.
Published: (2025)
On the Robustness of Agentic Function Calling
by: Rabinovich, Ella, et al.
Published: (2025)
by: Rabinovich, Ella, et al.
Published: (2025)
Similar Items
-
FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement
by: Hao, Bingguang, et al.
Published: (2025) -
FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
by: Xu, Zengzhuang, et al.
Published: (2025) -
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
by: Hao, Bingguang, et al.
Published: (2026) -
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026) -
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
by: Chen, Jikai, et al.
Published: (2026)