A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Qianben, Cao, Jingyi, Zhang, Jiayu, Qin, Tianrui, Li, Xiaowan, Zhu, King, Shi, Dingfeng, Zhu, He, Liu, Minghao, Liang, Xiaobo, Gui, Xin, Zhang, Ge, Yang, Jian, Jiang, Yuchen Eleanor, Zhou, Wangchunshu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912662285189120
author Chen, Qianben
Cao, Jingyi
Zhang, Jiayu
Qin, Tianrui
Li, Xiaowan
Zhu, King
Shi, Dingfeng
Zhu, He
Liu, Minghao
Liang, Xiaobo
Gui, Xin
Zhang, Ge
Yang, Jian
Jiang, Yuchen Eleanor
Zhou, Wangchunshu
author_facet Chen, Qianben
Cao, Jingyi
Zhang, Jiayu
Qin, Tianrui
Li, Xiaowan
Zhu, King
Shi, Dingfeng
Zhu, He
Liu, Minghao
Liang, Xiaobo
Gui, Xin
Zhang, Ge
Yang, Jian
Jiang, Yuchen Eleanor
Zhou, Wangchunshu
contents Large language models split into two families: reasoning-centric LLMs, which strengthen internal chain-of-thought reasoning but cannot invoke external tools, and agentic LLMs, which learn to interact with environments and leverage tools but often lag in deep reasoning. This divide arises from fundamentally different training objectives, leading to mismatched strengths and inefficiency on simple queries, where both families tend to overthink or over-call tools. In this work, we present Adaptive Agent Foundation Model (A$^2$FM), a unified framework that follows a route-then-align principle: the model first learns task-aware routing and then aligns mode-specific trajectories under a shared backbone. To address the inefficiency gap, we introduce a third mode-instant-that handles simple queries directly, preventing unnecessary reasoning or tool calls while complementing the agentic and reasoning modes. To jointly enhance accuracy and efficiency, we propose Adaptive Policy Optimization (APO), which enforces adaptive sampling across modes and applies a cost-regularized reward. On the 32B scale, A$^2$FM achieves 13.4% on BrowseComp, 70.4% on AIME25, and 16.7% on HLE, setting new SOTA among comparable models and performing competitively with frontier LLMs across agentic, reasoning, and general benchmarks. Notably, the adaptive execution achieves a cost of pass of only $0.00487 per correct answer-cutting cost by 45.2% relative to reasoning and 33.5% relative to agentic, thus delivering substantially higher cost efficiency while maintaining comparable accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12838
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
Chen, Qianben
Cao, Jingyi
Zhang, Jiayu
Qin, Tianrui
Li, Xiaowan
Zhu, King
Shi, Dingfeng
Zhu, He
Liu, Minghao
Liang, Xiaobo
Gui, Xin
Zhang, Ge
Yang, Jian
Jiang, Yuchen Eleanor
Zhou, Wangchunshu
Computation and Language
Artificial Intelligence
Large language models split into two families: reasoning-centric LLMs, which strengthen internal chain-of-thought reasoning but cannot invoke external tools, and agentic LLMs, which learn to interact with environments and leverage tools but often lag in deep reasoning. This divide arises from fundamentally different training objectives, leading to mismatched strengths and inefficiency on simple queries, where both families tend to overthink or over-call tools. In this work, we present Adaptive Agent Foundation Model (A$^2$FM), a unified framework that follows a route-then-align principle: the model first learns task-aware routing and then aligns mode-specific trajectories under a shared backbone. To address the inefficiency gap, we introduce a third mode-instant-that handles simple queries directly, preventing unnecessary reasoning or tool calls while complementing the agentic and reasoning modes. To jointly enhance accuracy and efficiency, we propose Adaptive Policy Optimization (APO), which enforces adaptive sampling across modes and applies a cost-regularized reward. On the 32B scale, A$^2$FM achieves 13.4% on BrowseComp, 70.4% on AIME25, and 16.7% on HLE, setting new SOTA among comparable models and performing competitively with frontier LLMs across agentic, reasoning, and general benchmarks. Notably, the adaptive execution achieves a cost of pass of only $0.00487 per correct answer-cutting cost by 45.2% relative to reasoning and 33.5% relative to agentic, thus delivering substantially higher cost efficiency while maintaining comparable accuracy.
title A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.12838