AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Yuanjie, Wang, Chengyu, Zheng, Haonan, Yue, Yuanhao, Yan, Junbing, Wang, Ming, Huang, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards
by: Lyu, Yuanjie, et al.
Published: (2026)
by: Lyu, Yuanjie, et al.
Published: (2026)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
by: Yue, Yuanhao, et al.
Published: (2026)
by: Yue, Yuanhao, et al.
Published: (2026)
DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
by: Chen, Aili, et al.
Published: (2026)
by: Chen, Aili, et al.
Published: (2026)
Do Large Language Models Understand Logic or Just Mimick Context?
by: Yan, Junbing, et al.
Published: (2024)
by: Yan, Junbing, et al.
Published: (2024)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
From Correction to Mastery: Reinforced Distillation of Large Language Model Agents
by: Lyu, Yuanjie, et al.
Published: (2025)
by: Lyu, Yuanjie, et al.
Published: (2025)
Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
Distilling Instruction-following Abilities of Large Language Models with Task-aware Curriculum Planning
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
Agentic Tool Use in Large Language Models
by: Hu, Jinchao, et al.
Published: (2026)
by: Hu, Jinchao, et al.
Published: (2026)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
by: Jiang, Dongfu, et al.
Published: (2025)
by: Jiang, Dongfu, et al.
Published: (2025)
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
by: Tan, Rongbin, et al.
Published: (2026)
by: Tan, Rongbin, et al.
Published: (2026)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
by: Xu, Siyuan, et al.
Published: (2026)
by: Xu, Siyuan, et al.
Published: (2026)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
ASTER: Agentic Scaling with Tool-integrated Extended Reasoning
by: Zhang, Xuqin, et al.
Published: (2026)
by: Zhang, Xuqin, et al.
Published: (2026)
SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
by: Jiang, Yanna, et al.
Published: (2026)
by: Jiang, Yanna, et al.
Published: (2026)
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
by: Wang, Shaokang, et al.
Published: (2026)
by: Wang, Shaokang, et al.
Published: (2026)
ToolRM: Towards Agentic Tool-Use Reward Modeling
by: Li, Renhao, et al.
Published: (2025)
by: Li, Renhao, et al.
Published: (2025)
MerLean: An Agentic Framework for Autoformalization in Quantum Computation
by: Ren, Yuanjie, et al.
Published: (2026)
by: Ren, Yuanjie, et al.
Published: (2026)
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
by: Yan, Shilin, et al.
Published: (2026)
by: Yan, Shilin, et al.
Published: (2026)
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Budget-Constrained Agentic Large Language Models: Intention-Based Planning for Costly Tool Use
by: Liu, Hanbing, et al.
Published: (2026)
by: Liu, Hanbing, et al.
Published: (2026)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning
by: Su, Jiaming, et al.
Published: (2026)
by: Su, Jiaming, et al.
Published: (2026)
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
by: Gu, Yi, et al.
Published: (2026)
by: Gu, Yi, et al.
Published: (2026)
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
by: Liu, Tong, et al.
Published: (2026)
by: Liu, Tong, et al.
Published: (2026)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Agentic Critical Training
by: Liu, Weize, et al.
Published: (2026)
by: Liu, Weize, et al.
Published: (2026)
On the Role of Long-tail Knowledge in Retrieval Augmented Large Language Models
by: Li, Dongyang, et al.
Published: (2024)
by: Li, Dongyang, et al.
Published: (2024)
TRELM: Towards Robust and Efficient Pre-training for Knowledge-Enhanced Language Models
by: Yan, Junbing, et al.
Published: (2024)
by: Yan, Junbing, et al.
Published: (2024)
Green R&D Efficiency and China's Industrial Carbon Intensity: New Evidence From a Spatial Perspective
by: Changhua Chen, et al.
Published: (2025)
by: Changhua Chen, et al.
Published: (2025)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
6GAgentGym: Tool Use, Data Synthesis, and Agentic Learning for Network Management
by: Chen, Jiao, et al.
Published: (2026)
by: Chen, Jiao, et al.
Published: (2026)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
by: Zhang, Yabo, et al.
Published: (2025)
by: Zhang, Yabo, et al.
Published: (2025)
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
by: Deng, Mengjie, et al.
Published: (2025)
by: Deng, Mengjie, et al.
Published: (2025)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
by: Lei, Fei, et al.
Published: (2025)
by: Lei, Fei, et al.
Published: (2025)
Similar Items
-
DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
by: Wang, Chengyu, et al.
Published: (2025) -
Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards
by: Lyu, Yuanjie, et al.
Published: (2026) -
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025) -
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
by: Cai, Wenrui, et al.
Published: (2025) -
OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models
by: Yue, Yuanhao, et al.
Published: (2026)