AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qian, Kangan, Jiang, Sicong, Zhong, Yang, Luo, Ziang, Huang, Zilin, Zhu, Tianze, Jiang, Kun, Yang, Mengmeng, Fu, Zheng, Miao, Jinyu, Shi, Yining, Lim, He Zhe, Liu, Li, Zhou, Tianbao, Yu, Huang, Hu, Yifei, Li, Guang, Chen, Guang, Ye, Hao, Sun, Lijun, Yang, Diange
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912610665889792
author Qian, Kangan
Jiang, Sicong
Zhong, Yang
Luo, Ziang
Huang, Zilin
Zhu, Tianze
Jiang, Kun
Yang, Mengmeng
Fu, Zheng
Miao, Jinyu
Shi, Yining
Lim, He Zhe
Liu, Li
Zhou, Tianbao
Yu, Huang
Hu, Yifei
Li, Guang
Chen, Guang
Ye, Hao
Sun, Lijun
Yang, Diange
author_facet Qian, Kangan
Jiang, Sicong
Zhong, Yang
Luo, Ziang
Huang, Zilin
Zhu, Tianze
Jiang, Kun
Yang, Mengmeng
Fu, Zheng
Miao, Jinyu
Shi, Yining
Lim, He Zhe
Liu, Li
Zhou, Tianbao
Yu, Huang
Hu, Yifei
Li, Guang
Chen, Guang
Ye, Hao
Sun, Lijun
Yang, Diange
contents Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. To overcome this, we introduce \textbf{AgentThink}, a pioneering unified framework that integrates Chain-of-Thought (CoT) reasoning with dynamic, agent-style tool invocation for autonomous driving tasks. AgentThink's core innovations include: \textbf{(i) Structured Data Generation}, which establishes an autonomous driving tool library to automatically construct structured, self-verified reasoning data explicitly incorporating tool usage for diverse driving scenarios; \textbf{(ii) A Two-stage Training Pipeline}, employing Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO) to equip VLMs with the capability for autonomous tool invocation; and \textbf{(iii) Agent-style Tool-Usage Evaluation}, introducing a novel multi-tool assessment protocol to rigorously evaluate the model's tool invocation and utilization. Experiments on the DriveLMM-o1 benchmark demonstrate that AgentThink significantly boosts overall reasoning scores by \textbf{53.91%} and enhances answer accuracy by \textbf{33.54%}, while markedly improving reasoning quality and consistency. Furthermore, ablation studies and robust zero-shot/few-shot generalization experiments across various benchmarks underscore its powerful capabilities. These findings highlight a promising trajectory for developing trustworthy and tool-aware autonomous driving models. Code is available at https://github.com/curryqka/AgentThink.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
Qian, Kangan
Jiang, Sicong
Zhong, Yang
Luo, Ziang
Huang, Zilin
Zhu, Tianze
Jiang, Kun
Yang, Mengmeng
Fu, Zheng
Miao, Jinyu
Shi, Yining
Lim, He Zhe
Liu, Li
Zhou, Tianbao
Yu, Huang
Hu, Yifei
Li, Guang
Chen, Guang
Ye, Hao
Sun, Lijun
Yang, Diange
Robotics
Computation and Language
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. To overcome this, we introduce \textbf{AgentThink}, a pioneering unified framework that integrates Chain-of-Thought (CoT) reasoning with dynamic, agent-style tool invocation for autonomous driving tasks. AgentThink's core innovations include: \textbf{(i) Structured Data Generation}, which establishes an autonomous driving tool library to automatically construct structured, self-verified reasoning data explicitly incorporating tool usage for diverse driving scenarios; \textbf{(ii) A Two-stage Training Pipeline}, employing Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO) to equip VLMs with the capability for autonomous tool invocation; and \textbf{(iii) Agent-style Tool-Usage Evaluation}, introducing a novel multi-tool assessment protocol to rigorously evaluate the model's tool invocation and utilization. Experiments on the DriveLMM-o1 benchmark demonstrate that AgentThink significantly boosts overall reasoning scores by \textbf{53.91%} and enhances answer accuracy by \textbf{33.54%}, while markedly improving reasoning quality and consistency. Furthermore, ablation studies and robust zero-shot/few-shot generalization experiments across various benchmarks underscore its powerful capabilities. These findings highlight a promising trajectory for developing trustworthy and tool-aware autonomous driving models. Code is available at https://github.com/curryqka/AgentThink.
title AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.15298