AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xiang, Kun, Liu, Zhili, Zhang, Terry Jingchen, Huang, Yinya, Nie, Yunshuang, Cai, Kaixin, Yin, Yiyang, Huang, Runhui, Li, Hanhui, Zeng, Yihan, Yuan, Yu-Jie, Han, Jianhua, Hong, Lanqing, Xu, Hang, Liang, Xiaodan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908754184765440
author Xiang, Kun
Liu, Zhili
Zhang, Terry Jingchen
Huang, Yinya
Nie, Yunshuang
Cai, Kaixin
Yin, Yiyang
Huang, Runhui
Li, Hanhui
Zeng, Yihan
Yuan, Yu-Jie
Han, Jianhua
Hong, Lanqing
Xu, Hang
Liang, Xiaodan
author_facet Xiang, Kun
Liu, Zhili
Zhang, Terry Jingchen
Huang, Yinya
Nie, Yunshuang
Cai, Kaixin
Yin, Yiyang
Huang, Runhui
Li, Hanhui
Zeng, Yihan
Yuan, Yu-Jie
Han, Jianhua
Hong, Lanqing
Xu, Hang
Liang, Xiaodan
contents In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of ``slow thinking'' into multimodal large language models (MLLMs). Our core idea is that models can learn to adaptively use different levels of reasoning to tackle questions of varying complexity. We propose a novel paradigm of Self-structured Chain of Thought (SCoT), which consists of minimal semantic atomic steps. Unlike existing methods that rely on structured templates or free-form paradigms, our method not only generates flexible CoT structures for various complex tasks but also mitigates the phenomenon of overthinking for easier tasks. To introduce structured reasoning into visual cognition, we design a novel AtomThink framework with four key modules: (i) a data engine to generate high-quality multimodal reasoning paths; (ii) a supervised fine-tuning (SFT) process with serialized inference data; (iii) a policy-guided multi-turn inference method; and (iv) an atomic capability metric to evaluate the single-step utilization rate. Extensive experiments demonstrate that the proposed AtomThink significantly improves the performance of baseline MLLMs, achieving more than 10\% average accuracy gains on MathVista and MathVerse. Compared to state-of-the-art structured CoT approaches, our method not only achieves higher accuracy but also improves data utilization by 5 $\times$ and boosts inference efficiency by 85.3\%. Our code is publicly available at https://github.com/Kun-Xiang/AtomThink.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11930
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
Xiang, Kun
Liu, Zhili
Zhang, Terry Jingchen
Huang, Yinya
Nie, Yunshuang
Cai, Kaixin
Yin, Yiyang
Huang, Runhui
Li, Hanhui
Zeng, Yihan
Yuan, Yu-Jie
Han, Jianhua
Hong, Lanqing
Xu, Hang
Liang, Xiaodan
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of ``slow thinking'' into multimodal large language models (MLLMs). Our core idea is that models can learn to adaptively use different levels of reasoning to tackle questions of varying complexity. We propose a novel paradigm of Self-structured Chain of Thought (SCoT), which consists of minimal semantic atomic steps. Unlike existing methods that rely on structured templates or free-form paradigms, our method not only generates flexible CoT structures for various complex tasks but also mitigates the phenomenon of overthinking for easier tasks. To introduce structured reasoning into visual cognition, we design a novel AtomThink framework with four key modules: (i) a data engine to generate high-quality multimodal reasoning paths; (ii) a supervised fine-tuning (SFT) process with serialized inference data; (iii) a policy-guided multi-turn inference method; and (iv) an atomic capability metric to evaluate the single-step utilization rate. Extensive experiments demonstrate that the proposed AtomThink significantly improves the performance of baseline MLLMs, achieving more than 10\% average accuracy gains on MathVista and MathVerse. Compared to state-of-the-art structured CoT approaches, our method not only achieves higher accuracy but also improves data utilization by 5 $\times$ and boosts inference efficiency by 85.3\%. Our code is publicly available at https://github.com/Kun-Xiang/AtomThink.
title AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.11930