LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Guowei, Jin, Peng, Wu, Ziang, Li, Hao, Song, Yibing, Sun, Lichao, Yuan, Li
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911065288212480
author Xu, Guowei
Jin, Peng
Wu, Ziang
Li, Hao
Song, Yibing
Sun, Lichao
Yuan, Li
author_facet Xu, Guowei
Jin, Peng
Wu, Ziang
Li, Hao
Song, Yibing
Sun, Lichao
Yuan, Li
contents Large language models have demonstrated substantial advancements in reasoning capabilities. However, current Vision-Language Models (VLMs) often struggle to perform systematic and structured reasoning, especially when handling complex visual question-answering tasks. In this work, we introduce LLaVA-CoT, a large VLM designed to conduct autonomous multistage reasoning. Unlike chain-of-thought prompting, LLaVA-CoT independently engages in sequential stages of summarization, visual interpretation, logical reasoning, and conclusion generation. This structured approach enables LLaVA-CoT to achieve marked improvements on reasoning-intensive tasks. To accomplish this, we construct the LLaVA-CoT-100k dataset, integrating samples from various visual question answering sources and providing structured reasoning annotations. Besides, we propose a test-time stage-wise retracing search method (SWIRES), which enables effective and efficient test-time scaling. Remarkably, with only 100k training samples and test-time scaling, LLaVA-CoT not only outperforms its base model by 9.4% on a wide range of multimodal reasoning benchmarks, but also surpasses the performance of larger and even closed-source models, such as Gemini-1.5-pro, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct. The code, dataset, and pre-trained weights are publicly available at https://github.com/PKU-YuanGroup/LLaVA-CoT.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10440
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Xu, Guowei
Jin, Peng
Wu, Ziang
Li, Hao
Song, Yibing
Sun, Lichao
Yuan, Li
Computer Vision and Pattern Recognition
Large language models have demonstrated substantial advancements in reasoning capabilities. However, current Vision-Language Models (VLMs) often struggle to perform systematic and structured reasoning, especially when handling complex visual question-answering tasks. In this work, we introduce LLaVA-CoT, a large VLM designed to conduct autonomous multistage reasoning. Unlike chain-of-thought prompting, LLaVA-CoT independently engages in sequential stages of summarization, visual interpretation, logical reasoning, and conclusion generation. This structured approach enables LLaVA-CoT to achieve marked improvements on reasoning-intensive tasks. To accomplish this, we construct the LLaVA-CoT-100k dataset, integrating samples from various visual question answering sources and providing structured reasoning annotations. Besides, we propose a test-time stage-wise retracing search method (SWIRES), which enables effective and efficient test-time scaling. Remarkably, with only 100k training samples and test-time scaling, LLaVA-CoT not only outperforms its base model by 9.4% on a wide range of multimodal reasoning benchmarks, but also surpasses the performance of larger and even closed-source models, such as Gemini-1.5-pro, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct. The code, dataset, and pre-trained weights are publicly available at https://github.com/PKU-YuanGroup/LLaVA-CoT.
title LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.10440