Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Bingshuai, Lyu, Chenyang, Min, Zijun, Wang, Zhanyu, Su, Jinsong, Wang, Longyue
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917602236825600
author Liu, Bingshuai
Lyu, Chenyang
Min, Zijun
Wang, Zhanyu
Su, Jinsong
Wang, Longyue
author_facet Liu, Bingshuai
Lyu, Chenyang
Min, Zijun
Wang, Zhanyu
Su, Jinsong
Wang, Longyue
contents The advancement of Large Language Models (LLMs) has brought substantial attention to the Chain of Thought (CoT) approach, primarily due to its ability to enhance the capability of LLMs on complex reasoning tasks. Moreover, the significance of CoT approaches extends to the application of LLMs for multi-modal tasks. However, the selection of optimal CoT demonstration examples in multi-modal reasoning remains less explored for LLMs due to the inherent complexity of multi-modal examples. In this paper, we introduce a novel approach that addresses this challenge by using retrieval mechanisms to dynamically and automatically select demonstration examples based on cross-modal and intra-modal similarities. Furthermore, we employ a Stratified Sampling method of categorising demonstration examples into groups based on their types and then retrieving examples from different groups respectively to promote the diversity of demonstration examples. Through a series of experiments on two popular benchmark datasets: ScienceQA and MathVista, we demonstrate that our approach significantly improves the performance of GPT-4 by 6% on ScienceQA and 12.9% on MathVista, and enhances the performance of GPT-4V on two datasets by 2.7%, substantially improving the performance of the most advanced LLMs and LMMs for complex multi-modal reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_01714
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models
Liu, Bingshuai
Lyu, Chenyang
Min, Zijun
Wang, Zhanyu
Su, Jinsong
Wang, Longyue
Computation and Language
The advancement of Large Language Models (LLMs) has brought substantial attention to the Chain of Thought (CoT) approach, primarily due to its ability to enhance the capability of LLMs on complex reasoning tasks. Moreover, the significance of CoT approaches extends to the application of LLMs for multi-modal tasks. However, the selection of optimal CoT demonstration examples in multi-modal reasoning remains less explored for LLMs due to the inherent complexity of multi-modal examples. In this paper, we introduce a novel approach that addresses this challenge by using retrieval mechanisms to dynamically and automatically select demonstration examples based on cross-modal and intra-modal similarities. Furthermore, we employ a Stratified Sampling method of categorising demonstration examples into groups based on their types and then retrieving examples from different groups respectively to promote the diversity of demonstration examples. Through a series of experiments on two popular benchmark datasets: ScienceQA and MathVista, we demonstrate that our approach significantly improves the performance of GPT-4 by 6% on ScienceQA and 12.9% on MathVista, and enhances the performance of GPT-4V on two datasets by 2.7%, substantially improving the performance of the most advanced LLMs and LMMs for complex multi-modal reasoning tasks.
title Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2312.01714