VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zejun, Luo, Ruipu, Zhang, Jiwen, Qiu, Minghui, Huang, Xuanjing, Wei, Zhongyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929747985956864
author Li, Zejun
Luo, Ruipu
Zhang, Jiwen
Qiu, Minghui
Huang, Xuanjing
Wei, Zhongyu
author_facet Li, Zejun
Luo, Ruipu
Zhang, Jiwen
Qiu, Minghui
Huang, Xuanjing
Wei, Zhongyu
contents While large multi-modal models (LMMs) have exhibited impressive capabilities across diverse tasks, their effectiveness in handling complex tasks has been limited by the prevailing single-step reasoning paradigm. To this end, this paper proposes VoCoT, a multi-step Visually grounded object-centric Chain-of-Thought reasoning framework tailored for inference with LMMs. VoCoT is characterized by two key features: (1) object-centric reasoning paths that revolve around cross-modal shared object-level information, and (2) visually grounded representation of object concepts in a multi-modal interleaved and aligned manner, which effectively bridges the modality gap within LMMs during long-term generation. To adapt LMMs in reasoning with VoCoT, we further construct an instruction-tuning dataset. By combining VoCoT with the prevalent open-source LMM architectures, we develop a VoCoT-based model, VolCano. With only 7B parameters and limited input image resolution, VolCano demonstrates excellent performance across various scenarios. In benchmarks like CLEVR and EmbSpatial, which highly require complex reasoning capabilities, VolCano outperforms SOTA models, including powerful GPT-4V. Related code, data and models are released in https://github.com/RupertLuo/VoCoT.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16919
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
Li, Zejun
Luo, Ruipu
Zhang, Jiwen
Qiu, Minghui
Huang, Xuanjing
Wei, Zhongyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
While large multi-modal models (LMMs) have exhibited impressive capabilities across diverse tasks, their effectiveness in handling complex tasks has been limited by the prevailing single-step reasoning paradigm. To this end, this paper proposes VoCoT, a multi-step Visually grounded object-centric Chain-of-Thought reasoning framework tailored for inference with LMMs. VoCoT is characterized by two key features: (1) object-centric reasoning paths that revolve around cross-modal shared object-level information, and (2) visually grounded representation of object concepts in a multi-modal interleaved and aligned manner, which effectively bridges the modality gap within LMMs during long-term generation. To adapt LMMs in reasoning with VoCoT, we further construct an instruction-tuning dataset. By combining VoCoT with the prevalent open-source LMM architectures, we develop a VoCoT-based model, VolCano. With only 7B parameters and limited input image resolution, VolCano demonstrates excellent performance across various scenarios. In benchmarks like CLEVR and EmbSpatial, which highly require complex reasoning capabilities, VolCano outperforms SOTA models, including powerful GPT-4V. Related code, data and models are released in https://github.com/RupertLuo/VoCoT.
title VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.16919