Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dong, Yuhao, Liu, Zuyan, Sun, Hai-Long, Yang, Jingkang, Hu, Winston, Rao, Yongming, Liu, Ziwei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912357370822656
author Dong, Yuhao
Liu, Zuyan
Sun, Hai-Long
Yang, Jingkang
Hu, Winston
Rao, Yongming
Liu, Ziwei
author_facet Dong, Yuhao
Liu, Zuyan
Sun, Hai-Long
Yang, Jingkang
Hu, Winston
Rao, Yongming
Liu, Ziwei
contents Large Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought prompting to product-level solutions like OpenAI o1. Despite various efforts to improve LLM reasoning, high-quality long-chain reasoning data and optimized training pipelines still remain inadequately explored in vision-language tasks. In this paper, we present Insight-V, an early effort to 1) scalably produce long and robust reasoning data for complex multi-modal tasks, and 2) an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs). Specifically, to create long and structured reasoning data without human labor, we design a two-step pipeline with a progressive strategy to generate sufficiently long and diverse reasoning paths and a multi-granularity assessment method to ensure data quality. We observe that directly supervising MLLMs with such long and complex reasoning data will not yield ideal reasoning ability. To tackle this problem, we design a multi-agent system consisting of a reasoning agent dedicated to performing long-chain reasoning and a summary agent trained to judge and summarize reasoning results. We further incorporate an iterative DPO algorithm to enhance the reasoning agent's generation stability and quality. Based on the popular LLaVA-NeXT model and our stronger base MLLM, we demonstrate significant performance gains across challenging multi-modal benchmarks requiring visual reasoning. Benefiting from our multi-agent system, Insight-V can also easily maintain or improve performance on perception-focused multi-modal tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14432
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Dong, Yuhao
Liu, Zuyan
Sun, Hai-Long
Yang, Jingkang
Hu, Winston
Rao, Yongming
Liu, Ziwei
Computer Vision and Pattern Recognition
Large Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought prompting to product-level solutions like OpenAI o1. Despite various efforts to improve LLM reasoning, high-quality long-chain reasoning data and optimized training pipelines still remain inadequately explored in vision-language tasks. In this paper, we present Insight-V, an early effort to 1) scalably produce long and robust reasoning data for complex multi-modal tasks, and 2) an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs). Specifically, to create long and structured reasoning data without human labor, we design a two-step pipeline with a progressive strategy to generate sufficiently long and diverse reasoning paths and a multi-granularity assessment method to ensure data quality. We observe that directly supervising MLLMs with such long and complex reasoning data will not yield ideal reasoning ability. To tackle this problem, we design a multi-agent system consisting of a reasoning agent dedicated to performing long-chain reasoning and a summary agent trained to judge and summarize reasoning results. We further incorporate an iterative DPO algorithm to enhance the reasoning agent's generation stability and quality. Based on the popular LLaVA-NeXT model and our stronger base MLLM, we demonstrate significant performance gains across challenging multi-modal benchmarks requiring visual reasoning. Benefiting from our multi-agent system, Insight-V can also easily maintain or improve performance on perception-focused multi-modal tasks.
title Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.14432