LATTE: Learning to Think with Vision Specialists

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Zixian, Zhang, Jianguo, Liu, Zhiwei, Zhang, Jieyu, Tan, Juntao, Shu, Manli, Niebles, Juan Carlos, Heinecke, Shelby, Wang, Huan, Xiong, Caiming, Krishna, Ranjay, Savarese, Silvio
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916949392359424
author Ma, Zixian
Zhang, Jianguo
Liu, Zhiwei
Zhang, Jieyu
Tan, Juntao
Shu, Manli
Niebles, Juan Carlos
Heinecke, Shelby
Wang, Huan
Xiong, Caiming
Krishna, Ranjay
Savarese, Silvio
author_facet Ma, Zixian
Zhang, Jianguo
Liu, Zhiwei
Zhang, Jieyu
Tan, Juntao
Shu, Manli
Niebles, Juan Carlos
Heinecke, Shelby
Wang, Huan
Xiong, Caiming
Krishna, Ranjay
Savarese, Silvio
contents While open-source vision-language models perform well on simple question-answering, they still struggle with complex questions that require both perceptual and reasoning capabilities. We propose LATTE, a family of vision-language models that have LeArned to Think wiTh vision spEcialists. By offloading perception to state-of-the-art vision models, our approach enables vision-language models to focus solely on reasoning over high-quality perceptual information. To train LATTE, we synthesize and filter a large dataset of 293K multi-modal reasoning traces over perceptual outputs of vision specialists. LATTE trained on this data achieves significant 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. Ablation studies reveal that the effectiveness of multi-modal reasoning traces depends on the data sources, formats, and quality of thoughts.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LATTE: Learning to Think with Vision Specialists
Ma, Zixian
Zhang, Jianguo
Liu, Zhiwei
Zhang, Jieyu
Tan, Juntao
Shu, Manli
Niebles, Juan Carlos
Heinecke, Shelby
Wang, Huan
Xiong, Caiming
Krishna, Ranjay
Savarese, Silvio
Computer Vision and Pattern Recognition
While open-source vision-language models perform well on simple question-answering, they still struggle with complex questions that require both perceptual and reasoning capabilities. We propose LATTE, a family of vision-language models that have LeArned to Think wiTh vision spEcialists. By offloading perception to state-of-the-art vision models, our approach enables vision-language models to focus solely on reasoning over high-quality perceptual information. To train LATTE, we synthesize and filter a large dataset of 293K multi-modal reasoning traces over perceptual outputs of vision specialists. LATTE trained on this data achieves significant 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. Ablation studies reveal that the effectiveness of multi-modal reasoning traces depends on the data sources, formats, and quality of thoughts.
title LATTE: Learning to Think with Vision Specialists
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.05479