LATTE: Learning to Think with Vision Specialists
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916949392359424 |
|---|---|
| author | Ma, Zixian Zhang, Jianguo Liu, Zhiwei Zhang, Jieyu Tan, Juntao Shu, Manli Niebles, Juan Carlos Heinecke, Shelby Wang, Huan Xiong, Caiming Krishna, Ranjay Savarese, Silvio |
| author_facet | Ma, Zixian Zhang, Jianguo Liu, Zhiwei Zhang, Jieyu Tan, Juntao Shu, Manli Niebles, Juan Carlos Heinecke, Shelby Wang, Huan Xiong, Caiming Krishna, Ranjay Savarese, Silvio |
| contents | While open-source vision-language models perform well on simple question-answering, they still struggle with complex questions that require both perceptual and reasoning capabilities. We propose LATTE, a family of vision-language models that have LeArned to Think wiTh vision spEcialists. By offloading perception to state-of-the-art vision models, our approach enables vision-language models to focus solely on reasoning over high-quality perceptual information. To train LATTE, we synthesize and filter a large dataset of 293K multi-modal reasoning traces over perceptual outputs of vision specialists. LATTE trained on this data achieves significant 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. Ablation studies reveal that the effectiveness of multi-modal reasoning traces depends on the data sources, formats, and quality of thoughts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_05479 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LATTE: Learning to Think with Vision Specialists Ma, Zixian Zhang, Jianguo Liu, Zhiwei Zhang, Jieyu Tan, Juntao Shu, Manli Niebles, Juan Carlos Heinecke, Shelby Wang, Huan Xiong, Caiming Krishna, Ranjay Savarese, Silvio Computer Vision and Pattern Recognition While open-source vision-language models perform well on simple question-answering, they still struggle with complex questions that require both perceptual and reasoning capabilities. We propose LATTE, a family of vision-language models that have LeArned to Think wiTh vision spEcialists. By offloading perception to state-of-the-art vision models, our approach enables vision-language models to focus solely on reasoning over high-quality perceptual information. To train LATTE, we synthesize and filter a large dataset of 293K multi-modal reasoning traces over perceptual outputs of vision specialists. LATTE trained on this data achieves significant 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. Ablation studies reveal that the effectiveness of multi-modal reasoning traces depends on the data sources, formats, and quality of thoughts. |
| title | LATTE: Learning to Think with Vision Specialists |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.05479 |