Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Mingyuan, Jiang, Jize, Zheng, Haozhen, Li, Meitang, Li, Zhaoheng, Tian, Beitong, Chen, Bo, Park, Yongjoo, Zhang, Minjia, Zhai, Chengxiang, Nahrstedt, Klara
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915502891204608
author Wu, Mingyuan
Jiang, Jize
Zheng, Haozhen
Li, Meitang
Li, Zhaoheng
Tian, Beitong
Chen, Bo
Park, Yongjoo
Zhang, Minjia
Zhai, Chengxiang
Nahrstedt, Klara
author_facet Wu, Mingyuan
Jiang, Jize
Zheng, Haozhen
Li, Meitang
Li, Zhaoheng
Tian, Beitong
Chen, Bo
Park, Yongjoo
Zhang, Minjia
Zhai, Chengxiang
Nahrstedt, Klara
contents Vision Language Models (VLMs) have achieved remarkable success in a wide range of vision applications of increasing complexity and scales, yet choosing the right VLM model size involves a trade-off between response quality and cost. While smaller VLMs are cheaper to run, they typically produce responses only marginally better than random guessing on benchmarks such as MMMU. In this paper, we propose Cache of Thought (CoT), a master apprentice framework for collaborative inference between large and small VLMs. CoT manages high quality query results from large VLMs (master) in a cache, which are then selected via a novel multi modal retrieval and in-context learning to aid the performance of small VLMs (apprentice). We extensively evaluate CoT on various widely recognized and challenging general reasoning benchmarks, and show that CoT increases overall reasoning performance by up to 7.7% under the same budget, and specifically boosts the performance of apprentice VLMs by up to 36.6%. Our code is available at https://github.com/UIUC-MONET/Cache-of-Thoughts
format Preprint
id arxiv_https___arxiv_org_abs_2502_20587
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
Wu, Mingyuan
Jiang, Jize
Zheng, Haozhen
Li, Meitang
Li, Zhaoheng
Tian, Beitong
Chen, Bo
Park, Yongjoo
Zhang, Minjia
Zhai, Chengxiang
Nahrstedt, Klara
Machine Learning
Vision Language Models (VLMs) have achieved remarkable success in a wide range of vision applications of increasing complexity and scales, yet choosing the right VLM model size involves a trade-off between response quality and cost. While smaller VLMs are cheaper to run, they typically produce responses only marginally better than random guessing on benchmarks such as MMMU. In this paper, we propose Cache of Thought (CoT), a master apprentice framework for collaborative inference between large and small VLMs. CoT manages high quality query results from large VLMs (master) in a cache, which are then selected via a novel multi modal retrieval and in-context learning to aid the performance of small VLMs (apprentice). We extensively evaluate CoT on various widely recognized and challenging general reasoning benchmarks, and show that CoT increases overall reasoning performance by up to 7.7% under the same budget, and specifically boosts the performance of apprentice VLMs by up to 36.6%. Our code is available at https://github.com/UIUC-MONET/Cache-of-Thoughts
title Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
topic Machine Learning
url https://arxiv.org/abs/2502.20587