Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benhammou, Yassir, Tiberio, Alessandro, Trautmann, Gabriel, Kalyan, Suman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913801262071808
author Benhammou, Yassir
Tiberio, Alessandro
Trautmann, Gabriel
Kalyan, Suman
author_facet Benhammou, Yassir
Tiberio, Alessandro
Trautmann, Gabriel
Kalyan, Suman
contents MILS (Multimodal Iterative LLM Solver) is a recently published framework that claims "LLMs can see and hear without any training" by leveraging an iterative, LLM-CLIP based approach for zero-shot image captioning. While this MILS approach demonstrates good performance, our investigation reveals that this success comes at a hidden, substantial computational cost due to its expensive multi-step refinement process. In contrast, alternative models such as BLIP-2 and GPT-4V achieve competitive results through a streamlined, single-pass approach. We hypothesize that the significant overhead inherent in MILS's iterative process may undermine its practical benefits, thereby challenging the narrative that zero-shot performance can be attained without incurring heavy resource demands. This work is the first to expose and quantify the trade-offs between output quality and computational cost in MILS, providing critical insights for the design of more efficient multimodal models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15199
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
Benhammou, Yassir
Tiberio, Alessandro
Trautmann, Gabriel
Kalyan, Suman
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Performance
MILS (Multimodal Iterative LLM Solver) is a recently published framework that claims "LLMs can see and hear without any training" by leveraging an iterative, LLM-CLIP based approach for zero-shot image captioning. While this MILS approach demonstrates good performance, our investigation reveals that this success comes at a hidden, substantial computational cost due to its expensive multi-step refinement process. In contrast, alternative models such as BLIP-2 and GPT-4V achieve competitive results through a streamlined, single-pass approach. We hypothesize that the significant overhead inherent in MILS's iterative process may undermine its practical benefits, thereby challenging the narrative that zero-shot performance can be attained without incurring heavy resource demands. This work is the first to expose and quantify the trade-offs between output quality and computational cost in MILS, providing critical insights for the design of more efficient multimodal models.
title Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Performance
url https://arxiv.org/abs/2504.15199