Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Benhammou, Yassir, Tiberio, Alessandro, Trautmann, Gabriel, Kalyan, Suman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
by: Benhammou, Yassir, et al.
Published: (2025)
by: Benhammou, Yassir, et al.
Published: (2025)
A Diffusion-Based Framework for Configurable and Realistic Multi-Storage Trace Generation
by: Kim, Seohyun, et al.
Published: (2025)
by: Kim, Seohyun, et al.
Published: (2025)
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025)
by: Bianchi, Lorenzo, et al.
Published: (2025)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
by: Lu, Zimao, et al.
Published: (2025)
by: Lu, Zimao, et al.
Published: (2025)
Dual-Image Enhanced CLIP for Zero-Shot Anomaly Detection
by: Zhang, Zhaoxiang, et al.
Published: (2024)
by: Zhang, Zhaoxiang, et al.
Published: (2024)
Performance is not All You Need: Sustainability Considerations for Algorithms
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
A Methodology to Evaluate Strategies Predicting Rankings on Unseen Domains
by: Piérard, Sébastien, et al.
Published: (2025)
by: Piérard, Sébastien, et al.
Published: (2025)
TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization
by: Zhong, Chonghao, et al.
Published: (2026)
by: Zhong, Chonghao, et al.
Published: (2026)
Enhancing Traffic Sign Recognition On The Performance Based On Yolov8
by: Ibrahim, Baba, et al.
Published: (2025)
by: Ibrahim, Baba, et al.
Published: (2025)
Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices
by: Ghafouri, Saeid, et al.
Published: (2025)
by: Ghafouri, Saeid, et al.
Published: (2025)
Resource-Efficient RGB-Only Action Recognition for Edge Deployment
by: Yoon, Dongsik, et al.
Published: (2026)
by: Yoon, Dongsik, et al.
Published: (2026)
Towards Efficient Multi-Scale Deformable Attention on NPU
by: Huang, Chenghuan, et al.
Published: (2025)
by: Huang, Chenghuan, et al.
Published: (2025)
CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment
by: Zhang, Yan, et al.
Published: (2026)
by: Zhang, Yan, et al.
Published: (2026)
TakuNet: an Energy-Efficient CNN for Real-Time Inference on Embedded UAV systems in Emergency Response Scenarios
by: Rossi, Daniel, et al.
Published: (2025)
by: Rossi, Daniel, et al.
Published: (2025)
ConvBench: A Comprehensive Benchmark for 2D Convolution Primitive Evaluation
by: Alvarenga, Lucas, et al.
Published: (2024)
by: Alvarenga, Lucas, et al.
Published: (2024)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
Transductive Zero-Shot and Few-Shot CLIP
by: Martin, Ségolène, et al.
Published: (2024)
by: Martin, Ségolène, et al.
Published: (2024)
A Closer Look at Data Augmentation Strategies for Finetuning-Based Low/Few-Shot Object Detection
by: Li, Vladislav, et al.
Published: (2024)
by: Li, Vladislav, et al.
Published: (2024)
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge
by: Sammani, Fawaz, et al.
Published: (2024)
by: Sammani, Fawaz, et al.
Published: (2024)
DISTWAR: Fast Differentiable Rendering on Raster-based Rendering Pipelines
by: Durvasula, Sankeerth, et al.
Published: (2023)
by: Durvasula, Sankeerth, et al.
Published: (2023)
How Cars Move: Analyzing Driving Dynamics for Safer Urban Traffic
by: Qian, Kangan, et al.
Published: (2024)
by: Qian, Kangan, et al.
Published: (2024)
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
by: Lebailly, Tim, et al.
Published: (2025)
by: Lebailly, Tim, et al.
Published: (2025)
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024)
by: Yu, Eric Yang, et al.
Published: (2024)
Zero-Shot Class Unlearning in CLIP with Synthetic Samples
by: Kravets, A., et al.
Published: (2024)
by: Kravets, A., et al.
Published: (2024)
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
by: Bendou, Yassir, et al.
Published: (2024)
by: Bendou, Yassir, et al.
Published: (2024)
CLAIR: CLIP-Aided Weakly Supervised Zero-Shot Cross-Domain Image Retrieval
by: Tan, Chor Boon, et al.
Published: (2025)
by: Tan, Chor Boon, et al.
Published: (2025)
What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely $F_1$
by: Piérard, Sébastien, et al.
Published: (2025)
by: Piérard, Sébastien, et al.
Published: (2025)
AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
by: Fang, Qingqing, et al.
Published: (2025)
by: Fang, Qingqing, et al.
Published: (2025)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
by: Ali, Muhammad, et al.
Published: (2024)
by: Ali, Muhammad, et al.
Published: (2024)
AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
TPCap: Unlocking Zero-Shot Image Captioning with Trigger-Augmented and Multi-Modal Purification Modules
by: Zhang, Ruoyu, et al.
Published: (2025)
by: Zhang, Ruoyu, et al.
Published: (2025)
Cost-Effective Model Evaluation with Meta-Learning
by: Pham, Trinh, et al.
Published: (2026)
by: Pham, Trinh, et al.
Published: (2026)
The Curious Case of End Token: A Zero-Shot Disentangled Image Editing using CLIP
by: Yesiltepe, Hidir, et al.
Published: (2024)
by: Yesiltepe, Hidir, et al.
Published: (2024)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
by: Vellaisamy, Prabhu, et al.
Published: (2026)
by: Vellaisamy, Prabhu, et al.
Published: (2026)
Implications of Noise in Resistive Memory on Deep Neural Networks for Image Classification
by: Emonds, Yannick, et al.
Published: (2024)
by: Emonds, Yannick, et al.
Published: (2024)
A Hitchhiker's Guide to Understanding Performances of Two-Class Classifiers
by: Halin, Anaïs, et al.
Published: (2024)
by: Halin, Anaïs, et al.
Published: (2024)
PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference
by: Fang, Jiarui, et al.
Published: (2024)
by: Fang, Jiarui, et al.
Published: (2024)
ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
by: Jafari, Sadegh, et al.
Published: (2025)
by: Jafari, Sadegh, et al.
Published: (2025)
Similar Items
-
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
by: Benhammou, Yassir, et al.
Published: (2025) -
A Diffusion-Based Framework for Configurable and Realistic Multi-Storage Trace Generation
by: Kim, Seohyun, et al.
Published: (2025) -
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024) -
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
by: Bianchi, Lorenzo, et al.
Published: (2025) -
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
by: Lu, Zimao, et al.
Published: (2025)