Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Poggi, Nicolas, Agnihotri, Shashank, Keuper, Margret
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912494617886720
author Poggi, Nicolas
Agnihotri, Shashank
Keuper, Margret
author_facet Poggi, Nicolas
Agnihotri, Shashank
Keuper, Margret
contents Terahertz (THz) imaging enables non-invasive analysis for applications such as security screening and material classification, but effective image classification remains challenging due to limited annotations, low resolution, and visual ambiguity. We introduce In-Context Learning (ICL) with Vision-Language Models (VLMs) as a flexible, interpretable alternative that requires no fine-tuning. Using a modality-aligned prompting framework, we adapt two open-weight VLMs to the THz domain and evaluate them under zero-shot and one-shot settings. Our results show that ICL improves classification and interpretability in low-data regimes. This is the first application of ICL-enhanced VLMs to THz imaging, offering a promising direction for resource-constrained scientific domains. Code: \href{https://github.com/Nicolas-Poggi/Project_THz_Classification/tree/main}{GitHub repository}.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
Poggi, Nicolas
Agnihotri, Shashank
Keuper, Margret
Computation and Language
Computer Vision and Pattern Recognition
Terahertz (THz) imaging enables non-invasive analysis for applications such as security screening and material classification, but effective image classification remains challenging due to limited annotations, low resolution, and visual ambiguity. We introduce In-Context Learning (ICL) with Vision-Language Models (VLMs) as a flexible, interpretable alternative that requires no fine-tuning. Using a modality-aligned prompting framework, we adapt two open-weight VLMs to the THz domain and evaluate them under zero-shot and one-shot settings. Our results show that ICL improves classification and interpretability in low-data regimes. This is the first application of ICL-enhanced VLMs to THz imaging, offering a promising direction for resource-constrained scientific domains. Code: \href{https://github.com/Nicolas-Poggi/Project_THz_Classification/tree/main}{GitHub repository}.
title Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.15576