Concept-based Analysis of Neural Networks via Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mangal, Ravi, Narodytska, Nina, Gopinath, Divya, Hu, Boyue Caroline, Roy, Anirban, Jha, Susmit, Pasareanu, Corina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911835386544128
author Mangal, Ravi
Narodytska, Nina
Gopinath, Divya
Hu, Boyue Caroline
Roy, Anirban
Jha, Susmit
Pasareanu, Corina
author_facet Mangal, Ravi
Narodytska, Nina
Gopinath, Divya
Hu, Boyue Caroline
Roy, Anirban
Jha, Susmit
Pasareanu, Corina
contents The analysis of vision-based deep neural networks (DNNs) is highly desirable but it is very challenging due to the difficulty of expressing formal specifications for vision tasks and the lack of efficient verification procedures. In this paper, we propose to leverage emerging multimodal, vision-language, foundation models (VLMs) as a lens through which we can reason about vision models. VLMs have been trained on a large body of images accompanied by their textual description, and are thus implicitly aware of high-level, human-understandable concepts describing the images. We describe a logical specification language $\texttt{Con}_{\texttt{spec}}$ designed to facilitate writing specifications in terms of these concepts. To define and formally check $\texttt{Con}_{\texttt{spec}}$ specifications, we build a map between the internal representations of a given vision model and a VLM, leading to an efficient verification procedure of natural-language properties for vision models. We demonstrate our techniques on a ResNet-based classifier trained on the RIVAL-10 dataset using CLIP as the multimodal model.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19837
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Concept-based Analysis of Neural Networks via Vision-Language Models
Mangal, Ravi
Narodytska, Nina
Gopinath, Divya
Hu, Boyue Caroline
Roy, Anirban
Jha, Susmit
Pasareanu, Corina
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Logic in Computer Science
The analysis of vision-based deep neural networks (DNNs) is highly desirable but it is very challenging due to the difficulty of expressing formal specifications for vision tasks and the lack of efficient verification procedures. In this paper, we propose to leverage emerging multimodal, vision-language, foundation models (VLMs) as a lens through which we can reason about vision models. VLMs have been trained on a large body of images accompanied by their textual description, and are thus implicitly aware of high-level, human-understandable concepts describing the images. We describe a logical specification language $\texttt{Con}_{\texttt{spec}}$ designed to facilitate writing specifications in terms of these concepts. To define and formally check $\texttt{Con}_{\texttt{spec}}$ specifications, we build a map between the internal representations of a given vision model and a VLM, leading to an efficient verification procedure of natural-language properties for vision models. We demonstrate our techniques on a ResNet-based classifier trained on the RIVAL-10 dataset using CLIP as the multimodal model.
title Concept-based Analysis of Neural Networks via Vision-Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Logic in Computer Science
url https://arxiv.org/abs/2403.19837