Attacking Attention of Foundation Models Disrupts Downstream Tasks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Silva, Hondamunige Prasanna, Becattini, Federico, Seidenari, Lorenzo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909784295342080
author Silva, Hondamunige Prasanna
Becattini, Federico
Seidenari, Lorenzo
author_facet Silva, Hondamunige Prasanna
Becattini, Federico
Seidenari, Lorenzo
contents Foundation models represent the most prominent and recent paradigm shift in artificial intelligence. Foundation models are large models, trained on broad data that deliver high accuracy in many downstream tasks, often without fine-tuning. For this reason, models such as CLIP , DINO or Vision Transfomers (ViT), are becoming the bedrock of many industrial AI-powered applications. However, the reliance on pre-trained foundation models also introduces significant security concerns, as these models are vulnerable to adversarial attacks. Such attacks involve deliberately crafted inputs designed to deceive AI systems, jeopardizing their reliability. This paper studies the vulnerabilities of vision foundation models, focusing specifically on CLIP and ViTs, and explores the transferability of adversarial attacks to downstream tasks. We introduce a novel attack, targeting the structure of transformer-based architectures in a task-agnostic fashion. We demonstrate the effectiveness of our attack on several downstream tasks: classification, captioning, image/text retrieval, segmentation and depth estimation. Code available at:https://github.com/HondamunigePrasannaSilva/attack-attention
format Preprint
id arxiv_https___arxiv_org_abs_2506_05394
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attacking Attention of Foundation Models Disrupts Downstream Tasks
Silva, Hondamunige Prasanna
Becattini, Federico
Seidenari, Lorenzo
Cryptography and Security
Machine Learning
Foundation models represent the most prominent and recent paradigm shift in artificial intelligence. Foundation models are large models, trained on broad data that deliver high accuracy in many downstream tasks, often without fine-tuning. For this reason, models such as CLIP , DINO or Vision Transfomers (ViT), are becoming the bedrock of many industrial AI-powered applications. However, the reliance on pre-trained foundation models also introduces significant security concerns, as these models are vulnerable to adversarial attacks. Such attacks involve deliberately crafted inputs designed to deceive AI systems, jeopardizing their reliability. This paper studies the vulnerabilities of vision foundation models, focusing specifically on CLIP and ViTs, and explores the transferability of adversarial attacks to downstream tasks. We introduce a novel attack, targeting the structure of transformer-based architectures in a task-agnostic fashion. We demonstrate the effectiveness of our attack on several downstream tasks: classification, captioning, image/text retrieval, segmentation and depth estimation. Code available at:https://github.com/HondamunigePrasannaSilva/attack-attention
title Attacking Attention of Foundation Models Disrupts Downstream Tasks
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.05394