Jailbreaking Large Vision Language Models in Intelligent Transportation Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Das, Badhan Chandra, Jawad, Md Tasnim, Mia, Md Jueal, Amini, M. Hadi, Wu, Yanzhao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918206877204480
author Das, Badhan Chandra
Jawad, Md Tasnim
Mia, Md Jueal
Amini, M. Hadi
Wu, Yanzhao
author_facet Das, Badhan Chandra
Jawad, Md Tasnim
Mia, Md Jueal
Amini, M. Hadi
Wu, Yanzhao
contents Large Vision Language Models (LVLMs) demonstrate strong capabilities in multimodal reasoning and many real-world applications, such as visual question answering. However, LVLMs are highly vulnerable to jailbreaking attacks. This paper systematically analyzes the vulnerabilities of LVLMs integrated in Intelligent Transportation Systems (ITS) under carefully crafted jailbreaking attacks. First, we carefully construct a dataset with harmful queries relevant to transportation, following OpenAI's prohibited categories to which the LVLMs should not respond. Second, we introduce a novel jailbreaking attack that exploits the vulnerabilities of LVLMs through image typography manipulation and multi-turn prompting. Third, we propose a multi-layered response filtering defense technique to prevent the model from generating inappropriate responses. We perform extensive experiments with the proposed attack and defense on the state-of-the-art LVLMs (both open-source and closed-source). To evaluate the attack method and defense technique, we use GPT-4's judgment to determine the toxicity score of the generated responses, as well as manual verification. Further, we compare our proposed jailbreaking method with existing jailbreaking techniques and highlight severe security risks involved with jailbreaking attacks with image typography manipulation and multi-turn prompting in the LVLMs integrated in ITS.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13892
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
Das, Badhan Chandra
Jawad, Md Tasnim
Mia, Md Jueal
Amini, M. Hadi
Wu, Yanzhao
Artificial Intelligence
Large Vision Language Models (LVLMs) demonstrate strong capabilities in multimodal reasoning and many real-world applications, such as visual question answering. However, LVLMs are highly vulnerable to jailbreaking attacks. This paper systematically analyzes the vulnerabilities of LVLMs integrated in Intelligent Transportation Systems (ITS) under carefully crafted jailbreaking attacks. First, we carefully construct a dataset with harmful queries relevant to transportation, following OpenAI's prohibited categories to which the LVLMs should not respond. Second, we introduce a novel jailbreaking attack that exploits the vulnerabilities of LVLMs through image typography manipulation and multi-turn prompting. Third, we propose a multi-layered response filtering defense technique to prevent the model from generating inappropriate responses. We perform extensive experiments with the proposed attack and defense on the state-of-the-art LVLMs (both open-source and closed-source). To evaluate the attack method and defense technique, we use GPT-4's judgment to determine the toxicity score of the generated responses, as well as manual verification. Further, we compare our proposed jailbreaking method with existing jailbreaking techniques and highlight severe security risks involved with jailbreaking attacks with image typography manipulation and multi-turn prompting in the LVLMs integrated in ITS.
title Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
topic Artificial Intelligence
url https://arxiv.org/abs/2511.13892