Preventing Robotic Jailbreaking via Multimodal Domain Adaptation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Marchiori, Francesco, Sinha, Rohan, Agia, Christopher, Robey, Alexander, Pappas, George J., Conti, Mauro, Pavone, Marco
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912611738583040
author Marchiori, Francesco
Sinha, Rohan
Agia, Christopher
Robey, Alexander
Pappas, George J.
Conti, Mauro
Pavone, Marco
author_facet Marchiori, Francesco
Sinha, Rohan
Agia, Christopher
Robey, Alexander
Pappas, George J.
Conti, Mauro
Pavone, Marco
contents Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly deployed in robotic environments but remain vulnerable to jailbreaking attacks that bypass safety mechanisms and drive unsafe or physically harmful behaviors in the real world. Data-driven defenses such as jailbreak classifiers show promise, yet they struggle to generalize in domains where specialized datasets are scarce, limiting their effectiveness in robotics and other safety-critical contexts. To address this gap, we introduce J-DAPT, a lightweight framework for multimodal jailbreak detection through attention-based fusion and domain adaptation. J-DAPT integrates textual and visual embeddings to capture both semantic intent and environmental grounding, while aligning general-purpose jailbreak datasets with domain-specific reference data. Evaluations across autonomous driving, maritime robotics, and quadruped navigation show that J-DAPT boosts detection accuracy to nearly 100% with minimal overhead. These results demonstrate that J-DAPT provides a practical defense for securing VLMs in robotic applications. Additional materials are made available at: https://j-dapt.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23281
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
Marchiori, Francesco
Sinha, Rohan
Agia, Christopher
Robey, Alexander
Pappas, George J.
Conti, Mauro
Pavone, Marco
Robotics
I.2.6; I.2.9
Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly deployed in robotic environments but remain vulnerable to jailbreaking attacks that bypass safety mechanisms and drive unsafe or physically harmful behaviors in the real world. Data-driven defenses such as jailbreak classifiers show promise, yet they struggle to generalize in domains where specialized datasets are scarce, limiting their effectiveness in robotics and other safety-critical contexts. To address this gap, we introduce J-DAPT, a lightweight framework for multimodal jailbreak detection through attention-based fusion and domain adaptation. J-DAPT integrates textual and visual embeddings to capture both semantic intent and environmental grounding, while aligning general-purpose jailbreak datasets with domain-specific reference data. Evaluations across autonomous driving, maritime robotics, and quadruped navigation show that J-DAPT boosts detection accuracy to nearly 100% with minimal overhead. These results demonstrate that J-DAPT provides a practical defense for securing VLMs in robotic applications. Additional materials are made available at: https://j-dapt.github.io.
title Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
topic Robotics
I.2.6; I.2.9
url https://arxiv.org/abs/2509.23281