Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Ju-Young, Park, Ji-Hong, Kim, Myeongjun, Kim, Gun-Woo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918247673102336
author Kim, Ju-Young
Park, Ji-Hong
Kim, Myeongjun
Kim, Gun-Woo
author_facet Kim, Ju-Young
Park, Ji-Hong
Kim, Myeongjun
Kim, Gun-Woo
contents Smart farming has emerged as a key technology for advancing modern agriculture through automation and intelligent control. However, systems relying on RGB cameras for perception and robotic manipulators for control, common in smart farming, are vulnerable to photometric perturbations such as hue, illumination, and noise changes, which can cause malfunction under adversarial attacks. To address this issue, we propose an explainable adversarial-robust Vision-Language-Action model based on the OpenVLA-OFT framework. The model integrates an Evidence-3 module that detects photometric perturbations and generates natural language explanations of their causes and effects. Experiments show that the proposed model reduces Current Action L1 loss by 21.7% and Next Actions L1 loss by 18.4% compared to the baseline, demonstrating improved action prediction accuracy and explainability under adversarial conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11865
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
Kim, Ju-Young
Park, Ji-Hong
Kim, Myeongjun
Kim, Gun-Woo
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Smart farming has emerged as a key technology for advancing modern agriculture through automation and intelligent control. However, systems relying on RGB cameras for perception and robotic manipulators for control, common in smart farming, are vulnerable to photometric perturbations such as hue, illumination, and noise changes, which can cause malfunction under adversarial attacks. To address this issue, we propose an explainable adversarial-robust Vision-Language-Action model based on the OpenVLA-OFT framework. The model integrates an Evidence-3 module that detects photometric perturbations and generates natural language explanations of their causes and effects. Experiments show that the proposed model reduces Current Action L1 loss by 21.7% and Next Actions L1 loss by 18.4% compared to the baseline, demonstrating improved action prediction accuracy and explainability under adversarial conditions.
title Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2512.11865