Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lu, Hong, Li, Hengxu, Shahani, Prithviraj Singh, Herbers, Stephanie, Scheutz, Matthias
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909482362077184
author Lu, Hong
Li, Hengxu
Shahani, Prithviraj Singh
Herbers, Stephanie
Scheutz, Matthias
author_facet Lu, Hong
Li, Hengxu
Shahani, Prithviraj Singh
Herbers, Stephanie
Scheutz, Matthias
contents Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to their black-box nature and sensitivity to environmental changes. In contrast, cognitive architectures (CA) excel in symbolic reasoning and state monitoring but are constrained by rigid predefined execution. This work bridges these approaches by probing OpenVLA's hidden layers to uncover symbolic representations of object properties, relations, and action states, enabling integration with a CA for enhanced interpretability and robustness. Through experiments on LIBERO-spatial pick-and-place tasks, we analyze the encoding of symbolic states across different layers of OpenVLA's Llama backbone. Our probing results show consistently high accuracies (> 0.90) for both object and action states across most layers, though contrary to our hypotheses, we did not observe the expected pattern of object states being encoded earlier than action states. We demonstrate an integrated DIARC-OpenVLA system that leverages these symbolic representations for real-time state monitoring, laying the foundation for more interpretable and reliable robotic manipulation.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
Lu, Hong
Li, Hengxu
Shahani, Prithviraj Singh
Herbers, Stephanie
Scheutz, Matthias
Robotics
Artificial Intelligence
Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to their black-box nature and sensitivity to environmental changes. In contrast, cognitive architectures (CA) excel in symbolic reasoning and state monitoring but are constrained by rigid predefined execution. This work bridges these approaches by probing OpenVLA's hidden layers to uncover symbolic representations of object properties, relations, and action states, enabling integration with a CA for enhanced interpretability and robustness. Through experiments on LIBERO-spatial pick-and-place tasks, we analyze the encoding of symbolic states across different layers of OpenVLA's Llama backbone. Our probing results show consistently high accuracies (> 0.90) for both object and action states across most layers, though contrary to our hypotheses, we did not observe the expected pattern of object states being encoded earlier than action states. We demonstrate an integrated DIARC-OpenVLA system that leverages these symbolic representations for real-time state monitoring, laying the foundation for more interpretable and reliable robotic manipulation.
title Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2502.04558