Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vishwanathan, Manoj, Subramanian, Suvinay, Raghunathan, Anand
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917308189900800
author Vishwanathan, Manoj
Subramanian, Suvinay
Raghunathan, Anand
author_facet Vishwanathan, Manoj
Subramanian, Suvinay
Raghunathan, Anand
contents Vision-Language-Action (VLA) models are an emerging class of workloads critical for robotics and embodied AI at the edge. As these models scale, they demonstrate significant capability gains, yet they must be deployed locally to meet the strict latency requirements of real-time applications. This paper characterizes VLA performance on two generations of edge hardware, viz. the Nvidia Jetson Orin and Thor platforms. Using MolmoAct-7B, a state-of-the-art VLA model, we identify a primary execution bottleneck: up to 75% of end-to-end latency is consumed by the memory-bound action-generation phase. Through analytical modeling and simulations, we project the hardware requirements for scaling to 100B parameter models. We also explore the impact of high-bandwidth memory technologies and processing-in-memory (PIM) as promising future pathways in edge systems for embodied AI.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02271
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
Vishwanathan, Manoj
Subramanian, Suvinay
Raghunathan, Anand
Performance
Artificial Intelligence
Hardware Architecture
Robotics
Vision-Language-Action (VLA) models are an emerging class of workloads critical for robotics and embodied AI at the edge. As these models scale, they demonstrate significant capability gains, yet they must be deployed locally to meet the strict latency requirements of real-time applications. This paper characterizes VLA performance on two generations of edge hardware, viz. the Nvidia Jetson Orin and Thor platforms. Using MolmoAct-7B, a state-of-the-art VLA model, we identify a primary execution bottleneck: up to 75% of end-to-end latency is consumed by the memory-bound action-generation phase. Through analytical modeling and simulations, we project the hardware requirements for scaling to 100B parameter models. We also explore the impact of high-bandwidth memory technologies and processing-in-memory (PIM) as promising future pathways in edge systems for embodied AI.
title Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
topic Performance
Artificial Intelligence
Hardware Architecture
Robotics
url https://arxiv.org/abs/2603.02271