Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taherin, Amir, Lin, Juyi, Akbari, Arash, Akbari, Arman, Zhao, Pu, Chen, Weiwei, Kaeli, David, Wang, Yanzhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908787676282880
author Taherin, Amir
Lin, Juyi
Akbari, Arash
Akbari, Arman
Zhao, Pu
Chen, Weiwei
Kaeli, David
Wang, Yanzhi
author_facet Taherin, Amir
Lin, Juyi
Akbari, Arash
Akbari, Arman
Zhao, Pu
Chen, Weiwei
Kaeli, David
Wang, Yanzhi
contents Vision-Language-Action (VLA) models have emerged as powerful generalist policies for robotic control, yet their performance scaling across model architectures and hardware platforms, as well as their associated power budgets, remain poorly understood. This work presents an evaluation of five representative VLA models -- spanning state-of-the-art baselines and two newly proposed architectures -- targeting edge and datacenter GPU platforms. Using the LIBERO benchmark, we measure accuracy alongside system-level metrics, including latency, throughput, and peak memory usage, under varying edge power constraints and high-performance datacenter GPU configurations. Our results identify distinct scaling trends: (1) architectural choices, such as action tokenization and model backbone size, strongly influence throughput and memory footprint; (2) power-constrained edge devices exhibit non-linear performance degradation, with some configurations matching or exceeding older datacenter GPUs; and (3) high-throughput variants can be achieved without significant accuracy loss. These findings provide actionable insights when selecting and optimizing VLAs across a range of deployment constraints. Our work challenges current assumptions about the superiority of datacenter hardware for robotic inference.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11480
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
Taherin, Amir
Lin, Juyi
Akbari, Arash
Akbari, Arman
Zhao, Pu
Chen, Weiwei
Kaeli, David
Wang, Yanzhi
Artificial Intelligence
Computer Vision and Pattern Recognition
Emerging Technologies
Machine Learning
Robotics
Vision-Language-Action (VLA) models have emerged as powerful generalist policies for robotic control, yet their performance scaling across model architectures and hardware platforms, as well as their associated power budgets, remain poorly understood. This work presents an evaluation of five representative VLA models -- spanning state-of-the-art baselines and two newly proposed architectures -- targeting edge and datacenter GPU platforms. Using the LIBERO benchmark, we measure accuracy alongside system-level metrics, including latency, throughput, and peak memory usage, under varying edge power constraints and high-performance datacenter GPU configurations. Our results identify distinct scaling trends: (1) architectural choices, such as action tokenization and model backbone size, strongly influence throughput and memory footprint; (2) power-constrained edge devices exhibit non-linear performance degradation, with some configurations matching or exceeding older datacenter GPUs; and (3) high-throughput variants can be achieved without significant accuracy loss. These findings provide actionable insights when selecting and optimizing VLAs across a range of deployment constraints. Our work challenges current assumptions about the superiority of datacenter hardware for robotic inference.
title Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Emerging Technologies
Machine Learning
Robotics
url https://arxiv.org/abs/2509.11480