Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Seongmin, Kim, Hyungmin, Kim, Sangwoo, Jeon, Wonseok, Yang, Juyoung, Jeon, Byeongwook, Oh, Yoonseon, Choi, Jungwook
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913867368497152
author Park, Seongmin
Kim, Hyungmin
Kim, Sangwoo
Jeon, Wonseok
Yang, Juyoung
Jeon, Byeongwook
Oh, Yoonseon
Choi, Jungwook
author_facet Park, Seongmin
Kim, Hyungmin
Kim, Sangwoo
Jeon, Wonseok
Yang, Juyoung
Jeon, Byeongwook
Oh, Yoonseon
Choi, Jungwook
contents Deep neural network (DNN)-based policy models, such as vision-language-action (VLA) models, excel at automating complex decision-making from multi-modal inputs. However, scaling these models greatly increases computational overhead, complicating deployment in resource-constrained settings like robot manipulation and autonomous driving. To address this, we propose Saliency-Aware Quantized Imitation Learning (SQIL), which combines quantization-aware training with a selective loss-weighting strategy for mission-critical states. By identifying these states via saliency scores and emphasizing them in the training loss, SQIL preserves decision fidelity under low-bit precision. We validate SQIL's generalization capability across extensive simulation benchmarks with environment variations, real-world tasks, and cross-domain tasks (self-driving, physics simulation), consistently recovering full-precision performance. Notably, a 4-bit weight-quantized VLA model for robotic manipulation achieves up to 2.5x speedup and 2.5x energy savings on an edge GPU with minimal accuracy loss. These results underline SQIL's potential for efficiently deploying large IL-based policy models on resource-limited devices.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control
Park, Seongmin
Kim, Hyungmin
Kim, Sangwoo
Jeon, Wonseok
Yang, Juyoung
Jeon, Byeongwook
Oh, Yoonseon
Choi, Jungwook
Robotics
Deep neural network (DNN)-based policy models, such as vision-language-action (VLA) models, excel at automating complex decision-making from multi-modal inputs. However, scaling these models greatly increases computational overhead, complicating deployment in resource-constrained settings like robot manipulation and autonomous driving. To address this, we propose Saliency-Aware Quantized Imitation Learning (SQIL), which combines quantization-aware training with a selective loss-weighting strategy for mission-critical states. By identifying these states via saliency scores and emphasizing them in the training loss, SQIL preserves decision fidelity under low-bit precision. We validate SQIL's generalization capability across extensive simulation benchmarks with environment variations, real-world tasks, and cross-domain tasks (self-driving, physics simulation), consistently recovering full-precision performance. Notably, a 4-bit weight-quantized VLA model for robotic manipulation achieves up to 2.5x speedup and 2.5x energy savings on an edge GPU with minimal accuracy loss. These results underline SQIL's potential for efficiently deploying large IL-based policy models on resource-limited devices.
title Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control
topic Robotics
url https://arxiv.org/abs/2505.15304