Quantization-Aware Collaborative Inference for Large Embodied AI Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lyu, Zhonghao, Xiao, Ming, Skoglund, Mikael, Debbah, Merouane, Poor, H. Vincent
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915796964343808
author Lyu, Zhonghao
Xiao, Ming
Skoglund, Mikael
Debbah, Merouane
Poor, H. Vincent
author_facet Lyu, Zhonghao
Xiao, Ming
Skoglund, Mikael
Debbah, Merouane
Poor, H. Vincent
contents Large artificial intelligence models (LAIMs) are increasingly regarded as a core intelligence engine for embodied AI applications. However, the massive parameter scale and computational demands of LAIMs pose significant challenges for resource-limited embodied agents. To address this issue, we investigate quantization-aware collaborative inference (co-inference) for embodied AI systems. First, we develop a tractable approximation for quantization-induced inference distortion. Based on this approximation, we derive lower and upper bounds on the quantization rate-inference distortion function, characterizing its dependence on LAIM statistics, including the quantization bit-width. Next, we formulate a joint quantization bit-width and computation frequency design problem under delay and energy constraints, aiming to minimize the distortion upper bound while ensuring tightness through the corresponding lower bound. Extensive evaluations validate the proposed distortion approximation, the derived rate-distortion bounds, and the effectiveness of the proposed joint design. Particularly, simulations and real-world testbed experiments demonstrate the effectiveness of the proposed joint design in balancing inference quality, latency, and energy consumption in edge embodied AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13052
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantization-Aware Collaborative Inference for Large Embodied AI Models
Lyu, Zhonghao
Xiao, Ming
Skoglund, Mikael
Debbah, Merouane
Poor, H. Vincent
Machine Learning
Signal Processing
Large artificial intelligence models (LAIMs) are increasingly regarded as a core intelligence engine for embodied AI applications. However, the massive parameter scale and computational demands of LAIMs pose significant challenges for resource-limited embodied agents. To address this issue, we investigate quantization-aware collaborative inference (co-inference) for embodied AI systems. First, we develop a tractable approximation for quantization-induced inference distortion. Based on this approximation, we derive lower and upper bounds on the quantization rate-inference distortion function, characterizing its dependence on LAIM statistics, including the quantization bit-width. Next, we formulate a joint quantization bit-width and computation frequency design problem under delay and energy constraints, aiming to minimize the distortion upper bound while ensuring tightness through the corresponding lower bound. Extensive evaluations validate the proposed distortion approximation, the derived rate-distortion bounds, and the effectiveness of the proposed joint design. Particularly, simulations and real-world testbed experiments demonstrate the effectiveness of the proposed joint design in balancing inference quality, latency, and energy consumption in edge embodied AI systems.
title Quantization-Aware Collaborative Inference for Large Embodied AI Models
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2602.13052