MVVM: Deploy Your AI Agents-Securely, Efficiently, Everywhere

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yiwei, Hu, Aibo, Zheng, Yusheng, Zhao, Brian, Zhang, Xinqi, Xiang, Dawei, Chu, Kexin, Zhang, Wei, Quinn, Andi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909802681073664
author Yang, Yiwei
Hu, Aibo
Zheng, Yusheng
Zhao, Brian
Zhang, Xinqi
Xiang, Dawei
Chu, Kexin
Zhang, Wei
Quinn, Andi
author_facet Yang, Yiwei
Hu, Aibo
Zheng, Yusheng
Zhao, Brian
Zhang, Xinqi
Xiang, Dawei
Chu, Kexin
Zhang, Wei
Quinn, Andi
contents The rise of AI agents powered by Large Language Models (LLMs) presents critical challenges: how to securely execute and migrate these agents across heterogeneous environments while protecting sensitive user data, maintaining availability during network failures, minimizing response latency for time-critical decisions, and ensuring output safety in mission-critical applications. We present MVVM, a WebAssembly-based secure container framework that enables transparent live migration of LLM agent workspaces between edge devices and cloud servers with end-to-end privacy guarantees, resilient multi-tier replication, speculative execution for latency optimization, and integrated validation for safety assurance. MVVM introduces two key innovations: (1) a two-way sandboxing framework leveraging hardware enclaves and accelerator extensions that protects both the agent from malicious hosts and the host from compromised agents; (2) an efficient cross platform migration mechanism using WebAssembly and WASI's platform-agnostic design, enabling seamless movement across ARM phones, RISC-V MCUs, x86 servers, and heterogeneous accelerators; and three astonishing use cases: (1) privacy-aware daemon that automatically determines whether to execute locally or remotely based on data sensitivity and resource availability; (2) multi-tier replication with intelligent quality degradation that maintains service availability despite network failures or resource constraints; (3) a comprehensive execution framework combining speculative execution for 10x latency reduction with parallel validation that ensures output safety without compromising responsiveness. Our evaluation demonstrates that MVVM is validated on three separate devices across 18 workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15894
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MVVM: Deploy Your AI Agents-Securely, Efficiently, Everywhere
Yang, Yiwei
Hu, Aibo
Zheng, Yusheng
Zhao, Brian
Zhang, Xinqi
Xiang, Dawei
Chu, Kexin
Zhang, Wei
Quinn, Andi
Operating Systems
The rise of AI agents powered by Large Language Models (LLMs) presents critical challenges: how to securely execute and migrate these agents across heterogeneous environments while protecting sensitive user data, maintaining availability during network failures, minimizing response latency for time-critical decisions, and ensuring output safety in mission-critical applications. We present MVVM, a WebAssembly-based secure container framework that enables transparent live migration of LLM agent workspaces between edge devices and cloud servers with end-to-end privacy guarantees, resilient multi-tier replication, speculative execution for latency optimization, and integrated validation for safety assurance. MVVM introduces two key innovations: (1) a two-way sandboxing framework leveraging hardware enclaves and accelerator extensions that protects both the agent from malicious hosts and the host from compromised agents; (2) an efficient cross platform migration mechanism using WebAssembly and WASI's platform-agnostic design, enabling seamless movement across ARM phones, RISC-V MCUs, x86 servers, and heterogeneous accelerators; and three astonishing use cases: (1) privacy-aware daemon that automatically determines whether to execute locally or remotely based on data sensitivity and resource availability; (2) multi-tier replication with intelligent quality degradation that maintains service availability despite network failures or resource constraints; (3) a comprehensive execution framework combining speculative execution for 10x latency reduction with parallel validation that ensures output safety without compromising responsiveness. Our evaluation demonstrates that MVVM is validated on three separate devices across 18 workloads.
title MVVM: Deploy Your AI Agents-Securely, Efficiently, Everywhere
topic Operating Systems
url https://arxiv.org/abs/2410.15894