Trust-Aware Routing for Distributed Generative AI Inference at the Edge
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Chanh, Elmroth, Erik |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
por: Zheng, Peirong, et al.
Publicado: (2026)
por: Zheng, Peirong, et al.
Publicado: (2026)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
por: Reddy, Tella Rajashekhar, et al.
Publicado: (2025)
por: Reddy, Tella Rajashekhar, et al.
Publicado: (2025)
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
por: Morabito, Roberto, et al.
Publicado: (2025)
por: Morabito, Roberto, et al.
Publicado: (2025)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
por: Ye, Shengyuan, et al.
Publicado: (2025)
por: Ye, Shengyuan, et al.
Publicado: (2025)
Generative AI on the Edge: Architecture and Performance Evaluation
por: Nezami, Zeinab, et al.
Publicado: (2024)
por: Nezami, Zeinab, et al.
Publicado: (2024)
Design and Optimization of Hierarchical Gradient Coding for Distributed Learning at Edge Devices
por: Tang, Weiheng, et al.
Publicado: (2024)
por: Tang, Weiheng, et al.
Publicado: (2024)
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
por: Yang, Jin, et al.
Publicado: (2025)
por: Yang, Jin, et al.
Publicado: (2025)
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
por: Sun, Tingyang, et al.
Publicado: (2025)
por: Sun, Tingyang, et al.
Publicado: (2025)
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
por: Wang, Zhanwei, et al.
Publicado: (2026)
por: Wang, Zhanwei, et al.
Publicado: (2026)
Early-Exit meets Model-Distributed Inference at Edge Networks
por: Colocrese, Marco, et al.
Publicado: (2024)
por: Colocrese, Marco, et al.
Publicado: (2024)
Towards Edge General Intelligence via Large Language Models: Opportunities and Challenges
por: Chen, Handi, et al.
Publicado: (2024)
por: Chen, Handi, et al.
Publicado: (2024)
Context-Aware Orchestration of Energy-Efficient Gossip Learning Schemes
por: Dinani, Mina Aghaei, et al.
Publicado: (2024)
por: Dinani, Mina Aghaei, et al.
Publicado: (2024)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
por: Jang, SiYoung, et al.
Publicado: (2025)
por: Jang, SiYoung, et al.
Publicado: (2025)
Agentic Performance at the Edge: Insights from Benchmarking
por: Wang, Shiqiang, et al.
Publicado: (2026)
por: Wang, Shiqiang, et al.
Publicado: (2026)
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
por: Ye, Shengyuan, et al.
Publicado: (2024)
por: Ye, Shengyuan, et al.
Publicado: (2024)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
por: Luo, Haoxiang, et al.
Publicado: (2025)
por: Luo, Haoxiang, et al.
Publicado: (2025)
Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
por: Sivtsov, Danil, et al.
Publicado: (2025)
por: Sivtsov, Danil, et al.
Publicado: (2025)
Digital Twinning of a Pressurized Water Reactor Startup Operation and Partial Computational Offloading in In-network Computing-Assisted Multiaccess Edge Computing
por: Aliyu, Ibrahim, et al.
Publicado: (2024)
por: Aliyu, Ibrahim, et al.
Publicado: (2024)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
por: Reddy, Tella Rajashekhar, et al.
Publicado: (2026)
por: Reddy, Tella Rajashekhar, et al.
Publicado: (2026)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
por: Du, Chengze, et al.
Publicado: (2025)
por: Du, Chengze, et al.
Publicado: (2025)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
por: Seo, Eunil, et al.
Publicado: (2025)
por: Seo, Eunil, et al.
Publicado: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
por: Liu, Zedong, et al.
Publicado: (2026)
por: Liu, Zedong, et al.
Publicado: (2026)
High-speed Networking for Giga-Scale AI Factories
por: Khashab, Sajy, et al.
Publicado: (2026)
por: Khashab, Sajy, et al.
Publicado: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
por: Chen, Zixuan, et al.
Publicado: (2024)
por: Chen, Zixuan, et al.
Publicado: (2024)
An Open API Architecture to Discover the Trustworthy Explanation of Cloud AI Services
por: Wang, Zerui, et al.
Publicado: (2024)
por: Wang, Zerui, et al.
Publicado: (2024)
Towards Net-Zero Carbon Emissions in Network AI for 6G and Beyond
por: Zhang, Peng, et al.
Publicado: (2023)
por: Zhang, Peng, et al.
Publicado: (2023)
Enabling Intelligent Vehicular Networks Through Distributed Learning in the Non-Terrestrial Networks 6G Vision
por: Naseh, David, et al.
Publicado: (2023)
por: Naseh, David, et al.
Publicado: (2023)
Policy Design in Zero-Trust Distributed Networks: Challenges and Solutions
por: Sandjaja, Fannya R., et al.
Publicado: (2025)
por: Sandjaja, Fannya R., et al.
Publicado: (2025)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
por: Li, Minghao, et al.
Publicado: (2026)
por: Li, Minghao, et al.
Publicado: (2026)
Communication Optimization for Decentralized Learning atop Bandwidth-limited Edge Networks
por: Sun, Tingyang, et al.
Publicado: (2025)
por: Sun, Tingyang, et al.
Publicado: (2025)
Dynamic and Distributed Routing in IoT Networks based on Multi-Objective Q-Learning
por: Vaishnav, Shubham, et al.
Publicado: (2025)
por: Vaishnav, Shubham, et al.
Publicado: (2025)
Network Anomaly Detection in Distributed Edge Computing Infrastructure
por: Marfo, William, et al.
Publicado: (2025)
por: Marfo, William, et al.
Publicado: (2025)
Federated Continual Learning for Edge-AI: A Comprehensive Survey
por: Wang, Zi, et al.
Publicado: (2024)
por: Wang, Zi, et al.
Publicado: (2024)
Resilient by Design -- Active Inference for Distributed Continuum Intelligence
por: Donta, Praveen Kumar, et al.
Publicado: (2025)
por: Donta, Praveen Kumar, et al.
Publicado: (2025)
Contention-Aware Microservice Deployment in Collaborative Mobile Edge Networks
por: Ge, Xinlei, et al.
Publicado: (2024)
por: Ge, Xinlei, et al.
Publicado: (2024)
Deep Tech to Space: Space Data Centers and AI Revolution at the Edge
por: Weiss, Jonas, et al.
Publicado: (2026)
por: Weiss, Jonas, et al.
Publicado: (2026)
Implementation of Big AI Models for Wireless Networks with Collaborative Edge Computing
por: Zeng, Liekang, et al.
Publicado: (2024)
por: Zeng, Liekang, et al.
Publicado: (2024)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
por: Yang, Zheming, et al.
Publicado: (2024)
por: Yang, Zheming, et al.
Publicado: (2024)
A Multi-Layered Distributed Computing Framework for Enhanced Edge Computing
por: Ma, Ke, et al.
Publicado: (2024)
por: Ma, Ke, et al.
Publicado: (2024)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
por: Kasnavieh, Hossein Hosseini, et al.
Publicado: (2026)
Ejemplares similares
-
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
por: Zheng, Peirong, et al.
Publicado: (2026) -
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
por: Reddy, Tella Rajashekhar, et al.
Publicado: (2025) -
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
por: Morabito, Roberto, et al.
Publicado: (2025) -
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
por: Ye, Shengyuan, et al.
Publicado: (2025) -
Generative AI on the Edge: Architecture and Performance Evaluation
por: Nezami, Zeinab, et al.
Publicado: (2024)