Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haiyuan, Madhukumar, Hari, Yan, Shuangyi, Wu, Yulei, Simeonidou, Dimitra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deadline-Driven Hierarchical Agentic Resource Sharing for AI Services and RAN Functions in AI-RAN
von: Li, Haiyuan, et al.
Veröffentlicht: (2026)
von: Li, Haiyuan, et al.
Veröffentlicht: (2026)
Towards Practical Operation of Deep Reinforcement Learning Agents in Real-World Network Management at Open RAN Edges
von: Li, Haiyuan, et al.
Veröffentlicht: (2024)
von: Li, Haiyuan, et al.
Veröffentlicht: (2024)
Profiling AI Models: Towards Efficient Computation Offloading in Heterogeneous Edge AI Systems
von: Parra-Ullauri, Juan Marcelo, et al.
Veröffentlicht: (2024)
von: Parra-Ullauri, Juan Marcelo, et al.
Veröffentlicht: (2024)
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024)
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
Data Sharing at the Edge of the Network: A Disturbance Resilient Multi-modal ITS
von: Mikolasek, Igor, et al.
Veröffentlicht: (2024)
von: Mikolasek, Igor, et al.
Veröffentlicht: (2024)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
von: Liang, Wenyi, et al.
Veröffentlicht: (2024)
von: Liang, Wenyi, et al.
Veröffentlicht: (2024)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
von: Li, Senyao, et al.
Veröffentlicht: (2026)
von: Li, Senyao, et al.
Veröffentlicht: (2026)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026)
von: Chen, Zhixiong, et al.
Veröffentlicht: (2026)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
von: Sheng, Jinhao, et al.
Veröffentlicht: (2025)
von: Sheng, Jinhao, et al.
Veröffentlicht: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
von: Yoon, JinYi, et al.
Veröffentlicht: (2025)
von: Yoon, JinYi, et al.
Veröffentlicht: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
von: Rajashekar, Kolichala, et al.
Veröffentlicht: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
von: O'Quinn, Austin, et al.
Veröffentlicht: (2025)
von: O'Quinn, Austin, et al.
Veröffentlicht: (2025)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Knowledge-driven Reasoning for Mobile Agentic AI: Concepts, Approaches, and Directions
von: Liu, Guangyuan, et al.
Veröffentlicht: (2026)
von: Liu, Guangyuan, et al.
Veröffentlicht: (2026)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
von: Nguyen, Thanh-Tung, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh-Tung, et al.
Veröffentlicht: (2025)
Mobile Edge Computing
von: Ahmed, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmed, Sohaib, et al.
Veröffentlicht: (2024)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
Priority-Aware Model-Distributed Inference at Edge Networks
von: Li, Teng, et al.
Veröffentlicht: (2024)
von: Li, Teng, et al.
Veröffentlicht: (2024)
Energy-Efficient Real-Time Job Mapping and Resource Management in Mobile-Edge Computing
von: Gao, Chuanchao, et al.
Veröffentlicht: (2025)
von: Gao, Chuanchao, et al.
Veröffentlicht: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
von: Wolfrath, Joel, et al.
Veröffentlicht: (2025)
von: Wolfrath, Joel, et al.
Veröffentlicht: (2025)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
von: Wang, Kewei, et al.
Veröffentlicht: (2025)
von: Wang, Kewei, et al.
Veröffentlicht: (2025)
SynergAI: Edge-to-Cloud Synergy for Architecture-Driven High-Performance Orchestration for AI Inference
von: Stathopoulou, Foteini, et al.
Veröffentlicht: (2025)
von: Stathopoulou, Foteini, et al.
Veröffentlicht: (2025)
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
LLM-assisted Agentic Edge Intelligence Framework
von: Dehury, Chinmaya Kumar, et al.
Veröffentlicht: (2026)
von: Dehury, Chinmaya Kumar, et al.
Veröffentlicht: (2026)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
von: Papaioannou, Konstantinos, et al.
Veröffentlicht: (2026)
Accelerating the Delivery of Data Services over Uncertain Mobile Crowdsensing Networks
von: Liwang, Minghui, et al.
Veröffentlicht: (2022)
von: Liwang, Minghui, et al.
Veröffentlicht: (2022)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deadline-Driven Hierarchical Agentic Resource Sharing for AI Services and RAN Functions in AI-RAN
von: Li, Haiyuan, et al.
Veröffentlicht: (2026) -
Towards Practical Operation of Deep Reinforcement Learning Agents in Real-World Network Management at Open RAN Edges
von: Li, Haiyuan, et al.
Veröffentlicht: (2024) -
Profiling AI Models: Towards Efficient Computation Offloading in Heterogeneous Edge AI Systems
von: Parra-Ullauri, Juan Marcelo, et al.
Veröffentlicht: (2024) -
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
von: Chen, Shuangyi, et al.
Veröffentlicht: (2024) -
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)