Agentic Performance at the Edge: Insights from Benchmarking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Shiqiang, Woisetschläger, Herbert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative AI on the Edge: Architecture and Performance Evaluation
von: Nezami, Zeinab, et al.
Veröffentlicht: (2024)
von: Nezami, Zeinab, et al.
Veröffentlicht: (2024)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
von: Zheng, Peirong, et al.
Veröffentlicht: (2026)
von: Zheng, Peirong, et al.
Veröffentlicht: (2026)
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
von: Morabito, Roberto, et al.
Veröffentlicht: (2025)
von: Morabito, Roberto, et al.
Veröffentlicht: (2025)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
Design and Optimization of Hierarchical Gradient Coding for Distributed Learning at Edge Devices
von: Tang, Weiheng, et al.
Veröffentlicht: (2024)
von: Tang, Weiheng, et al.
Veröffentlicht: (2024)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
Towards Edge General Intelligence via Large Language Models: Opportunities and Challenges
von: Chen, Handi, et al.
Veröffentlicht: (2024)
von: Chen, Handi, et al.
Veröffentlicht: (2024)
Adaptive Rank Allocation for Federated Parameter-Efficient Fine-Tuning of Language Models
von: Wu, Fei, et al.
Veröffentlicht: (2025)
von: Wu, Fei, et al.
Veröffentlicht: (2025)
Digital Twinning of a Pressurized Water Reactor Startup Operation and Partial Computational Offloading in In-network Computing-Assisted Multiaccess Edge Computing
von: Aliyu, Ibrahim, et al.
Veröffentlicht: (2024)
von: Aliyu, Ibrahim, et al.
Veröffentlicht: (2024)
eACGM: Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems
von: Xu, Ruilin, et al.
Veröffentlicht: (2025)
von: Xu, Ruilin, et al.
Veröffentlicht: (2025)
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
von: Jang, SiYoung, et al.
Veröffentlicht: (2025)
von: Jang, SiYoung, et al.
Veröffentlicht: (2025)
An Open API Architecture to Discover the Trustworthy Explanation of Cloud AI Services
von: Wang, Zerui, et al.
Veröffentlicht: (2024)
von: Wang, Zerui, et al.
Veröffentlicht: (2024)
When Digital Twin Meets 6G: Concepts, Obstacles, and Research Prospects
von: Liu, Wenshuai, et al.
Veröffentlicht: (2024)
von: Liu, Wenshuai, et al.
Veröffentlicht: (2024)
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
von: Wang, Zhanwei, et al.
Veröffentlicht: (2026)
von: Wang, Zhanwei, et al.
Veröffentlicht: (2026)
NebulaFL: Effective Asynchronous Federated Learning for JointCloud Computing
von: Gao, Fei, et al.
Veröffentlicht: (2024)
von: Gao, Fei, et al.
Veröffentlicht: (2024)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
Edge Graph Intelligence: Reciprocally Empowering Edge Networks with Graph Intelligence
von: Zeng, Liekang, et al.
Veröffentlicht: (2024)
von: Zeng, Liekang, et al.
Veröffentlicht: (2024)
EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement Learning
von: Hao, Yijun, et al.
Veröffentlicht: (2024)
von: Hao, Yijun, et al.
Veröffentlicht: (2024)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
High-speed Networking for Giga-Scale AI Factories
von: Khashab, Sajy, et al.
Veröffentlicht: (2026)
von: Khashab, Sajy, et al.
Veröffentlicht: (2026)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
von: Li, Minghao, et al.
Veröffentlicht: (2026)
von: Li, Minghao, et al.
Veröffentlicht: (2026)
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2026)
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2026)
Collective Communication Profiling of Modern-day Machine Learning Workloads
von: Gupta, Jit, et al.
Veröffentlicht: (2025)
von: Gupta, Jit, et al.
Veröffentlicht: (2025)
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
von: Sun, Tingyang, et al.
Veröffentlicht: (2025)
von: Sun, Tingyang, et al.
Veröffentlicht: (2025)
Context-Aware Orchestration of Energy-Efficient Gossip Learning Schemes
von: Dinani, Mina Aghaei, et al.
Veröffentlicht: (2024)
von: Dinani, Mina Aghaei, et al.
Veröffentlicht: (2024)
Towards Net-Zero Carbon Emissions in Network AI for 6G and Beyond
von: Zhang, Peng, et al.
Veröffentlicht: (2023)
von: Zhang, Peng, et al.
Veröffentlicht: (2023)
Optimizing Split Learning Latency in TinyML-Based IoT Systems
von: Jenhani, Zied, et al.
Veröffentlicht: (2025)
von: Jenhani, Zied, et al.
Veröffentlicht: (2025)
When IoT Meet LLMs: Applications and Challenges
von: Kok, Ibrahim, et al.
Veröffentlicht: (2024)
von: Kok, Ibrahim, et al.
Veröffentlicht: (2024)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2025)
von: Reddy, Tella Rajashekhar, et al.
Veröffentlicht: (2025)
Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
The Implications of Decentralization in Blockchained Federated Learning: Evaluating the Impact of Model Staleness and Inconsistencies
von: Wilhelmi, Francesc, et al.
Veröffentlicht: (2023)
von: Wilhelmi, Francesc, et al.
Veröffentlicht: (2023)
Enabling Reconfiguration-Communication Overlap for Collective Communication in Optical Networks
von: Wu, Changbo, et al.
Veröffentlicht: (2025)
von: Wu, Changbo, et al.
Veröffentlicht: (2025)
Adaptive Parameter-Efficient Federated Fine-Tuning on Heterogeneous Devices
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
Teola: Towards End-to-End Optimization of LLM-based Applications
von: Tan, Xin, et al.
Veröffentlicht: (2024)
von: Tan, Xin, et al.
Veröffentlicht: (2024)
Enabling Intelligent Vehicular Networks Through Distributed Learning in the Non-Terrestrial Networks 6G Vision
von: Naseh, David, et al.
Veröffentlicht: (2023)
von: Naseh, David, et al.
Veröffentlicht: (2023)
Intelligent Task Offloading: Advanced MEC Task Offloading and Resource Management in 5G Networks
von: Ebrahimi, Alireza, et al.
Veröffentlicht: (2025)
von: Ebrahimi, Alireza, et al.
Veröffentlicht: (2025)
Federated Continual Learning for Edge-AI: A Comprehensive Survey
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
Urgent Edge Computing
von: Dazzi, Patrizio, et al.
Veröffentlicht: (2024)
von: Dazzi, Patrizio, et al.
Veröffentlicht: (2024)
UserCentrix: An Agentic Memory-augmented AI Framework for Smart Spaces
von: Saleh, Alaa, et al.
Veröffentlicht: (2025)
von: Saleh, Alaa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generative AI on the Edge: Architecture and Performance Evaluation
von: Nezami, Zeinab, et al.
Veröffentlicht: (2024) -
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
von: Zheng, Peirong, et al.
Veröffentlicht: (2026) -
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
von: Morabito, Roberto, et al.
Veröffentlicht: (2025) -
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026) -
Design and Optimization of Hierarchical Gradient Coding for Distributed Learning at Edge Devices
von: Tang, Weiheng, et al.
Veröffentlicht: (2024)