Embedded Distributed Inference of Deep Neural Networks: A Systematic Review
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peccia, Federico Nicolás, Bringmann, Oliver |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
von: Kreß, Fabian, et al.
Veröffentlicht: (2024)
von: Kreß, Fabian, et al.
Veröffentlicht: (2024)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
Optimizing Distributed Training Approaches for Scaling Neural Networks
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)
OptimES: Optimizing Federated Learning Using Remote Embeddings for Graph Neural Networks
von: Naman, Pranjal, et al.
Veröffentlicht: (2025)
von: Naman, Pranjal, et al.
Veröffentlicht: (2025)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
von: Ma, Xiaolong, et al.
Veröffentlicht: (2024)
von: Ma, Xiaolong, et al.
Veröffentlicht: (2024)
Spatiotemporal Traffic Prediction in Distributed Backend Systems via Graph Neural Networks
von: Qiu, Zhimin, et al.
Veröffentlicht: (2025)
von: Qiu, Zhimin, et al.
Veröffentlicht: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
von: He, Pujiang, et al.
Veröffentlicht: (2024)
von: He, Pujiang, et al.
Veröffentlicht: (2024)
Big Data Intelligence Using Distributed Deep Neural Networks
von: Ongati, Felix, et al.
Veröffentlicht: (2019)
von: Ongati, Felix, et al.
Veröffentlicht: (2019)
KV Cache Compression for Inference Efficiency in LLMs: A Review
von: Liu, Yanyu, et al.
Veröffentlicht: (2025)
von: Liu, Yanyu, et al.
Veröffentlicht: (2025)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
Edge AI: A Taxonomy, Systematic Review and Future Directions
von: Gill, Sukhpal Singh, et al.
Veröffentlicht: (2024)
von: Gill, Sukhpal Singh, et al.
Veröffentlicht: (2024)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Argus: Token Aware Distributed LLM Inference Optimization
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
von: Wu, Panlong, et al.
Veröffentlicht: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Sharding Distributed Databases: A Critical Review
von: Solat, Siamak
Veröffentlicht: (2024)
von: Solat, Siamak
Veröffentlicht: (2024)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
Characterizing Communication Patterns in Distributed Large Language Model Inference
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
von: Duan, Jiaang, et al.
Veröffentlicht: (2024)
von: Duan, Jiaang, et al.
Veröffentlicht: (2024)
Blockchain and Edge Computing Nexus: A Large-scale Systematic Literature Review
von: Nezami, Zeinab, et al.
Veröffentlicht: (2025)
von: Nezami, Zeinab, et al.
Veröffentlicht: (2025)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
Federated Neural Radiance Field for Distributed Intelligence
von: Zhang, Yintian, et al.
Veröffentlicht: (2024)
von: Zhang, Yintian, et al.
Veröffentlicht: (2024)
Cold Start Latency in Serverless Computing: A Systematic Review, Taxonomy, and Future Directions
von: Golec, Muhammed, et al.
Veröffentlicht: (2023)
von: Golec, Muhammed, et al.
Veröffentlicht: (2023)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
von: Chen, Jiu, et al.
Veröffentlicht: (2026)
von: Chen, Jiu, et al.
Veröffentlicht: (2026)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
von: Peng, Xian, et al.
Veröffentlicht: (2025)
von: Peng, Xian, et al.
Veröffentlicht: (2025)
A Systematic Literature Review on Task Allocation and Performance Management Techniques in Cloud Data Center
von: Chauhan, Nidhika, et al.
Veröffentlicht: (2024)
von: Chauhan, Nidhika, et al.
Veröffentlicht: (2024)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
von: Wu, Tian, et al.
Veröffentlicht: (2025)
von: Wu, Tian, et al.
Veröffentlicht: (2025)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
von: Khoshsirat, Aria, et al.
Veröffentlicht: (2024)
A Framework for Hybrid Collective Inference in Distributed Sensor Networks
von: Nash, Andrew, et al.
Veröffentlicht: (2026)
von: Nash, Andrew, et al.
Veröffentlicht: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
von: Zhang, Runhua, et al.
Veröffentlicht: (2025)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
von: Chen, June, et al.
Veröffentlicht: (2026)
von: Chen, June, et al.
Veröffentlicht: (2026)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
von: Zhang, Songge, et al.
Veröffentlicht: (2026)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
von: Peccia, Federico Nicolas, et al.
Veröffentlicht: (2024) -
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
von: Kreß, Fabian, et al.
Veröffentlicht: (2024) -
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024) -
Optimizing Distributed Training Approaches for Scaling Neural Networks
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025) -
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
von: Naman, Pranjal, et al.
Veröffentlicht: (2026)