Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
Fuente:
arXiv
Salvato in:
| Autori principali: | Arya, Mayank, Simmhan, Yogesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2023)
di: K., Prashanthi S., et al.
Pubblicazione: (2023)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
di: Naman, Pranjal, et al.
Pubblicazione: (2025)
di: Naman, Pranjal, et al.
Pubblicazione: (2025)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
di: Naman, Pranjal, et al.
Pubblicazione: (2026)
di: Naman, Pranjal, et al.
Pubblicazione: (2026)
Federated Learning within Global Energy Budget over Heterogeneous Edge Accelerators
di: Banerjee, Roopkatha, et al.
Pubblicazione: (2025)
di: Banerjee, Roopkatha, et al.
Pubblicazione: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
di: Raj, Suman, et al.
Pubblicazione: (2024)
di: Raj, Suman, et al.
Pubblicazione: (2024)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
OptimES: Optimizing Federated Learning Using Remote Embeddings for Graph Neural Networks
di: Naman, Pranjal, et al.
Pubblicazione: (2025)
di: Naman, Pranjal, et al.
Pubblicazione: (2025)
RIPPLE++: An Incremental Framework for Efficient GNN Inference on Evolving Graphs
di: Naman, Pranjal, et al.
Pubblicazione: (2026)
di: Naman, Pranjal, et al.
Pubblicazione: (2026)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
PowerTrain: Fast, Generalizable Time and Power Prediction Models to Optimize DNN Training on Accelerated Edges
di: K., Prashanthi S., et al.
Pubblicazione: (2024)
di: K., Prashanthi S., et al.
Pubblicazione: (2024)
AeroGen: Agentic Drone Autonomy through Single-Shot Structured Prompting & Drone SDK
di: Astu, Kautuk, et al.
Pubblicazione: (2026)
di: Astu, Kautuk, et al.
Pubblicazione: (2026)
AerialDB: A Federated Peer-to-Peer Spatio-temporal Edge Datastore for Drone Fleets
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
Scaling Real-Time Traffic Analytics on Edge-Cloud Fabrics for City-Scale Camera Networks
di: Sharma, Akash, et al.
Pubblicazione: (2026)
di: Sharma, Akash, et al.
Pubblicazione: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
di: Ng, Nathan, et al.
Pubblicazione: (2026)
di: Ng, Nathan, et al.
Pubblicazione: (2026)
Optimizing Federated Learning using Remote Embeddings for Graph Neural Networks
di: Naman, Pranjal, et al.
Pubblicazione: (2025)
di: Naman, Pranjal, et al.
Pubblicazione: (2025)
AeroDaaS: Towards an Application Programming Framework for Drones-as-a-Service
di: Raj, Suman, et al.
Pubblicazione: (2025)
di: Raj, Suman, et al.
Pubblicazione: (2025)
AeroDaaS: A Programmable Drones-as-a-Service Platform for Intelligent Aerial Systems
di: Astu, Kautuk, et al.
Pubblicazione: (2026)
di: Astu, Kautuk, et al.
Pubblicazione: (2026)
AeroResQ: Edge-Accelerated UAV Framework for Scalable, Resilient and Collaborative Escape Route Planning in Wildfire Scenarios
di: Raj, Suman, et al.
Pubblicazione: (2025)
di: Raj, Suman, et al.
Pubblicazione: (2025)
A Blockchain-Enabled Framework for Storage and Retrieval of Social Data
di: Parab, Aishwarya, et al.
Pubblicazione: (2025)
di: Parab, Aishwarya, et al.
Pubblicazione: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
di: Sun, Mingyu, et al.
Pubblicazione: (2025)
di: Sun, Mingyu, et al.
Pubblicazione: (2025)
End-to-End and Phase-Level Performance Optimization for Hyperledger Fabric
di: Sollu, Pavan, et al.
Pubblicazione: (2026)
di: Sollu, Pavan, et al.
Pubblicazione: (2026)
Characterizing FaaS Workflows on Public Clouds: The Good, the Bad and the Ugly
di: Kulkarni, Varad, et al.
Pubblicazione: (2025)
di: Kulkarni, Varad, et al.
Pubblicazione: (2025)
Optimizing FaaS Platforms for MCP-enabled Agentic Workflows
di: Kulkarni, Varad, et al.
Pubblicazione: (2026)
di: Kulkarni, Varad, et al.
Pubblicazione: (2026)
D3FL: Data Distribution and Detrending for Robust Federated Learning in Non-linear Time-series Data
di: Marisetty, Harsha Varun, et al.
Pubblicazione: (2025)
di: Marisetty, Harsha Varun, et al.
Pubblicazione: (2025)
AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
di: Tokal, Shiva Sai Krishna Anand, et al.
Pubblicazione: (2025)
di: Tokal, Shiva Sai Krishna Anand, et al.
Pubblicazione: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
di: Zhang, Mingjin, et al.
Pubblicazione: (2024)
di: Zhang, Mingjin, et al.
Pubblicazione: (2024)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
di: Raj, Suman, et al.
Pubblicazione: (2025)
di: Raj, Suman, et al.
Pubblicazione: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
Optimizing Federated Learning for Scalable Power-demand Forecasting in Microgrids
di: Banerjee, Roopkatha, et al.
Pubblicazione: (2025)
di: Banerjee, Roopkatha, et al.
Pubblicazione: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
di: Khoshsirat, Aria, et al.
Pubblicazione: (2024)
di: Khoshsirat, Aria, et al.
Pubblicazione: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
di: Wu, Tian, et al.
Pubblicazione: (2025)
di: Wu, Tian, et al.
Pubblicazione: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
di: Chow, Will
Pubblicazione: (2025)
di: Chow, Will
Pubblicazione: (2025)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
di: Bournias, Ilias, et al.
Pubblicazione: (2024)
di: Bournias, Ilias, et al.
Pubblicazione: (2024)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
di: Zhang, Runhua, et al.
Pubblicazione: (2025)
di: Zhang, Runhua, et al.
Pubblicazione: (2025)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
di: Butler, Branden, et al.
Pubblicazione: (2024)
di: Butler, Branden, et al.
Pubblicazione: (2024)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
di: Wei, Jinhui, et al.
Pubblicazione: (2025)
di: Wei, Jinhui, et al.
Pubblicazione: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
di: Yu, Shibo, et al.
Pubblicazione: (2025)
di: Yu, Shibo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025) -
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2023) -
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
di: K., Prashanthi S., et al.
Pubblicazione: (2025) -
Ripple: Scalable Incremental GNN Inferencing on Large Streaming Graphs
di: Naman, Pranjal, et al.
Pubblicazione: (2025) -
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
di: Naman, Pranjal, et al.
Pubblicazione: (2026)