Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
Fuente:
arXiv
Saved in:
| Main Authors: | Dagli, Ismet, Belviranli, Mehmet |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
H-EYE: Holistic Resource Modeling and Management for Diversely Scaled Edge-Cloud Systems
by: Dagli, Ismet, et al.
Published: (2024)
by: Dagli, Ismet, et al.
Published: (2024)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
by: Proaño, Andrès Rubio, et al.
Published: (2024)
by: Proaño, Andrès Rubio, et al.
Published: (2024)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025)
by: Suo, Jiashun, et al.
Published: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
by: Shen, Zixu, et al.
Published: (2025)
by: Shen, Zixu, et al.
Published: (2025)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
by: Taufique, Zain, et al.
Published: (2023)
by: Taufique, Zain, et al.
Published: (2023)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
by: Jayakody, Shakya, et al.
Published: (2026)
by: Jayakody, Shakya, et al.
Published: (2026)
Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum
by: Deng, Akuen Akoi, et al.
Published: (2026)
by: Deng, Akuen Akoi, et al.
Published: (2026)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
by: Chu, Ruifan, et al.
Published: (2025)
by: Chu, Ruifan, et al.
Published: (2025)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
by: Veeravalli, Bharadwaj
Published: (2026)
by: Veeravalli, Bharadwaj
Published: (2026)
Can Large Language Models Predict Parallel Code Performance?
by: Bolet, Gregory, et al.
Published: (2025)
by: Bolet, Gregory, et al.
Published: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
by: Sharma, Chandan, et al.
Published: (2025)
by: Sharma, Chandan, et al.
Published: (2025)
AGOCS -- Accurate Google Cloud Simulator Framework
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
Standardized Methods and Recommendations for Green Federated Learning
by: Tapp, Austin, et al.
Published: (2026)
by: Tapp, Austin, et al.
Published: (2026)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
by: Pinnock, Alyssa, et al.
Published: (2025)
by: Pinnock, Alyssa, et al.
Published: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
Binary Bleed: Fast Distributed and Parallel Method for Automatic Model Selection
by: Barron, Ryan, et al.
Published: (2024)
by: Barron, Ryan, et al.
Published: (2024)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
by: Sun, Bowen, et al.
Published: (2026)
by: Sun, Bowen, et al.
Published: (2026)
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
by: Bolet, Gregory, et al.
Published: (2025)
by: Bolet, Gregory, et al.
Published: (2025)
When AI Bends Metal: AI-Assisted Optimization of Design Parameters in Sheet Metal Forming
by: Tarraf, Ahmad, et al.
Published: (2025)
by: Tarraf, Ahmad, et al.
Published: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
by: Yu, Shan, et al.
Published: (2025)
by: Yu, Shan, et al.
Published: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
by: Zhang, Yaozheng, et al.
Published: (2025)
by: Zhang, Yaozheng, et al.
Published: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)
by: Lei, Yiran, et al.
Published: (2026)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)
by: Yang, Hanmei, et al.
Published: (2024)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
by: Bhattacharjee, Arijit, et al.
Published: (2025)
by: Bhattacharjee, Arijit, et al.
Published: (2025)
Orthrus: Accelerating Multi-BFT Consensus through Concurrent Partial Ordering of Transactions (Extended Version)
by: Lyu, Hanzheng, et al.
Published: (2024)
by: Lyu, Hanzheng, et al.
Published: (2024)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
THAPI: Tracing Heterogeneous APIs
by: Bekele, Solomon, et al.
Published: (2025)
by: Bekele, Solomon, et al.
Published: (2025)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
by: Kabamba, Herve Mbikayi, et al.
Published: (2023)
by: Kabamba, Herve Mbikayi, et al.
Published: (2023)
Scaling Multi Agent Reinforcement Learning for Underwater Acoustic Tracking via Autonomous Vehicles
by: Gallici, Matteo, et al.
Published: (2025)
by: Gallici, Matteo, et al.
Published: (2025)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
by: Jeffery, Andrew, et al.
Published: (2024)
by: Jeffery, Andrew, et al.
Published: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
by: Banegas-Luna, Antonio Jesús, et al.
Published: (2025)
by: Banegas-Luna, Antonio Jesús, et al.
Published: (2025)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
by: Poptani, Akash, et al.
Published: (2025)
by: Poptani, Akash, et al.
Published: (2025)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
by: Grbic, Dragana
Published: (2026)
by: Grbic, Dragana
Published: (2026)
Similar Items
-
H-EYE: Holistic Resource Modeling and Management for Diversely Scaled Edge-Cloud Systems
by: Dagli, Ismet, et al.
Published: (2024) -
Understanding Power Consumption Metric on Heterogeneous Memory Systems
by: Proaño, Andrès Rubio, et al.
Published: (2024) -
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025) -
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
by: Shen, Zixu, et al.
Published: (2025) -
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
by: Taufique, Zain, et al.
Published: (2023)