A Structure-Aware Framework for Learning Device Placements on Computation Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Shukai, Ping, Heng, Kanakaris, Nikos, Xiao, Xiongye, Kyriakis, Panagiotis, Ahmed, Nesreen K., Zhang, Peiyu, Ma, Guixiang, Capota, Mihai, Nazarian, Shahin, Willke, Theodore L., Bogdan, Paul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PerfRL: A Small Language Model Framework for Efficient Code Optimization
by: Duan, Shukai, et al.
Published: (2023)
by: Duan, Shukai, et al.
Published: (2023)
HDLCoRe: A Training-Free Framework for Mitigating Hallucinations in LLM-Generated HDL
by: Ping, Heng, et al.
Published: (2025)
by: Ping, Heng, et al.
Published: (2025)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
by: Bhattacharjee, Arijit, et al.
Published: (2025)
by: Bhattacharjee, Arijit, et al.
Published: (2025)
Network-informed Prompt Engineering against Organized Astroturf Campaigns under Extreme Class Imbalance
by: Kanakaris, Nikos, et al.
Published: (2025)
by: Kanakaris, Nikos, et al.
Published: (2025)
Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
by: Cheng, Anzhe, et al.
Published: (2025)
by: Cheng, Anzhe, et al.
Published: (2025)
AutoParLLM: GNN-guided Context Generation for Zero-Shot Code Parallelization using LLMs
by: Mahmud, Quazi Ishtiaque, et al.
Published: (2023)
by: Mahmud, Quazi Ishtiaque, et al.
Published: (2023)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025)
by: Ahmadvand, Hossein, et al.
Published: (2025)
Modeling Tradeoffs between mobility, cost, and performance in Edge Computing
by: Waseem, Muhammad Danish, et al.
Published: (2026)
by: Waseem, Muhammad Danish, et al.
Published: (2026)
Who Will Be Accountable for Accountability?
by: Dolmatch, Theodore B.
Published: (1970)
by: Dolmatch, Theodore B.
Published: (1970)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
by: Xu, Mingbin, et al.
Published: (2023)
by: Xu, Mingbin, et al.
Published: (2023)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025)
by: Zhang, Niansong, et al.
Published: (2025)
H$^2$GFM: Towards unifying Homogeneity and Heterogeneity on Text-Attributed Graphs
by: Nguyen, Trung-Kien, et al.
Published: (2025)
by: Nguyen, Trung-Kien, et al.
Published: (2025)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
by: Uhlig, Arno, et al.
Published: (2025)
by: Uhlig, Arno, et al.
Published: (2025)
Spatiotemporal Non-Uniformity-Aware Online Task Scheduling in Collaborative Edge Computing for Industrial Internet of Things
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models
by: Xiao, Xiongye, et al.
Published: (2024)
by: Xiao, Xiongye, et al.
Published: (2024)
Efficient Data-Driven Production Scheduling in Pharmaceutical Manufacturing
by: Balatsos, Ioannis, et al.
Published: (2026)
by: Balatsos, Ioannis, et al.
Published: (2026)
POET: Power-Oriented Evolutionary Tuning for LLM-Based RTL PPA Optimization
by: Ping, Heng, et al.
Published: (2026)
by: Ping, Heng, et al.
Published: (2026)
Energy-Aware Computing in the Year 2026
by: Tchakoute, Roblex Nana, et al.
Published: (2026)
by: Tchakoute, Roblex Nana, et al.
Published: (2026)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Spatially Correlated multi-RIS Communication: The Effect of Inter-Operator Interference
by: Miridakis, Nikolaos I., et al.
Published: (2025)
by: Miridakis, Nikolaos I., et al.
Published: (2025)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
by: Dong, Ximing, et al.
Published: (2026)
by: Dong, Ximing, et al.
Published: (2026)
Size-Aware Dispatching to Fluid Queues
by: Xie, Runhan, et al.
Published: (2025)
by: Xie, Runhan, et al.
Published: (2025)
Redundant Array Computation Elimination
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
Rethinking Temporal Models for TinyML: LSTM versus 1D-CNN in Resource-Constrained Devices
by: Saha, Bidyut, et al.
Published: (2026)
by: Saha, Bidyut, et al.
Published: (2026)
Structure Guided Prompt: Instructing Large Language Model in Multi-Step Reasoning by Exploring Graph Structure of the Text
by: Cheng, Kewei, et al.
Published: (2024)
by: Cheng, Kewei, et al.
Published: (2024)
Evaluating Learning Congestion control Schemes for LEO Constellations
by: Mazilu, Mihai, et al.
Published: (2025)
by: Mazilu, Mihai, et al.
Published: (2025)
An Experimental Study of Low-Latency Video Streaming over 5G
by: Khan, Imran, et al.
Published: (2024)
by: Khan, Imran, et al.
Published: (2024)
Performance Characterization of Containers in Edge Computing
by: Gupta, Ragini, et al.
Published: (2025)
by: Gupta, Ragini, et al.
Published: (2025)
Heavy-Traffic Optimal Size- and State-Aware Dispatching
by: Xie, Runhan, et al.
Published: (2023)
by: Xie, Runhan, et al.
Published: (2023)
CARINA: Carbon-Aware Execution of Recurrent Industrial Analytics
by: Farooq, Muhammad Umar
Published: (2026)
by: Farooq, Muhammad Umar
Published: (2026)
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
by: Hossain, Abrar, et al.
Published: (2025)
by: Hossain, Abrar, et al.
Published: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
Towards A Flexible Accuracy-Oriented Deep Learning Module Inference Latency Prediction Framework for Adaptive Optimization Algorithms
by: Shen, Jingran, et al.
Published: (2023)
by: Shen, Jingran, et al.
Published: (2023)
EMoE: Eigenbasis-Guided Routing for Mixture-of-Experts
by: Cheng, Anzhe, et al.
Published: (2026)
by: Cheng, Anzhe, et al.
Published: (2026)
Dependability of UAV-Based Networks and Computing Systems: A Survey
by: Zhang, Qingyang, et al.
Published: (2025)
by: Zhang, Qingyang, et al.
Published: (2025)
Similar Items
-
PerfRL: A Small Language Model Framework for Efficient Code Optimization
by: Duan, Shukai, et al.
Published: (2023) -
HDLCoRe: A Training-Free Framework for Mitigating Hallucinations in LLM-Generated HDL
by: Ping, Heng, et al.
Published: (2025) -
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
by: Bhattacharjee, Arijit, et al.
Published: (2025) -
Network-informed Prompt Engineering against Organized Astroturf Campaigns under Extreme Class Imbalance
by: Kanakaris, Nikos, et al.
Published: (2025) -
Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
by: Xiao, Yao, et al.
Published: (2025)