A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Puigdemont, Pol, Russo, Enrico, Wassington, Axel, Das, Abhijit, Abadal, Sergi, Palesi, Maurizio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators
von: Das, Abhijit, et al.
Veröffentlicht: (2022)
von: Das, Abhijit, et al.
Veröffentlicht: (2022)
CHAOS: Controlled Hardware fAult injectOr System for gem5
von: Vinciguerra, Elio, et al.
Veröffentlicht: (2026)
von: Vinciguerra, Elio, et al.
Veröffentlicht: (2026)
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
von: Blanco, Francesco G., et al.
Veröffentlicht: (2024)
von: Blanco, Francesco G., et al.
Veröffentlicht: (2024)
Communication Characterization of AI Workloads for Large-scale Multi-chiplet Accelerators
von: Musavi, Mariam, et al.
Veröffentlicht: (2024)
von: Musavi, Mariam, et al.
Veröffentlicht: (2024)
Exploring the Potential of Wireless-enabled Multi-Chip AI Accelerators
von: Irabor, Emmanuel, et al.
Veröffentlicht: (2025)
von: Irabor, Emmanuel, et al.
Veröffentlicht: (2025)
Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
von: Russo, Enrico, et al.
Veröffentlicht: (2024)
von: Russo, Enrico, et al.
Veröffentlicht: (2024)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
von: Li, Fang
Veröffentlicht: (2025)
von: Li, Fang
Veröffentlicht: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
Time Reversal for Near-Field Communications on Multi-chip Wireless Networks
von: Rodríguez-Galán, Fátima, et al.
Veröffentlicht: (2024)
von: Rodríguez-Galán, Fátima, et al.
Veröffentlicht: (2024)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
von: Zhao, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaotian, et al.
Veröffentlicht: (2025)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
von: Yang, Kuilian, et al.
Veröffentlicht: (2026)
von: Yang, Kuilian, et al.
Veröffentlicht: (2026)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
von: Chen, Yuehai, et al.
Veröffentlicht: (2025)
von: Chen, Yuehai, et al.
Veröffentlicht: (2025)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
Fast Cross-Operator Optimization of Attention Dataflow
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
Stream-HLS: Towards Automatic Dataflow Acceleration
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
von: Lin, Kuan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Kuan-Ting, et al.
Veröffentlicht: (2025)
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
von: Yang, Simei, et al.
Veröffentlicht: (2025)
von: Yang, Simei, et al.
Veröffentlicht: (2025)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
von: Tong, Jianming, et al.
Veröffentlicht: (2024)
von: Tong, Jianming, et al.
Veröffentlicht: (2024)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2026)
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2026)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
von: Ganti, Ravindra, et al.
Veröffentlicht: (2025)
von: Ganti, Ravindra, et al.
Veröffentlicht: (2025)
A Hierarchical Dataflow-Driven Heterogeneous Architecture for Wireless Baseband Processing
von: Jiang, Limin, et al.
Veröffentlicht: (2024)
von: Jiang, Limin, et al.
Veröffentlicht: (2024)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
von: Gu, Hang, et al.
Veröffentlicht: (2026)
von: Gu, Hang, et al.
Veröffentlicht: (2026)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators
von: Das, Abhijit, et al.
Veröffentlicht: (2022) -
CHAOS: Controlled Hardware fAult injectOr System for gem5
von: Vinciguerra, Elio, et al.
Veröffentlicht: (2026) -
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
von: Blanco, Francesco G., et al.
Veröffentlicht: (2024) -
Communication Characterization of AI Workloads for Large-scale Multi-chiplet Accelerators
von: Musavi, Mariam, et al.
Veröffentlicht: (2024) -
Exploring the Potential of Wireless-enabled Multi-Chip AI Accelerators
von: Irabor, Emmanuel, et al.
Veröffentlicht: (2025)