Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Colleman, Steven, Symons, Arne, Jung, Victor J. B., Verhelst, Marian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
von: Symons, Arne, et al.
Veröffentlicht: (2022)
von: Symons, Arne, et al.
Veröffentlicht: (2022)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
von: Jung, Victor J. B., et al.
Veröffentlicht: (2023)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2023)
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
von: Geens, Robin, et al.
Veröffentlicht: (2025)
von: Geens, Robin, et al.
Veröffentlicht: (2025)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
von: Zniber, Alaa, et al.
Veröffentlicht: (2025)
von: Zniber, Alaa, et al.
Veröffentlicht: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
von: Geens, Robin, et al.
Veröffentlicht: (2025)
von: Geens, Robin, et al.
Veröffentlicht: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
von: Houshmand, Pouya, et al.
Veröffentlicht: (2024)
von: Houshmand, Pouya, et al.
Veröffentlicht: (2024)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
von: Fang, Chao, et al.
Veröffentlicht: (2024)
von: Fang, Chao, et al.
Veröffentlicht: (2024)
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
von: Geens, Robin, et al.
Veröffentlicht: (2026)
von: Geens, Robin, et al.
Veröffentlicht: (2026)
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
von: Geens, Robin, et al.
Veröffentlicht: (2026)
von: Geens, Robin, et al.
Veröffentlicht: (2026)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2024)
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2024)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2025)
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
von: Cuyckens, Stef, et al.
Veröffentlicht: (2025)
von: Cuyckens, Stef, et al.
Veröffentlicht: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
von: Yi, Xiaoling, et al.
Veröffentlicht: (2026)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2026)
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
von: Cuyckens, Stef, et al.
Veröffentlicht: (2025)
von: Cuyckens, Stef, et al.
Veröffentlicht: (2025)
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
von: Bochem, Severin, et al.
Veröffentlicht: (2024)
von: Bochem, Severin, et al.
Veröffentlicht: (2024)
IMMSched: Interruptible Multi-DNN Scheduling via Parallel Multi-Particle Optimizing Subgraph Isomorphism
von: Zhao, Boran, et al.
Veröffentlicht: (2026)
von: Zhao, Boran, et al.
Veröffentlicht: (2026)
Analog or Digital In-memory Computing? Benchmarking through Quantitative Modeling
von: Sun, Jiacong, et al.
Veröffentlicht: (2024)
von: Sun, Jiacong, et al.
Veröffentlicht: (2024)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
von: Antonio, Ryan Albert, et al.
Veröffentlicht: (2025)
von: Antonio, Ryan Albert, et al.
Veröffentlicht: (2025)
FuseMax: Leveraging Extended Einsums to Optimize Attention Accelerator Design
von: Nayak, Nandeeka, et al.
Veröffentlicht: (2024)
von: Nayak, Nandeeka, et al.
Veröffentlicht: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
von: Yang, Simei, et al.
Veröffentlicht: (2025)
von: Yang, Simei, et al.
Veröffentlicht: (2025)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
Retrieve, Schedule, Reflect: LLM Agents for Chip QoR Optimization
von: ouyang, Yikang, et al.
Veröffentlicht: (2026)
von: ouyang, Yikang, et al.
Veröffentlicht: (2026)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
von: Jia, Shuao, et al.
Veröffentlicht: (2025)
von: Jia, Shuao, et al.
Veröffentlicht: (2025)
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
von: Yi, Xiaoling, et al.
Veröffentlicht: (2024)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2024)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
Predictive Software Scheduling as an Early-Warning Hint Layer for Optical Engine Thermal Drift in Heterogeneous SoIC Packaging
von: Chung, Chi Fei
Veröffentlicht: (2026)
von: Chung, Chi Fei
Veröffentlicht: (2026)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
von: Wang, Bowen, et al.
Veröffentlicht: (2025)
von: Wang, Bowen, et al.
Veröffentlicht: (2025)
MultiVic: A Time-Predictable RISC-V Multi-Core Processor Optimized for Neural Network Inference
von: Kirschner, Maximilian, et al.
Veröffentlicht: (2025)
von: Kirschner, Maximilian, et al.
Veröffentlicht: (2025)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
von: Li, Meng, et al.
Veröffentlicht: (2026)
von: Li, Meng, et al.
Veröffentlicht: (2026)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
von: Symons, Arne, et al.
Veröffentlicht: (2022) -
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
von: Jung, Victor J. B., et al.
Veröffentlicht: (2023) -
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
von: Geens, Robin, et al.
Veröffentlicht: (2025) -
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
von: Mei, Linyan, et al.
Veröffentlicht: (2022) -
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
von: Zniber, Alaa, et al.
Veröffentlicht: (2025)