FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lacouture, Rubens, Zhang, Nathan, Sharma, Ritvik, Siracusa, Marco, Kjolstad, Fredrik, Olukotun, Kunle, Hsu, Olivia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
von: Ko, Sho, et al.
Veröffentlicht: (2024)
von: Ko, Sho, et al.
Veröffentlicht: (2024)
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
von: Ko, Sho, et al.
Veröffentlicht: (2025)
von: Ko, Sho, et al.
Veröffentlicht: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
von: Sohn, Gina, et al.
Veröffentlicht: (2025)
von: Sohn, Gina, et al.
Veröffentlicht: (2025)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
von: Symons, Arne, et al.
Veröffentlicht: (2022)
von: Symons, Arne, et al.
Veröffentlicht: (2022)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
Stream-HLS: Towards Automatic Dataflow Acceleration
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
von: Gilbert, Michael, et al.
Veröffentlicht: (2024)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
von: Yang, Simei, et al.
Veröffentlicht: (2025)
von: Yang, Simei, et al.
Veröffentlicht: (2025)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
von: Yang, Kuilian, et al.
Veröffentlicht: (2026)
von: Yang, Kuilian, et al.
Veröffentlicht: (2026)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
von: Siracusa, Marco, et al.
Veröffentlicht: (2025)
von: Siracusa, Marco, et al.
Veröffentlicht: (2025)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
von: Qin, Shantian, et al.
Veröffentlicht: (2025)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
SpChar: Characterizing the Sparse Puzzle via Decision Trees
von: Sgherzi, Francesco, et al.
Veröffentlicht: (2023)
von: Sgherzi, Francesco, et al.
Veröffentlicht: (2023)
A Low-Power Sparse Deep Learning Accelerator with Optimized Data Reuse
von: Hsu, Kai-Chieh, et al.
Veröffentlicht: (2025)
von: Hsu, Kai-Chieh, et al.
Veröffentlicht: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
Fast Cross-Operator Optimization of Attention Dataflow
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
von: Gu, Hang, et al.
Veröffentlicht: (2026)
von: Gu, Hang, et al.
Veröffentlicht: (2026)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads
von: Sharma, Harsh, et al.
Veröffentlicht: (2024)
von: Sharma, Harsh, et al.
Veröffentlicht: (2024)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
von: Deshmukh, Sameer, et al.
Veröffentlicht: (2024)
von: Deshmukh, Sameer, et al.
Veröffentlicht: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
von: Zhao, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaotian, et al.
Veröffentlicht: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
von: Umuroglu, Yaman, et al.
Veröffentlicht: (2025)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
von: Chen, Yuehai, et al.
Veröffentlicht: (2025)
von: Chen, Yuehai, et al.
Veröffentlicht: (2025)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
von: Geens, Robin, et al.
Veröffentlicht: (2025)
von: Geens, Robin, et al.
Veröffentlicht: (2025)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
von: Tong, Jianming, et al.
Veröffentlicht: (2024)
von: Tong, Jianming, et al.
Veröffentlicht: (2024)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
von: Shi, Shangyi, et al.
Veröffentlicht: (2025)
von: Shi, Shangyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023) -
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
von: Sohn, Gina, et al.
Veröffentlicht: (2024) -
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
von: Ko, Sho, et al.
Veröffentlicht: (2024) -
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
von: Ko, Sho, et al.
Veröffentlicht: (2025) -
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)