Fast Cross-Operator Optimization of Attention Dataflow
Fuente:
arXiv
Salvato in:
| Autori principali: | Chang, Haodong, Hu, Hailiang, Wang, Zhenrui, Gong, Yu, Liang, Rongjian, Tang, Zhexiang, Yuan, Bo, Hu, Jiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
di: Sohn, Gina, et al.
Pubblicazione: (2024)
di: Sohn, Gina, et al.
Pubblicazione: (2024)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
di: Zhang, Weichuang, et al.
Pubblicazione: (2026)
di: Zhang, Weichuang, et al.
Pubblicazione: (2026)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
di: Chang, Kaiyan, et al.
Pubblicazione: (2025)
di: Chang, Kaiyan, et al.
Pubblicazione: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
di: Zhang, Chi, et al.
Pubblicazione: (2026)
di: Zhang, Chi, et al.
Pubblicazione: (2026)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
di: Chen, Yuehai, et al.
Pubblicazione: (2025)
di: Chen, Yuehai, et al.
Pubblicazione: (2025)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
di: Shi, Shangyi, et al.
Pubblicazione: (2025)
di: Shi, Shangyi, et al.
Pubblicazione: (2025)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
di: Gu, Hang, et al.
Pubblicazione: (2026)
di: Gu, Hang, et al.
Pubblicazione: (2026)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
di: He, Xiaolin, et al.
Pubblicazione: (2025)
di: He, Xiaolin, et al.
Pubblicazione: (2025)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
di: Ko, Sho, et al.
Pubblicazione: (2024)
di: Ko, Sho, et al.
Pubblicazione: (2024)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
di: Yang, Simei, et al.
Pubblicazione: (2025)
di: Yang, Simei, et al.
Pubblicazione: (2025)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
di: Hsiung, Ching-Lin, et al.
Pubblicazione: (2025)
di: Hsiung, Ching-Lin, et al.
Pubblicazione: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
di: Umuroglu, Yaman, et al.
Pubblicazione: (2025)
di: Umuroglu, Yaman, et al.
Pubblicazione: (2025)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
di: Han, Meng, et al.
Pubblicazione: (2023)
di: Han, Meng, et al.
Pubblicazione: (2023)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
di: Wu, Junyi, et al.
Pubblicazione: (2025)
di: Wu, Junyi, et al.
Pubblicazione: (2025)
Revet: A Language and Compiler for Dataflow Threads
di: Rucker, Alexander, et al.
Pubblicazione: (2023)
di: Rucker, Alexander, et al.
Pubblicazione: (2023)
Stream-HLS: Towards Automatic Dataflow Acceleration
di: Basalama, Suhail, et al.
Pubblicazione: (2025)
di: Basalama, Suhail, et al.
Pubblicazione: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
di: Wang, Zhican, et al.
Pubblicazione: (2025)
di: Wang, Zhican, et al.
Pubblicazione: (2025)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
di: Chen, Xingzhen, et al.
Pubblicazione: (2026)
di: Chen, Xingzhen, et al.
Pubblicazione: (2026)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
di: Gupta, Neelesh, et al.
Pubblicazione: (2026)
di: Gupta, Neelesh, et al.
Pubblicazione: (2026)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
di: He, Siyuan, et al.
Pubblicazione: (2025)
di: He, Siyuan, et al.
Pubblicazione: (2025)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
di: Xue, Chenhao, et al.
Pubblicazione: (2026)
di: Xue, Chenhao, et al.
Pubblicazione: (2026)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
di: Qin, Shantian, et al.
Pubblicazione: (2025)
di: Qin, Shantian, et al.
Pubblicazione: (2025)
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
di: Zhao, Xiaotian, et al.
Pubblicazione: (2025)
di: Zhao, Xiaotian, et al.
Pubblicazione: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
di: Ye, Hanchen, et al.
Pubblicazione: (2025)
di: Ye, Hanchen, et al.
Pubblicazione: (2025)
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
di: Jiang, Wenqi, et al.
Pubblicazione: (2024)
di: Jiang, Wenqi, et al.
Pubblicazione: (2024)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
di: Zhang, Junming, et al.
Pubblicazione: (2026)
di: Zhang, Junming, et al.
Pubblicazione: (2026)
Dataflow Optimized Reconfigurable Acceleration for FEM-based CFD Simulations
di: Kapetanakis, Anastassis, et al.
Pubblicazione: (2024)
di: Kapetanakis, Anastassis, et al.
Pubblicazione: (2024)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
di: Gilbert, Michael, et al.
Pubblicazione: (2024)
di: Gilbert, Michael, et al.
Pubblicazione: (2024)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2024)
di: Zhao, Zhiyuan, et al.
Pubblicazione: (2024)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
di: Zheng, Jianing, et al.
Pubblicazione: (2025)
di: Zheng, Jianing, et al.
Pubblicazione: (2025)
SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated Tiling
di: Wang, Huizheng, et al.
Pubblicazione: (2024)
di: Wang, Huizheng, et al.
Pubblicazione: (2024)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
di: Wang, Yimin, et al.
Pubblicazione: (2025)
di: Wang, Yimin, et al.
Pubblicazione: (2025)
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
di: Deshmukh, Sameer, et al.
Pubblicazione: (2024)
di: Deshmukh, Sameer, et al.
Pubblicazione: (2024)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
di: Symons, Arne, et al.
Pubblicazione: (2022)
di: Symons, Arne, et al.
Pubblicazione: (2022)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
di: Yin, Ruokai, et al.
Pubblicazione: (2025)
di: Yin, Ruokai, et al.
Pubblicazione: (2025)
Learning Library Cell Representations in Vector Space
di: Liang, Rongjian, et al.
Pubblicazione: (2025)
di: Liang, Rongjian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
di: Sohn, Gina, et al.
Pubblicazione: (2024) -
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
di: Zhang, Weichuang, et al.
Pubblicazione: (2026) -
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
di: Chang, Kaiyan, et al.
Pubblicazione: (2025) -
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
di: Zhang, Chi, et al.
Pubblicazione: (2026) -
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
di: Wu, Haibin, et al.
Pubblicazione: (2024)