SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Ko, Sho, Olukotun, Kunle |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
por: Ko, Sho, et al.
Publicado: (2024)
por: Ko, Sho, et al.
Publicado: (2024)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
por: Sohn, Gina, et al.
Publicado: (2024)
por: Sohn, Gina, et al.
Publicado: (2024)
Revet: A Language and Compiler for Dataflow Threads
por: Rucker, Alexander, et al.
Publicado: (2023)
por: Rucker, Alexander, et al.
Publicado: (2023)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
por: Lacouture, Rubens, et al.
Publicado: (2025)
por: Lacouture, Rubens, et al.
Publicado: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
por: Kim, Jungwoo, et al.
Publicado: (2026)
por: Kim, Jungwoo, et al.
Publicado: (2026)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
por: Zhou, Zikai, et al.
Publicado: (2025)
por: Zhou, Zikai, et al.
Publicado: (2025)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
por: Tong, Jianming, et al.
Publicado: (2024)
por: Tong, Jianming, et al.
Publicado: (2024)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
por: Gilbert, Michael, et al.
Publicado: (2024)
por: Gilbert, Michael, et al.
Publicado: (2024)
Dataflow Optimized Reconfigurable Acceleration for FEM-based CFD Simulations
por: Kapetanakis, Anastassis, et al.
Publicado: (2024)
por: Kapetanakis, Anastassis, et al.
Publicado: (2024)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
por: Symons, Arne, et al.
Publicado: (2022)
por: Symons, Arne, et al.
Publicado: (2022)
Stream-HLS: Towards Automatic Dataflow Acceleration
por: Basalama, Suhail, et al.
Publicado: (2025)
por: Basalama, Suhail, et al.
Publicado: (2025)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
por: Zhang, Weichuang, et al.
Publicado: (2026)
por: Zhang, Weichuang, et al.
Publicado: (2026)
Fast Cross-Operator Optimization of Attention Dataflow
por: Chang, Haodong, et al.
Publicado: (2026)
por: Chang, Haodong, et al.
Publicado: (2026)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
por: Chang, Kaiyan, et al.
Publicado: (2025)
por: Chang, Kaiyan, et al.
Publicado: (2025)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
por: Chen, Xingzhen, et al.
Publicado: (2026)
por: Chen, Xingzhen, et al.
Publicado: (2026)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
por: Zheng, Jianing, et al.
Publicado: (2025)
por: Zheng, Jianing, et al.
Publicado: (2025)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
por: Gu, Hang, et al.
Publicado: (2026)
por: Gu, Hang, et al.
Publicado: (2026)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
por: Zhao, Zhiyuan, et al.
Publicado: (2024)
por: Zhao, Zhiyuan, et al.
Publicado: (2024)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
por: Sohn, Gina, et al.
Publicado: (2025)
por: Sohn, Gina, et al.
Publicado: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
por: Ye, Hanchen, et al.
Publicado: (2025)
por: Ye, Hanchen, et al.
Publicado: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
por: Li, Jonathan, et al.
Publicado: (2025)
por: Li, Jonathan, et al.
Publicado: (2025)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
por: Yin, Ruokai, et al.
Publicado: (2025)
por: Yin, Ruokai, et al.
Publicado: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
por: Cheng, Jianyi, et al.
Publicado: (2023)
por: Cheng, Jianyi, et al.
Publicado: (2023)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
por: Xue, Chenhao, et al.
Publicado: (2026)
por: Xue, Chenhao, et al.
Publicado: (2026)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
por: He, Xiaolin, et al.
Publicado: (2025)
por: He, Xiaolin, et al.
Publicado: (2025)
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
por: Zhao, Xiaotian, et al.
Publicado: (2025)
por: Zhao, Xiaotian, et al.
Publicado: (2025)
Evaluation of Posits for Spectral Analysis Using a Software-Defined Dataflow Architecture
por: Deshmukh, Sameer, et al.
Publicado: (2024)
por: Deshmukh, Sameer, et al.
Publicado: (2024)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
por: Li, Fang
Publicado: (2025)
por: Li, Fang
Publicado: (2025)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
por: Gao, Yizhao, et al.
Publicado: (2024)
por: Gao, Yizhao, et al.
Publicado: (2024)
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
por: Zhou, Weiyu, et al.
Publicado: (2025)
por: Zhou, Weiyu, et al.
Publicado: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
por: Gupta, Neelesh, et al.
Publicado: (2026)
por: Gupta, Neelesh, et al.
Publicado: (2026)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
por: Umuroglu, Yaman, et al.
Publicado: (2025)
por: Umuroglu, Yaman, et al.
Publicado: (2025)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
por: Chen, Yuehai, et al.
Publicado: (2025)
por: Chen, Yuehai, et al.
Publicado: (2025)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
por: Yang, Kuilian, et al.
Publicado: (2026)
por: Yang, Kuilian, et al.
Publicado: (2026)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
por: Han, Meng, et al.
Publicado: (2023)
por: Han, Meng, et al.
Publicado: (2023)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
por: Yang, Simei, et al.
Publicado: (2025)
por: Yang, Simei, et al.
Publicado: (2025)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
por: Shi, Shangyi, et al.
Publicado: (2025)
por: Shi, Shangyi, et al.
Publicado: (2025)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
por: Wang, Jiayi, et al.
Publicado: (2026)
por: Wang, Jiayi, et al.
Publicado: (2026)
Ejemplares similares
-
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
por: Ko, Sho, et al.
Publicado: (2024) -
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
por: Sohn, Gina, et al.
Publicado: (2024) -
Revet: A Language and Compiler for Dataflow Threads
por: Rucker, Alexander, et al.
Publicado: (2023) -
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
por: Lacouture, Rubens, et al.
Publicado: (2025) -
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
por: Kim, Jungwoo, et al.
Publicado: (2026)