StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Hanchen, Chen, Deming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stream-HLS: Towards Automatic Dataflow Acceleration
by: Basalama, Suhail, et al.
Published: (2025)
by: Basalama, Suhail, et al.
Published: (2025)
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
by: Qin, Shantian, et al.
Published: (2025)
by: Qin, Shantian, et al.
Published: (2025)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022)
by: Symons, Arne, et al.
Published: (2022)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
by: Yang, Kuilian, et al.
Published: (2026)
by: Yang, Kuilian, et al.
Published: (2026)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
by: Sohn, Gina, et al.
Published: (2024)
by: Sohn, Gina, et al.
Published: (2024)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
by: Sohn, Gina, et al.
Published: (2025)
by: Sohn, Gina, et al.
Published: (2025)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
by: Li, Fang
Published: (2025)
by: Li, Fang
Published: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
by: Li, Jonathan, et al.
Published: (2025)
by: Li, Jonathan, et al.
Published: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
by: Huang, Sixiao, et al.
Published: (2025)
by: Huang, Sixiao, et al.
Published: (2025)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
Open-source Stand-Alone Versatile Tensor Accelerator
by: Faure-Gignoux, Anthony, et al.
Published: (2025)
by: Faure-Gignoux, Anthony, et al.
Published: (2025)
ATLAAS: Automatic Tensor-Level Abstraction of Accelerator Semantics
by: Gao, Ruijie, et al.
Published: (2026)
by: Gao, Ruijie, et al.
Published: (2026)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
by: Jia, Shuao, et al.
Published: (2025)
by: Jia, Shuao, et al.
Published: (2025)
Accelerating Recommender Model ETL with a Streaming FPGA-GPU Dataflow
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
by: Zhao, Zhiyuan, et al.
Published: (2024)
by: Zhao, Zhiyuan, et al.
Published: (2024)
Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacity
by: Xue, Zi Yu, et al.
Published: (2023)
by: Xue, Zi Yu, et al.
Published: (2023)
AME-PIM: Can Memory be Your Next Tensor Accelerator?
by: Venieri, Emanuele, et al.
Published: (2026)
by: Venieri, Emanuele, et al.
Published: (2026)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
TeAAL: A Declarative Framework for Modeling Sparse Tensor Accelerators
by: Nayak, Nandeeka, et al.
Published: (2023)
by: Nayak, Nandeeka, et al.
Published: (2023)
RTeAAL Sim: Using Tensor Algebra to Represent and Accelerate RTL Simulation (Extended Version)
by: Zhu, Yan, et al.
Published: (2026)
by: Zhu, Yan, et al.
Published: (2026)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
by: Xue, Chenhao, et al.
Published: (2026)
by: Xue, Chenhao, et al.
Published: (2026)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
by: Gilbert, Michael, et al.
Published: (2024)
by: Gilbert, Michael, et al.
Published: (2024)
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
GTA: a new General Tensor Accelerator with Better Area Efficiency and Data Reuse
by: Ai, Chenyang, et al.
Published: (2024)
by: Ai, Chenyang, et al.
Published: (2024)
Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
by: Zhou, Weiyu, et al.
Published: (2025)
by: Zhou, Weiyu, et al.
Published: (2025)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
by: He, Xiaolin, et al.
Published: (2025)
by: He, Xiaolin, et al.
Published: (2025)
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
by: Umuroglu, Yaman, et al.
Published: (2025)
by: Umuroglu, Yaman, et al.
Published: (2025)
StreamGrid: Streaming Point Cloud Analytics via Compulsory Splitting and Deterministic Termination
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
by: Kulp, Gabriel, et al.
Published: (2024)
by: Kulp, Gabriel, et al.
Published: (2024)
Demystifying the 7-D Convolution Loop Nest for Data and Instruction Streaming in Reconfigurable AI Accelerators
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2025)
Reconfigurable Stream Network Architecture
by: Wang, Chengyue, et al.
Published: (2024)
by: Wang, Chengyue, et al.
Published: (2024)
Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
by: Tong, Jianming, et al.
Published: (2024)
by: Tong, Jianming, et al.
Published: (2024)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
by: Manoni, Simone, et al.
Published: (2025)
by: Manoni, Simone, et al.
Published: (2025)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
by: Wu, Ka Wai
Published: (2024)
by: Wu, Ka Wai
Published: (2024)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
Similar Items
-
Stream-HLS: Towards Automatic Dataflow Acceleration
by: Basalama, Suhail, et al.
Published: (2025) -
StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal Transformer
by: Qin, Shantian, et al.
Published: (2025) -
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
by: Symons, Arne, et al.
Published: (2022) -
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
by: Yang, Kuilian, et al.
Published: (2026) -
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
by: Sohn, Gina, et al.
Published: (2024)