Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Sohn, Gina, Zhang, Genghan, Hossfeld, Konstantin, Kim, Jungwoo, Sobotka, Nathan, Zhang, Nathan, Hsu, Olivia, Olukotun, Kunle |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
by: Sohn, Gina, et al.
Published: (2024)
by: Sohn, Gina, et al.
Published: (2024)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
by: Lacouture, Rubens, et al.
Published: (2025)
by: Lacouture, Rubens, et al.
Published: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
by: Kim, Jungwoo, et al.
Published: (2026)
by: Kim, Jungwoo, et al.
Published: (2026)
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
by: Ko, Sho, et al.
Published: (2024)
by: Ko, Sho, et al.
Published: (2024)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)
by: Zhou, Zikai, et al.
Published: (2025)
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models
by: Ko, Sho, et al.
Published: (2025)
by: Ko, Sho, et al.
Published: (2025)
Allo: A Programming Model for Composable Accelerator Design
by: Chen, Hongzheng, et al.
Published: (2024)
by: Chen, Hongzheng, et al.
Published: (2024)
LLM-Aided Compilation for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2024)
by: Hong, Charles, et al.
Published: (2024)
Dato: A Task-Based Programming Model for Dataflow Accelerators
by: Fang, Shihan, et al.
Published: (2025)
by: Fang, Shihan, et al.
Published: (2025)
Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
by: Soi, Rupanshu, et al.
Published: (2025)
by: Soi, Rupanshu, et al.
Published: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
by: Ye, Hanchen, et al.
Published: (2025)
by: Ye, Hanchen, et al.
Published: (2025)
WaveCert: Translation Validation for Asynchronous Dataflow Programs via Dynamic Fractional Permissions
by: Lin, Zhengyao, et al.
Published: (2023)
by: Lin, Zhengyao, et al.
Published: (2023)
ACT: Automatically Generating Compiler Backends from Tensor Accelerator ISA Descriptions
by: Jain, Devansh, et al.
Published: (2025)
by: Jain, Devansh, et al.
Published: (2025)
VerilogMonkey: Exploring Parallel Scaling for Automated Verilog Code Generation with LLMs
by: Niu, Juxin, et al.
Published: (2025)
by: Niu, Juxin, et al.
Published: (2025)
Partial Cross-Compilation and Mixed Execution for Accelerating Dynamic Binary Translation
by: Gu, Yuhao, et al.
Published: (2025)
by: Gu, Yuhao, et al.
Published: (2025)
QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression for Circuit Design
by: Huang, Lei, et al.
Published: (2025)
by: Huang, Lei, et al.
Published: (2025)
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
by: Zhu, Yaoyu, et al.
Published: (2025)
by: Zhu, Yaoyu, et al.
Published: (2025)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
Scaling Program Synthesis Based Technology Mapping with Equality Saturation
by: Smith, Gus Henry, et al.
Published: (2024)
by: Smith, Gus Henry, et al.
Published: (2024)
FPGA Technology Mapping Using Sketch-Guided Program Synthesis
by: Smith, Gus Henry, et al.
Published: (2024)
by: Smith, Gus Henry, et al.
Published: (2024)
Tawa: Automatic Warp Specialization for Modern GPUs with Asynchronous References
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
The Program Hypergraph: Multi-Way Relational Structure for Geometric Algebra, Spatial Compute, and Physics-Aware Compilation
by: Haynes, Houston
Published: (2026)
by: Haynes, Houston
Published: (2026)
RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution
by: Liu, Shang, et al.
Published: (2023)
by: Liu, Shang, et al.
Published: (2023)
Quantum Register Machine: Efficient Implementation of Quantum Recursive Programs
by: Zhang, Zhicheng, et al.
Published: (2024)
by: Zhang, Zhicheng, et al.
Published: (2024)
Revet: A Language and Compiler for Dataflow Threads
by: Rucker, Alexander, et al.
Published: (2023)
by: Rucker, Alexander, et al.
Published: (2023)
SpecLoop: An Agentic RTL-to-Specification Framework with Formal Verification Feedback Loop
by: Chang, Fu-Chieh, et al.
Published: (2026)
by: Chang, Fu-Chieh, et al.
Published: (2026)
RTL++: Graph-enhanced LLM for RTL Code Generation
by: Akyash, Mohammad, et al.
Published: (2025)
by: Akyash, Mohammad, et al.
Published: (2025)
DecoRTL: A Run-time Decoding Framework for RTL Code Generation with LLMs
by: Akyash, Mohammad, et al.
Published: (2025)
by: Akyash, Mohammad, et al.
Published: (2025)
Leveraging High-Level Synthesis and Large Language Models to Generate, Simulate, and Deploy a Uniform Random Number Generator Hardware Design
by: Meech, James T.
Published: (2023)
by: Meech, James T.
Published: (2023)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024)
by: Zhang, Weichuang, et al.
Published: (2024)
E-Syn: E-Graph Rewriting with Technology-Aware Cost Functions for Logic Synthesis
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
VRank: Enhancing Verilog Code Generation from Large Language Models via Self-Consistency
by: Zhao, Zhuorui, et al.
Published: (2025)
by: Zhao, Zhuorui, et al.
Published: (2025)
HaVen: Hallucination-Mitigated LLM for Verilog Code Generation Aligned with HDL Engineers
by: Yang, Yiyao, et al.
Published: (2025)
by: Yang, Yiyao, et al.
Published: (2025)
NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures
by: Li, Shangkun, et al.
Published: (2026)
by: Li, Shangkun, et al.
Published: (2026)
A2H-MAS: An Algorithm-to-HLS Multi-Agent System for Automated and Reliable FPGA Implementation
by: Lei, Jie, et al.
Published: (2025)
by: Lei, Jie, et al.
Published: (2025)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
by: Tian, Jiayi, et al.
Published: (2025)
by: Tian, Jiayi, et al.
Published: (2025)
Reconfigurable Stream Network Architecture
by: Wang, Chengyue, et al.
Published: (2024)
by: Wang, Chengyue, et al.
Published: (2024)
Zoozve: A Strip-Mining-Free RISC-V Vector Extension with Arbitrary Register Grouping Compilation Support (WIP)
by: Xu, Siyi, et al.
Published: (2025)
by: Xu, Siyi, et al.
Published: (2025)
Register Aggregation for Hardware Decompilation
by: Rao, Varun, et al.
Published: (2024)
by: Rao, Varun, et al.
Published: (2024)
Similar Items
-
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
by: Sohn, Gina, et al.
Published: (2024) -
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
by: Lacouture, Rubens, et al.
Published: (2025) -
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
by: Kim, Jungwoo, et al.
Published: (2026) -
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings
by: Ko, Sho, et al.
Published: (2024) -
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)