PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Ruokai, Li, Yuhang, Panda, Priyadarshini |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
von: Yin, Ruokai, et al.
Veröffentlicht: (2024)
von: Yin, Ruokai, et al.
Veröffentlicht: (2024)
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
von: Ghosh, Arkapravo, et al.
Veröffentlicht: (2025)
von: Ghosh, Arkapravo, et al.
Veröffentlicht: (2025)
Banked Memories for Soft SIMT Processors
von: Langhammer, Martin, et al.
Veröffentlicht: (2025)
von: Langhammer, Martin, et al.
Veröffentlicht: (2025)
A 950 MHz SIMT Soft Processor
von: Langhammer, Martin, et al.
Veröffentlicht: (2025)
von: Langhammer, Martin, et al.
Veröffentlicht: (2025)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
A Flexible Instruction Set Architecture for Efficient GEMMs
von: Santana, Alexandre de Limas, et al.
Veröffentlicht: (2025)
von: Santana, Alexandre de Limas, et al.
Veröffentlicht: (2025)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
von: Li, Fang
Veröffentlicht: (2025)
von: Li, Fang
Veröffentlicht: (2025)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
von: Zheng, Jianing, et al.
Veröffentlicht: (2025)
Benchmarking for Single Feature Attribution with Microarchitecture Cliffs
von: Zhen, Hao, et al.
Veröffentlicht: (2026)
von: Zhen, Hao, et al.
Veröffentlicht: (2026)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
von: Cheng, Jianyi, et al.
Veröffentlicht: (2023)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
von: He, Xiaolin, et al.
Veröffentlicht: (2025)
Revet: A Language and Compiler for Dataflow Threads
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
von: Rucker, Alexander, et al.
Veröffentlicht: (2023)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2026)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
von: Gao, Yizhao, et al.
Veröffentlicht: (2024)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
von: Gu, Hang, et al.
Veröffentlicht: (2026)
von: Gu, Hang, et al.
Veröffentlicht: (2026)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
Automatic Microarchitecture-Aware Custom Instruction Design for RISC-V Processors
von: Rezunov, Evgenii, et al.
Veröffentlicht: (2025)
von: Rezunov, Evgenii, et al.
Veröffentlicht: (2025)
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
von: Zhang, Jie, et al.
Veröffentlicht: (2026)
von: Zhang, Jie, et al.
Veröffentlicht: (2026)
Evaluating the Effectiveness of Microarchitectural Hardware Fault Detection for Application-Specific Requirements
von: Papadopoulos, Konstantinos-Nikolaos, et al.
Veröffentlicht: (2024)
von: Papadopoulos, Konstantinos-Nikolaos, et al.
Veröffentlicht: (2024)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
Stream-HLS: Towards Automatic Dataflow Acceleration
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
Fast Cross-Operator Optimization of Attention Dataflow
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
von: Chang, Haodong, et al.
Veröffentlicht: (2026)
SemanticBBV: A Semantic Signature for Cross-Program Knowledge Reuse in Microarchitecture Simulation
von: Liu, Zhenguo, et al.
Veröffentlicht: (2025)
von: Liu, Zhenguo, et al.
Veröffentlicht: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
von: Yi, Xiaoling, et al.
Veröffentlicht: (2025)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
von: Park, Haneul, et al.
Veröffentlicht: (2025)
von: Park, Haneul, et al.
Veröffentlicht: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python Tensors
von: Leitersdorf, Orian, et al.
Veröffentlicht: (2023)
von: Leitersdorf, Orian, et al.
Veröffentlicht: (2023)
DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness
von: Zhao, Xiaotian, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaotian, et al.
Veröffentlicht: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
von: Ye, Hanchen, et al.
Veröffentlicht: (2025)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
von: Sohn, Gina, et al.
Veröffentlicht: (2024)
Supporting Secured Integration of Microarchitectural Defenses
von: Ramkrishnan, Kartik, et al.
Veröffentlicht: (2026)
von: Ramkrishnan, Kartik, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
von: Yin, Ruokai, et al.
Veröffentlicht: (2024) -
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
von: Ghosh, Arkapravo, et al.
Veröffentlicht: (2025) -
Banked Memories for Soft SIMT Processors
von: Langhammer, Martin, et al.
Veröffentlicht: (2025) -
A 950 MHz SIMT Soft Processor
von: Langhammer, Martin, et al.
Veröffentlicht: (2025) -
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)