MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Moitra, Abhishek, Ghosh, Arkapravo, Agarwal, Shrey, Amarnath, Aporva, Swaminathan, Karthik, Panda, Priyadarshini |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
by: Ghosh, Arkapravo, et al.
Published: (2025)
by: Ghosh, Arkapravo, et al.
Published: (2025)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
by: Yin, Ruokai, et al.
Published: (2024)
by: Yin, Ruokai, et al.
Published: (2024)
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
Computing-In-Memory Dataflow for Minimal Buffer Traffic
by: Song, Choongseok, et al.
Published: (2025)
by: Song, Choongseok, et al.
Published: (2025)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
by: Chen, Chun-Ting, et al.
Published: (2025)
by: Chen, Chun-Ting, et al.
Published: (2025)
Low-power Spike-based Wearable Analytics on RRAM Crossbars
by: Bhattacharjee, Abhiroop, et al.
Published: (2025)
by: Bhattacharjee, Abhiroop, et al.
Published: (2025)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMs
by: Ye, Hanchen, et al.
Published: (2025)
by: Ye, Hanchen, et al.
Published: (2025)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
by: Xue, Chenhao, et al.
Published: (2026)
by: Xue, Chenhao, et al.
Published: (2026)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
by: Tong, Jianming, et al.
Published: (2024)
by: Tong, Jianming, et al.
Published: (2024)
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
by: He, Xiaolin, et al.
Published: (2025)
by: He, Xiaolin, et al.
Published: (2025)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
by: Hsiung, Ching-Lin, et al.
Published: (2025)
by: Hsiung, Ching-Lin, et al.
Published: (2025)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
by: Min, Kyeongpil, et al.
Published: (2026)
by: Min, Kyeongpil, et al.
Published: (2026)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
by: Shi, Shangyi, et al.
Published: (2025)
by: Shi, Shangyi, et al.
Published: (2025)
TrIM, Triangular Input Movement Systolic Array for Convolutional Neural Networks: Dataflow and Analytical Modelling
by: Sestito, Cristian, et al.
Published: (2024)
by: Sestito, Cristian, et al.
Published: (2024)
Neuromorphic Computing for Low-Power Artificial Intelligence
by: Katti, Keshava, et al.
Published: (2026)
by: Katti, Keshava, et al.
Published: (2026)
FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
by: Raha, Arnab, et al.
Published: (2024)
by: Raha, Arnab, et al.
Published: (2024)
NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention Dataflow
by: Chen, Yuehai, et al.
Published: (2025)
by: Chen, Yuehai, et al.
Published: (2025)
HyDRA: Deadline and Reuse-Aware Cacheability for Hardware Accelerators
by: Agarwal, Ayushi, et al.
Published: (2026)
by: Agarwal, Ayushi, et al.
Published: (2026)
A Compact, Low Power Transprecision ALU for Smart Edge Devices
by: Dube, Ayushi, et al.
Published: (2025)
by: Dube, Ayushi, et al.
Published: (2025)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
by: Zhang, Weichuang, et al.
Published: (2026)
by: Zhang, Weichuang, et al.
Published: (2026)
Revet: A Language and Compiler for Dataflow Threads
by: Rucker, Alexander, et al.
Published: (2023)
by: Rucker, Alexander, et al.
Published: (2023)
Fast Cross-Operator Optimization of Attention Dataflow
by: Chang, Haodong, et al.
Published: (2026)
by: Chang, Haodong, et al.
Published: (2026)
Stream-HLS: Towards Automatic Dataflow Acceleration
by: Basalama, Suhail, et al.
Published: (2025)
by: Basalama, Suhail, et al.
Published: (2025)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
by: Jeon, Sangmin, et al.
Published: (2025)
by: Jeon, Sangmin, et al.
Published: (2025)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
by: Yang, Simei, et al.
Published: (2025)
by: Yang, Simei, et al.
Published: (2025)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
DORA: Dataflow-Instruction Orchestration Architecture for DNN Acceleration
by: Chen, Xingzhen, et al.
Published: (2026)
by: Chen, Xingzhen, et al.
Published: (2026)
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads
by: Sharma, Harsh, et al.
Published: (2024)
by: Sharma, Harsh, et al.
Published: (2024)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
by: Sohn, Gina, et al.
Published: (2024)
by: Sohn, Gina, et al.
Published: (2024)
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
by: Negi, Shubham, et al.
Published: (2025)
by: Negi, Shubham, et al.
Published: (2025)
Hermes: A Unified High-Performance NTT Architecture with Hybrid Dataflow
by: Gu, Hang, et al.
Published: (2026)
by: Gu, Hang, et al.
Published: (2026)
LoopTree: Exploring the Fused-layer Dataflow Accelerator Design Space
by: Gilbert, Michael, et al.
Published: (2024)
by: Gilbert, Michael, et al.
Published: (2024)
Similar Items
-
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
by: Ghosh, Arkapravo, et al.
Published: (2025) -
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024) -
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
by: Moitra, Abhishek, et al.
Published: (2024) -
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025) -
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
by: Yin, Ruokai, et al.
Published: (2024)