TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
Fuente:
arXiv
Saved in:
| Main Authors: | Moitra, Abhishek, Bhattacharjee, Abhiroop, Kim, Youngeun, Panda, Priyadarshini |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024)
by: Moitra, Abhishek, et al.
Published: (2024)
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
by: Ghosh, Arkapravo, et al.
Published: (2025)
by: Ghosh, Arkapravo, et al.
Published: (2025)
Low-power Spike-based Wearable Analytics on RRAM Crossbars
by: Bhattacharjee, Abhiroop, et al.
Published: (2025)
by: Bhattacharjee, Abhiroop, et al.
Published: (2025)
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
by: Yin, Ruokai, et al.
Published: (2024)
by: Yin, Ruokai, et al.
Published: (2024)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
by: Moitra, Abhishek, et al.
Published: (2025)
by: Moitra, Abhishek, et al.
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
by: He, Zifan, et al.
Published: (2025)
by: He, Zifan, et al.
Published: (2025)
BinSparX: Sparsified Binary Neural Networks for Reduced Hardware Non-Idealities in Xbar Arrays
by: Malhotra, Akul, et al.
Published: (2024)
by: Malhotra, Akul, et al.
Published: (2024)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
by: Yun, Sungmin, et al.
Published: (2025)
by: Yun, Sungmin, et al.
Published: (2025)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
by: Ji, Yuhao, et al.
Published: (2024)
by: Ji, Yuhao, et al.
Published: (2024)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)
by: Liang, Yanbiao, et al.
Published: (2024)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
by: You, Kang, et al.
Published: (2026)
by: You, Kang, et al.
Published: (2026)
Efficient Deployment of CNN Models on Multiple In-Memory Computing Units
by: Bougioukou, Eleni, et al.
Published: (2025)
by: Bougioukou, Eleni, et al.
Published: (2025)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
by: Zhu, Zhantong, et al.
Published: (2025)
by: Zhu, Zhantong, et al.
Published: (2025)
Full-stack evaluation of Machine Learning inference workloads for RISC-V systems
by: Bhattacharjee, Debjyoti, et al.
Published: (2024)
by: Bhattacharjee, Debjyoti, et al.
Published: (2024)
ApproXAI: Energy-Efficient Hardware Acceleration of Explainable AI using Approximate Computing
by: Siddique, Ayesha, et al.
Published: (2025)
by: Siddique, Ayesha, et al.
Published: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
by: Xia, Tianhua, et al.
Published: (2025)
by: Xia, Tianhua, et al.
Published: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM
by: Cha, SangHoon, et al.
Published: (2026)
by: Cha, SangHoon, et al.
Published: (2026)
MICSim: A Modular Simulator for Mixed-signal Compute-in-Memory based AI Accelerator
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
StoX-Net: Stochastic Processing of Partial Sums for Efficient In-Memory Computing DNN Accelerators
by: Rogers, Ethan G, et al.
Published: (2024)
by: Rogers, Ethan G, et al.
Published: (2024)
Characterizing Soft-Error Resiliency in Arm's Ethos-U55 Embedded Machine Learning Accelerator
by: Tyagi, Abhishek, et al.
Published: (2024)
by: Tyagi, Abhishek, et al.
Published: (2024)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
by: Ma, Haiyue, et al.
Published: (2024)
by: Ma, Haiyue, et al.
Published: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
Efficient Calibration for RRAM-based In-Memory Computing using DoRA
by: Dong, Weirong, et al.
Published: (2025)
by: Dong, Weirong, et al.
Published: (2025)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
by: Lee, Minjae, et al.
Published: (2023)
by: Lee, Minjae, et al.
Published: (2023)
Towards Forever Access for Implanted Brain-Computer Interfaces
by: Ugur, Muhammed, et al.
Published: (2024)
by: Ugur, Muhammed, et al.
Published: (2024)
Static IR Drop Prediction with Attention U-Net and Saliency-Based Explainability
by: Zhang, Lizi, et al.
Published: (2024)
by: Zhang, Lizi, et al.
Published: (2024)
Neuromorphic Computing for Low-Power Artificial Intelligence
by: Katti, Keshava, et al.
Published: (2026)
by: Katti, Keshava, et al.
Published: (2026)
Computing-In-Memory Dataflow for Minimal Buffer Traffic
by: Song, Choongseok, et al.
Published: (2025)
by: Song, Choongseok, et al.
Published: (2025)
MAGNet: A Multi-Scale Attention-Guided Graph Fusion Network for DRC Violation Detection
by: Lu, Weihan, et al.
Published: (2025)
by: Lu, Weihan, et al.
Published: (2025)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
by: Wan, Zishen, et al.
Published: (2024)
by: Wan, Zishen, et al.
Published: (2024)
HyDRA: Deadline and Reuse-Aware Cacheability for Hardware Accelerators
by: Agarwal, Ayushi, et al.
Published: (2026)
by: Agarwal, Ayushi, et al.
Published: (2026)
An ultra-low-power CGRA for accelerating Transformers at the edge
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
Similar Items
-
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
by: Moitra, Abhishek, et al.
Published: (2024) -
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
by: Moitra, Abhishek, et al.
Published: (2024) -
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
by: Ghosh, Arkapravo, et al.
Published: (2025) -
Low-power Spike-based Wearable Analytics on RRAM Crossbars
by: Bhattacharjee, Abhiroop, et al.
Published: (2025) -
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
by: Yin, Ruokai, et al.
Published: (2024)