HeTraX: Energy Efficient 3D Heterogeneous Manycore Architecture for Transformer Acceleration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dhingra, Pratyush, Doppa, Janardhan Rao, Pande, Partha Pratim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025)
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads
von: Sharma, Harsh, et al.
Veröffentlicht: (2024)
von: Sharma, Harsh, et al.
Veröffentlicht: (2024)
FARe: Fault-Aware GNN Training on ReRAM-based PIM Accelerators
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024)
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024)
Designing High-Performance and Thermally Feasible Multi-Chiplet Architectures enabled by Non-bendable Glass Interposer
von: Sharma, Harsh, et al.
Veröffentlicht: (2025)
von: Sharma, Harsh, et al.
Veröffentlicht: (2025)
HePGA: A Heterogeneous Processing-in-Memory based GNN Training Accelerator
von: Ogbogu, Chukwufumnanya, et al.
Veröffentlicht: (2025)
von: Ogbogu, Chukwufumnanya, et al.
Veröffentlicht: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
MFIT: Multi-Fidelity Thermal Modeling for 2.5D and 3D Multi-Chiplet Architectures
von: Pfromm, Lukas, et al.
Veröffentlicht: (2024)
von: Pfromm, Lukas, et al.
Veröffentlicht: (2024)
CHIPSIM: A Co-Simulation Framework for Deep Learning on Chiplet-Based Systems
von: Pfromm, Lukas, et al.
Veröffentlicht: (2025)
von: Pfromm, Lukas, et al.
Veröffentlicht: (2025)
Look-Up Table based Neural Network Hardware
von: Sen, Ovishake, et al.
Veröffentlicht: (2024)
von: Sen, Ovishake, et al.
Veröffentlicht: (2024)
A 2.5-nA Area-Efficient Temperature-Independent 176-/82-ppm/°C CMOS-Only Current Reference in 0.11-$μ$m Bulk and 22-nm FD-SOI
von: Lefebvre, Martin, et al.
Veröffentlicht: (2024)
von: Lefebvre, Martin, et al.
Veröffentlicht: (2024)
A nA-Range Area-Efficient Sub-100-ppm/°C Peaking Current Reference Using Forward Body Biasing in 0.11-$μ$m Bulk and 22-nm FD-SOI
von: Lefebvre, Martin, et al.
Veröffentlicht: (2024)
von: Lefebvre, Martin, et al.
Veröffentlicht: (2024)
Hourglass Sorting: A novel parallel sorting algorithm and its implementation
von: Bascones, Daniel, et al.
Veröffentlicht: (2025)
von: Bascones, Daniel, et al.
Veröffentlicht: (2025)
MANTIS: A Mixed-Signal Near-Sensor Convolutional Imager SoC Using Charge-Domain 4b-Weighted 5-to-84-TOPS/W MAC Operations for Feature Extraction and Region-of-Interest Detection
von: Lefebvre, Martin, et al.
Veröffentlicht: (2024)
von: Lefebvre, Martin, et al.
Veröffentlicht: (2024)
A 1.1- / 0.9-nA Temperature-Independent 213- / 565-ppm/$^\circ$C Self-Biased CMOS-Only Current Reference in 65-nm Bulk and 22-nm FDSOI
von: Lefebvre, Martin, et al.
Veröffentlicht: (2023)
von: Lefebvre, Martin, et al.
Veröffentlicht: (2023)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
Re-thinking Memory-Bound Limitations in CGRAs
von: Liu, Xiangfeng, et al.
Veröffentlicht: (2025)
von: Liu, Xiangfeng, et al.
Veröffentlicht: (2025)
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
von: Riedel, Samuel, et al.
Veröffentlicht: (2024)
von: Riedel, Samuel, et al.
Veröffentlicht: (2024)
Data Gravity and the Energy Limits of Computation
von: Lee, Wonsuk, et al.
Veröffentlicht: (2026)
von: Lee, Wonsuk, et al.
Veröffentlicht: (2026)
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design
von: Qian, Chao, et al.
Veröffentlicht: (2026)
von: Qian, Chao, et al.
Veröffentlicht: (2026)
Towards Generalized On-Chip Communication for Programmable Accelerators in Heterogeneous Architectures
von: Zuckerman, Joseph, et al.
Veröffentlicht: (2024)
von: Zuckerman, Joseph, et al.
Veröffentlicht: (2024)
A Flexible Instruction Set Architecture for Efficient GEMMs
von: Santana, Alexandre de Limas, et al.
Veröffentlicht: (2025)
von: Santana, Alexandre de Limas, et al.
Veröffentlicht: (2025)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
von: Wu, Jiajun, et al.
Veröffentlicht: (2024)
von: Wu, Jiajun, et al.
Veröffentlicht: (2024)
An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Weight/Output Stationarity
von: Chauvaux, Nicolas, et al.
Veröffentlicht: (2024)
von: Chauvaux, Nicolas, et al.
Veröffentlicht: (2024)
Energy-Aware Heterogeneous Federated Learning via Approximate DNN Accelerators
von: Pfeiffer, Kilian, et al.
Veröffentlicht: (2024)
von: Pfeiffer, Kilian, et al.
Veröffentlicht: (2024)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026)
von: Özyılmaz, Mustafa Mert
Veröffentlicht: (2026)
A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
von: Huang, Wei-Hsing, et al.
Veröffentlicht: (2025)
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
von: Mitra, Subhasish, et al.
Veröffentlicht: (2025)
von: Mitra, Subhasish, et al.
Veröffentlicht: (2025)
basic_RV32s: An Open-Source Microarchitectural Roadmap for RISC-V RV32I
von: Kang, Hyun Woo, et al.
Veröffentlicht: (2025)
von: Kang, Hyun Woo, et al.
Veröffentlicht: (2025)
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
von: Hunhold, Laslo
Veröffentlicht: (2025)
von: Hunhold, Laslo
Veröffentlicht: (2025)
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
von: Xu, Han, et al.
Veröffentlicht: (2024)
von: Xu, Han, et al.
Veröffentlicht: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
Trilinear Compute-in-Memory Architecture for Energy-Efficient Transformer Acceleration
von: Mia, Md Zesun Ahmed, et al.
Veröffentlicht: (2026)
von: Mia, Md Zesun Ahmed, et al.
Veröffentlicht: (2026)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
Energy-Efficient Hardware Acceleration of Whisper ASR on a CGLA
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
von: Ando, Takuto, et al.
Veröffentlicht: (2025)
Resource Optimized Quantum Squaring Circuit
von: Sultana, Afrin, et al.
Veröffentlicht: (2024)
von: Sultana, Afrin, et al.
Veröffentlicht: (2024)
HeLEx: A Heterogeneous Layout Explorer for Spatial Elastic Coarse-Grained Reconfigurable Arrays
von: Du, Alan Jia Bao, et al.
Veröffentlicht: (2025)
von: Du, Alan Jia Bao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025) -
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023) -
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads
von: Sharma, Harsh, et al.
Veröffentlicht: (2024) -
FARe: Fault-Aware GNN Training on ReRAM-based PIM Accelerators
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024) -
Designing High-Performance and Thermally Feasible Multi-Chiplet Architectures enabled by Non-bendable Glass Interposer
von: Sharma, Harsh, et al.
Veröffentlicht: (2025)