Guardado en:
| Autores principales: | Khabbazan, Bahareh, Riera, Marc, González, Antonio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.02142 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
por: Hong, Junguk, et al.
Publicado: (2026)
por: Hong, Junguk, et al.
Publicado: (2026)
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
por: Sabri, Mohammad, et al.
Publicado: (2024)
por: Sabri, Mohammad, et al.
Publicado: (2024)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
por: Han, Wontak, et al.
Publicado: (2024)
por: Han, Wontak, et al.
Publicado: (2024)
Hamun: An Approximate Computation Method to Prolong the Lifespan of ReRAM-Based Accelerators
por: Sabri, Mohammad, et al.
Publicado: (2025)
por: Sabri, Mohammad, et al.
Publicado: (2025)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
por: Lee, Dongjae, et al.
Publicado: (2025)
por: Lee, Dongjae, et al.
Publicado: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
por: Andronic, Marta, et al.
Publicado: (2023)
por: Andronic, Marta, et al.
Publicado: (2023)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
por: Lin, Ye, et al.
Publicado: (2026)
por: Lin, Ye, et al.
Publicado: (2026)
CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
por: Ahn, Bas, et al.
Publicado: (2026)
por: Ahn, Bas, et al.
Publicado: (2026)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
por: Kanani, Alish, et al.
Publicado: (2025)
por: Kanani, Alish, et al.
Publicado: (2025)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
por: Kim, Youngsuk, et al.
Publicado: (2024)
por: Kim, Youngsuk, et al.
Publicado: (2024)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
por: Duan, Cenlin, et al.
Publicado: (2024)
por: Duan, Cenlin, et al.
Publicado: (2024)
A Survey on LUT-based Deep Neural Networks Implemented in FPGAs
por: Guo, Zeyu
Publicado: (2025)
por: Guo, Zeyu
Publicado: (2025)
Power-Area Efficient Serial IMPLY-based 4:2 Compressor Applied in Data-Intensive Applications
por: Bagheralmoosavi, Bahareh, et al.
Publicado: (2024)
por: Bagheralmoosavi, Bahareh, et al.
Publicado: (2024)
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
por: Kabir, MD Arafat, et al.
Publicado: (2024)
por: Kabir, MD Arafat, et al.
Publicado: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
por: Hyun, Bongjoon, et al.
Publicado: (2023)
por: Hyun, Bongjoon, et al.
Publicado: (2023)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
por: Popovici, Doru Thom, et al.
Publicado: (2025)
por: Popovici, Doru Thom, et al.
Publicado: (2025)
Annotated PIM Bibliography
por: Kogge, Peter M.
Publicado: (2026)
por: Kogge, Peter M.
Publicado: (2026)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
por: Lee, Dongjae, et al.
Publicado: (2024)
por: Lee, Dongjae, et al.
Publicado: (2024)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
por: Jeon, Sangmin, et al.
Publicado: (2025)
por: Jeon, Sangmin, et al.
Publicado: (2025)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
por: Li, Xingyang, et al.
Publicado: (2025)
por: Li, Xingyang, et al.
Publicado: (2025)
Dataflow-Aware PIM-Enabled Manycore Architecture for Deep Learning Workloads
por: Sharma, Harsh, et al.
Publicado: (2024)
por: Sharma, Harsh, et al.
Publicado: (2024)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
por: Alsop, Johnathan, et al.
Publicado: (2023)
por: Alsop, Johnathan, et al.
Publicado: (2023)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
por: Huang, Zhirui, et al.
Publicado: (2025)
por: Huang, Zhirui, et al.
Publicado: (2025)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
por: Lou, Binglei, et al.
Publicado: (2024)
por: Lou, Binglei, et al.
Publicado: (2024)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
por: Wang, Yimin, et al.
Publicado: (2025)
por: Wang, Yimin, et al.
Publicado: (2025)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
por: He, Zifan, et al.
Publicado: (2025)
por: He, Zifan, et al.
Publicado: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
por: Seo, Minseok, et al.
Publicado: (2024)
por: Seo, Minseok, et al.
Publicado: (2024)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
por: Vahdatniya, Parmida, et al.
Publicado: (2025)
por: Vahdatniya, Parmida, et al.
Publicado: (2025)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
por: Duan, Cenlin, et al.
Publicado: (2025)
por: Duan, Cenlin, et al.
Publicado: (2025)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
por: Malekar, Jinendra, et al.
Publicado: (2025)
por: Malekar, Jinendra, et al.
Publicado: (2025)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
por: Li, Guoyu, et al.
Publicado: (2025)
por: Li, Guoyu, et al.
Publicado: (2025)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
por: Hayes, Oran, et al.
Publicado: (2026)
por: Hayes, Oran, et al.
Publicado: (2026)
DSLUT: An Asymmetric LUT and its Automatic Design Flow Based on Practical Functions
por: Yang, Moucheng, et al.
Publicado: (2025)
por: Yang, Moucheng, et al.
Publicado: (2025)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
por: Pun, Junius, et al.
Publicado: (2025)
por: Pun, Junius, et al.
Publicado: (2025)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
por: Wu, Yuting, et al.
Publicado: (2023)
por: Wu, Yuting, et al.
Publicado: (2023)
WaSP: Warp Scheduling to Mimic Prefetching in Graphics Workloads
por: Joseph, Diya, et al.
Publicado: (2024)
por: Joseph, Diya, et al.
Publicado: (2024)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
por: Purayil, Navaneeth Kunhi, et al.
Publicado: (2025)
Control Flow Management in Modern GPUs
por: Shoushtary, Mojtaba Abaie, et al.
Publicado: (2024)
por: Shoushtary, Mojtaba Abaie, et al.
Publicado: (2024)
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference
por: Sun, Chang, et al.
Publicado: (2026)
por: Sun, Chang, et al.
Publicado: (2026)
Ejemplares similares
-
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
por: Hong, Junguk, et al.
Publicado: (2026) -
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
por: Sabri, Mohammad, et al.
Publicado: (2024) -
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
por: Han, Wontak, et al.
Publicado: (2024) -
Hamun: An Approximate Computation Method to Prolong the Lifespan of ReRAM-Based Accelerators
por: Sabri, Mohammad, et al.
Publicado: (2025) -
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
por: Lee, Dongjae, et al.
Publicado: (2025)