LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Junguk, Shin, Changmin, Kim, Sukjin, Noh, Si Ung, Kwon, Taehee, Park, Seongyeon, Kim, Hanjun, Kim, Youngsok, Lee, Jinho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
von: Noh, Si Ung, et al.
Veröffentlicht: (2024)
von: Noh, Si Ung, et al.
Veröffentlicht: (2024)
A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
von: Jang, Hongsun, et al.
Veröffentlicht: (2025)
von: Jang, Hongsun, et al.
Veröffentlicht: (2025)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024)
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
von: Han, Wontak, et al.
Veröffentlicht: (2024)
von: Han, Wontak, et al.
Veröffentlicht: (2024)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
von: Shin, Changmin, et al.
Veröffentlicht: (2025)
von: Shin, Changmin, et al.
Veröffentlicht: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
von: Jang, Hongsun, et al.
Veröffentlicht: (2024)
von: Jang, Hongsun, et al.
Veröffentlicht: (2024)
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
von: Kiyawat, Khyati, et al.
Veröffentlicht: (2025)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
Towards Efficient LUT-based PIM: A Scalable and Low-Power Approach for Modern Workloads
von: Khabbazan, Bahareh, et al.
Veröffentlicht: (2025)
von: Khabbazan, Bahareh, et al.
Veröffentlicht: (2025)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
von: Kim, Sukjin, et al.
Veröffentlicht: (2025)
von: Kim, Sukjin, et al.
Veröffentlicht: (2025)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
von: Yang, Simei, et al.
Veröffentlicht: (2025)
von: Yang, Simei, et al.
Veröffentlicht: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
Securing DRAM at Scale: ARFM-Driven Row Hammer Defense with Unveiling the Threat of Short tRC Patterns
von: Joo, Nogeun, et al.
Veröffentlicht: (2025)
von: Joo, Nogeun, et al.
Veröffentlicht: (2025)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
von: Kwon, Omin, et al.
Veröffentlicht: (2025)
von: Kwon, Omin, et al.
Veröffentlicht: (2025)
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands
von: Nam, Hwayong, et al.
Veröffentlicht: (2024)
von: Nam, Hwayong, et al.
Veröffentlicht: (2024)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
von: Kim, Youngsuk, et al.
Veröffentlicht: (2024)
von: Kim, Youngsuk, et al.
Veröffentlicht: (2024)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
von: Shin, Yongwon, et al.
Veröffentlicht: (2024)
von: Shin, Yongwon, et al.
Veröffentlicht: (2024)
Low-Power Encoding for PAM-3 DRAM Bus
von: Nam, Jonghyeon, et al.
Veröffentlicht: (2024)
von: Nam, Jonghyeon, et al.
Veröffentlicht: (2024)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
von: Andronic, Marta, et al.
Veröffentlicht: (2023)
Shifting in-DRAM
von: Tegge, William C., et al.
Veröffentlicht: (2026)
von: Tegge, William C., et al.
Veröffentlicht: (2026)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
von: Huang, Zhirui, et al.
Veröffentlicht: (2025)
von: Huang, Zhirui, et al.
Veröffentlicht: (2025)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
Flexible In-NAND Cryptographic Processing for Secure Flash Storage
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM
von: Cha, SangHoon, et al.
Veröffentlicht: (2026)
von: Cha, SangHoon, et al.
Veröffentlicht: (2026)
Annotated PIM Bibliography
von: Kogge, Peter M.
Veröffentlicht: (2026)
von: Kogge, Peter M.
Veröffentlicht: (2026)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
von: Lou, Binglei, et al.
Veröffentlicht: (2024)
von: Lou, Binglei, et al.
Veröffentlicht: (2024)
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models
von: Yang, Jinho, et al.
Veröffentlicht: (2025)
von: Yang, Jinho, et al.
Veröffentlicht: (2025)
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
von: Shin, Hery, et al.
Veröffentlicht: (2024)
von: Shin, Hery, et al.
Veröffentlicht: (2024)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
von: Moon, Seungjae, et al.
Veröffentlicht: (2024)
von: Moon, Seungjae, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
von: Noh, Si Ung, et al.
Veröffentlicht: (2024) -
A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
von: Jang, Hongsun, et al.
Veröffentlicht: (2025) -
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
von: Kwon, Hyucksung, et al.
Veröffentlicht: (2024) -
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025) -
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
von: Han, Wontak, et al.
Veröffentlicht: (2024)