MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kubo, Tatsuya, Tokuda, Daichi, Nagatani, Tomoya, Usui, Masayuki, Qu, Lei, Cao, Ting, Takamaeda-Yamazaki, Shinya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
von: Ohno, Ayumi, et al.
Veröffentlicht: (2025)
von: Ohno, Ayumi, et al.
Veröffentlicht: (2025)
Exploring the Versal AI Engine for 3D Gaussian Splatting
von: Shimamura, Kotaro, et al.
Veröffentlicht: (2025)
von: Shimamura, Kotaro, et al.
Veröffentlicht: (2025)
Relational Hoare Logic for High-Level Synthesis of Hardware Accelerators
von: Tanaka, Izumi, et al.
Veröffentlicht: (2026)
von: Tanaka, Izumi, et al.
Veröffentlicht: (2026)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
von: Jeong, Minki, et al.
Veröffentlicht: (2022)
von: Jeong, Minki, et al.
Veröffentlicht: (2022)
Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
In-DRAM True Random Number Generation Using Simultaneous Multiple-Row Activation: An Experimental Study of Real DRAM Chips
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2025)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2025)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2024)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
Taking Cryptography Out of the Data Path via Near-Memory Processing in DRAM
von: Barcarolo, Nicola, et al.
Veröffentlicht: (2026)
von: Barcarolo, Nicola, et al.
Veröffentlicht: (2026)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
Achieving Dependability of AI Execution with Radiation Hardened Processors
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)
von: Taquichiri, Carlos Rafael Tordoya, et al.
Veröffentlicht: (2025)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)
Accelerated Execution of Bayesian Neural Networks using a Single Probabilistic Forward Pass and Code Generation
von: Klein, Bernhard, et al.
Veröffentlicht: (2025)
von: Klein, Bernhard, et al.
Veröffentlicht: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Enabling Accelerators for Graph Computing
von: Shivdikar, Kaustubh
Veröffentlicht: (2023)
von: Shivdikar, Kaustubh
Veröffentlicht: (2023)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
Leveraging SIMD for Accelerating Large-number Arithmetic
von: Das, Subhrajit, et al.
Veröffentlicht: (2026)
von: Das, Subhrajit, et al.
Veröffentlicht: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
von: Hsia, Samuel, et al.
Veröffentlicht: (2023)
von: Hsia, Samuel, et al.
Veröffentlicht: (2023)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
Enabling Mixed criticality applications for the Versal AI-Engines
von: Sprave, Vincent, et al.
Veröffentlicht: (2026)
von: Sprave, Vincent, et al.
Veröffentlicht: (2026)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025) -
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025) -
Accelerating Elliptic Curve Point Additions on Versal AI Engine for Multi-scalar Multiplication
von: Ohno, Ayumi, et al.
Veröffentlicht: (2025) -
Exploring the Versal AI Engine for 3D Gaussian Splatting
von: Shimamura, Kotaro, et al.
Veröffentlicht: (2025) -
Relational Hoare Logic for High-Level Synthesis of Hardware Accelerators
von: Tanaka, Izumi, et al.
Veröffentlicht: (2026)