Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kiyawat, Khyati, Fan, Zhenxing, Seneviratne, Yasas, Baradaran, Morteza, Shekar, Akhil, Xia, Zihan, Kang, Mingu, Skadron, Kevin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
von: Shekar, Akhil, et al.
Veröffentlicht: (2025)
Swift: A Multi-FPGA Framework for Scaling Up Accelerated Graph Analytics
von: Jaiyeoba, Oluwole, et al.
Veröffentlicht: (2024)
von: Jaiyeoba, Oluwole, et al.
Veröffentlicht: (2024)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
von: Gu, Yufeng, et al.
Veröffentlicht: (2025)
von: Gu, Yufeng, et al.
Veröffentlicht: (2025)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
PIMfused: Near-Bank DRAM-PIM with Fused-layer Dataflow for CNN Data Transfer Optimization
von: Yang, Simei, et al.
Veröffentlicht: (2025)
von: Yang, Simei, et al.
Veröffentlicht: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
Chiplet-Gym: Optimizing Chiplet-based AI Accelerator Design with Reinforcement Learning
von: Mishty, Kaniz, et al.
Veröffentlicht: (2024)
von: Mishty, Kaniz, et al.
Veröffentlicht: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
Toleo: Scaling Freshness to Tera-scale Memory using CXL and PIM
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2025)
von: Chong, Yue Jiet, et al.
Veröffentlicht: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
Mozart: A Chiplet Ecosystem-Accelerator Codesign Framework for Composable Bespoke Application Specific Integrated Circuits
von: Jin, Haoran, et al.
Veröffentlicht: (2025)
von: Jin, Haoran, et al.
Veröffentlicht: (2025)
The Survey of Chiplet-based Integrated Architecture: An EDA perspective
von: Chen, Shixin, et al.
Veröffentlicht: (2024)
von: Chen, Shixin, et al.
Veröffentlicht: (2024)
Monad: Towards Cost-effective Specialization for Chiplet-based Spatial Accelerators
von: Hao, Xiaochen, et al.
Veröffentlicht: (2023)
von: Hao, Xiaochen, et al.
Veröffentlicht: (2023)
AME-PIM: Can Memory be Your Next Tensor Accelerator?
von: Venieri, Emanuele, et al.
Veröffentlicht: (2026)
von: Venieri, Emanuele, et al.
Veröffentlicht: (2026)
Chiplets on Wheels: Review Paper on Holistic Chiplet Solutions for Autonomous Vehicles
von: Narashiman, Swathi, et al.
Veröffentlicht: (2024)
von: Narashiman, Swathi, et al.
Veröffentlicht: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
von: Shin, Yongwon, et al.
Veröffentlicht: (2024)
von: Shin, Yongwon, et al.
Veröffentlicht: (2024)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
von: Lin, Ye, et al.
Veröffentlicht: (2026)
von: Lin, Ye, et al.
Veröffentlicht: (2026)
Shifting in-DRAM
von: Tegge, William C., et al.
Veröffentlicht: (2026)
von: Tegge, William C., et al.
Veröffentlicht: (2026)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Chiplet Actuary: A Quantitative Cost Model and Multi-Chiplet Architecture Exploration
von: Feng, Yinxiao, et al.
Veröffentlicht: (2022)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2022)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
von: Iff, Patrick, et al.
Veröffentlicht: (2023)
von: Iff, Patrick, et al.
Veröffentlicht: (2023)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
von: Wu, Yuting, et al.
Veröffentlicht: (2023)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
An Analog and Digital Hybrid Attention Accelerator for Transformers with Charge-based In-memory Computing
von: Moradifirouzabadi, Ashkan, et al.
Veröffentlicht: (2024)
von: Moradifirouzabadi, Ashkan, et al.
Veröffentlicht: (2024)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
von: Kabir, MD Arafat, et al.
Veröffentlicht: (2024)
von: Kabir, MD Arafat, et al.
Veröffentlicht: (2024)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
von: Ma, Siyuan, et al.
Veröffentlicht: (2025)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
ProactivePIM: Accelerating Weight-Sharing Embedding Layer with PIM for Scalable Recommendation System
von: Kim, Youngsuk, et al.
Veröffentlicht: (2024)
von: Kim, Youngsuk, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Membrane: Accelerating Database Analytics with Bank-Level DRAM-PIM Filtering
von: Shekar, Akhil, et al.
Veröffentlicht: (2025) -
Swift: A Multi-FPGA Framework for Scaling Up Accelerated Graph Analytics
von: Jaiyeoba, Oluwole, et al.
Veröffentlicht: (2024) -
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
von: Hong, Junguk, et al.
Veröffentlicht: (2026) -
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
von: Gu, Yufeng, et al.
Veröffentlicht: (2025) -
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2024)