PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Ortega, Cristobal, Falevoz, Yann, Ayrignac, Renaud |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Managed-Retention Memory: A New Class of Memory for the AI Era
by: Legtchenko, Sergey, et al.
Published: (2025)
by: Legtchenko, Sergey, et al.
Published: (2025)
WaferLLM: Large Language Model Inference at Wafer Scale
by: He, Congjie, et al.
Published: (2025)
by: He, Congjie, et al.
Published: (2025)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
Transforming the Hybrid Cloud for Emerging AI Workloads
by: Chen, Deming, et al.
Published: (2024)
by: Chen, Deming, et al.
Published: (2024)
Experience Deploying Containerized GenAI Services at an HPC Center
by: Beltre, Angel M., et al.
Published: (2025)
by: Beltre, Angel M., et al.
Published: (2025)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
by: Call, Aaron, et al.
Published: (2025)
by: Call, Aaron, et al.
Published: (2025)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)
by: Puente, Valentin, et al.
Published: (2016)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
Reference Architecture of a Quantum-Centric Supercomputer
by: Seelam, Seetharami, et al.
Published: (2026)
by: Seelam, Seetharami, et al.
Published: (2026)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
by: Hou, Yuewen, et al.
Published: (2025)
by: Hou, Yuewen, et al.
Published: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
by: Filippov, Dmitry, et al.
Published: (2025)
by: Filippov, Dmitry, et al.
Published: (2025)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
by: Roy, Rohan Basu, et al.
Published: (2025)
by: Roy, Rohan Basu, et al.
Published: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
by: Lee, Benjamin C., et al.
Published: (2024)
by: Lee, Benjamin C., et al.
Published: (2024)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
by: Garg, Chirag, et al.
Published: (2025)
by: Garg, Chirag, et al.
Published: (2025)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024)
by: Qin, Ruoyu, et al.
Published: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Improving AI Efficiency in Data Centres by Power Dynamic Response
by: Marinoni, Andrea, et al.
Published: (2025)
by: Marinoni, Andrea, et al.
Published: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
by: Zhao, Yiren, et al.
Published: (2026)
by: Zhao, Yiren, et al.
Published: (2026)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
by: DeBole, Michael V., et al.
Published: (2025)
by: DeBole, Michael V., et al.
Published: (2025)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
by: Lu, Chien-Ping
Published: (2026)
by: Lu, Chien-Ping
Published: (2026)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
PASS: An Asynchronous Probabilistic Processor for Next Generation Intelligence
by: Patel, Saavan, et al.
Published: (2024)
by: Patel, Saavan, et al.
Published: (2024)
ZettaLith: An Architectural Exploration of Extreme-Scale AI Inference Acceleration
by: Silverbrook, Kia
Published: (2025)
by: Silverbrook, Kia
Published: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
by: Pan, Yudong, et al.
Published: (2026)
by: Pan, Yudong, et al.
Published: (2026)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
by: Graziano, Marco
Published: (2026)
by: Graziano, Marco
Published: (2026)
NPU Design for Diffusion Language Model Inference
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
Investigating Memory Failure Prediction Across CPU Architectures
by: Yu, Qiao, et al.
Published: (2024)
by: Yu, Qiao, et al.
Published: (2024)
Evaluating Kubernetes Performance for GenAI Inference: From Automatic Speech Recognition to LLM Summarization
by: Malleni, Sai Sindhur, et al.
Published: (2026)
by: Malleni, Sai Sindhur, et al.
Published: (2026)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
by: Dorairaj, Nij, et al.
Published: (2026)
by: Dorairaj, Nij, et al.
Published: (2026)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
by: Barkhordar, Marzieh, et al.
Published: (2026)
by: Barkhordar, Marzieh, et al.
Published: (2026)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
by: Renney, Harri, et al.
Published: (2026)
by: Renney, Harri, et al.
Published: (2026)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Deep Tech to Space: Space Data Centers and AI Revolution at the Edge
by: Weiss, Jonas, et al.
Published: (2026)
by: Weiss, Jonas, et al.
Published: (2026)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
by: Liu, Fangxin, et al.
Published: (2026)
by: Liu, Fangxin, et al.
Published: (2026)
Similar Items
-
Managed-Retention Memory: A New Class of Memory for the AI Era
by: Legtchenko, Sergey, et al.
Published: (2025) -
WaferLLM: Large Language Model Inference at Wafer Scale
by: He, Congjie, et al.
Published: (2025) -
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
by: Tithi, Jesmin Jahan, et al.
Published: (2025) -
Transforming the Hybrid Cloud for Emerging AI Workloads
by: Chen, Deming, et al.
Published: (2024) -
Experience Deploying Containerized GenAI Services at an HPC Center
by: Beltre, Angel M., et al.
Published: (2025)