TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhuoran, Bian, Zhuohang, Huang, Zihao, Sun, Guangyu, Liang, Yun, Zhuo, Youwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
von: Li, Peijing, et al.
Veröffentlicht: (2025)
von: Li, Peijing, et al.
Veröffentlicht: (2025)
Towards Efficient and Accurate Detection of On-Chip Fail-Slow Failures for Many-Core Accelerators
von: Wu, Junchi, et al.
Veröffentlicht: (2025)
von: Wu, Junchi, et al.
Veröffentlicht: (2025)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
von: Heo, Guseul, et al.
Veröffentlicht: (2024)
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
von: Palacios, Pedro, et al.
Veröffentlicht: (2024)
von: Palacios, Pedro, et al.
Veröffentlicht: (2024)
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2025)
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2025)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms
von: Waqar, Faaiq, et al.
Veröffentlicht: (2025)
von: Waqar, Faaiq, et al.
Veröffentlicht: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
The Monte Carlo Method and New Device and Architectural Techniques for Accelerating It
von: Petangoda, Janith, et al.
Veröffentlicht: (2025)
von: Petangoda, Janith, et al.
Veröffentlicht: (2025)
SynapticCore-X: A Modular Neural Processing Architecture for Low-Cost FPGA Acceleration
von: Parameshwara, Arya
Veröffentlicht: (2025)
von: Parameshwara, Arya
Veröffentlicht: (2025)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
A 65 nm Bayesian Neural Network Accelerator with 360 fJ/Sample In-Word GRNG for AI Uncertainty Estimation
von: Enciso, Zephan M., et al.
Veröffentlicht: (2025)
von: Enciso, Zephan M., et al.
Veröffentlicht: (2025)
LIPPEN: A Lightweight In-Place Pointer Encryption Architecture for Pointer Integrity
von: Iravani, Erfan, et al.
Veröffentlicht: (2026)
von: Iravani, Erfan, et al.
Veröffentlicht: (2026)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2026)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
von: Tahmasebi, Faraz, et al.
Veröffentlicht: (2024)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
von: Malekar, Jinendra, et al.
Veröffentlicht: (2025)
von: Malekar, Jinendra, et al.
Veröffentlicht: (2025)
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
von: Lin, Ye, et al.
Veröffentlicht: (2026)
von: Lin, Ye, et al.
Veröffentlicht: (2026)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
von: Liu, Lian, et al.
Veröffentlicht: (2025)
von: Liu, Lian, et al.
Veröffentlicht: (2025)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
von: Chen, Kangqi, et al.
Veröffentlicht: (2025)
von: Chen, Kangqi, et al.
Veröffentlicht: (2025)
FARe: Fault-Aware GNN Training on ReRAM-based PIM Accelerators
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024)
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2024)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
von: Ferreira, João Dinis, et al.
Veröffentlicht: (2021)
Anvil: A General-Purpose Timing-Safe Hardware Description Language
von: Yu, Jason Zhijingcheng, et al.
Veröffentlicht: (2025)
von: Yu, Jason Zhijingcheng, et al.
Veröffentlicht: (2025)
Biological Intuition on Digital Hardware: An RTL Implementation of Poisson-Encoded SNNs for Static Image Classification
von: Das, Debabrata, et al.
Veröffentlicht: (2026)
von: Das, Debabrata, et al.
Veröffentlicht: (2026)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
von: Giorgi, Roberto, et al.
Veröffentlicht: (2026)
von: Giorgi, Roberto, et al.
Veröffentlicht: (2026)
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
von: Jang, Yongjoo, et al.
Veröffentlicht: (2025)
von: Jang, Yongjoo, et al.
Veröffentlicht: (2025)
Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip Networks
von: Wu, Qizhe, et al.
Veröffentlicht: (2024)
von: Wu, Qizhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
von: Yuan, Aojie, et al.
Veröffentlicht: (2026) -
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023) -
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
von: He, Siyuan, et al.
Veröffentlicht: (2025) -
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
von: Sanovar, Rya, et al.
Veröffentlicht: (2024) -
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
von: Li, Peijing, et al.
Veröffentlicht: (2025)