WaferLLM: Large Language Model Inference at Wafer Scale
Fuente:
arXiv
Saved in:
| Main Authors: | He, Congjie, Huang, Yeqi, Mu, Pei, Miao, Ziming, Xue, Jilong, Ma, Lingxiao, Yang, Fan, Mai, Luo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
by: Ortega, Cristobal, et al.
Published: (2024)
by: Ortega, Cristobal, et al.
Published: (2024)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
by: Call, Aaron, et al.
Published: (2025)
by: Call, Aaron, et al.
Published: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)
by: Puente, Valentin, et al.
Published: (2016)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
by: Roy, Rohan Basu, et al.
Published: (2025)
by: Roy, Rohan Basu, et al.
Published: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
by: Tithi, Jesmin Jahan, et al.
Published: (2025)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
by: Hou, Yuewen, et al.
Published: (2025)
by: Hou, Yuewen, et al.
Published: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
by: Filippov, Dmitry, et al.
Published: (2025)
by: Filippov, Dmitry, et al.
Published: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
by: Lee, Benjamin C., et al.
Published: (2024)
by: Lee, Benjamin C., et al.
Published: (2024)
Reference Architecture of a Quantum-Centric Supercomputer
by: Seelam, Seetharami, et al.
Published: (2026)
by: Seelam, Seetharami, et al.
Published: (2026)
Managed-Retention Memory: A New Class of Memory for the AI Era
by: Legtchenko, Sergey, et al.
Published: (2025)
by: Legtchenko, Sergey, et al.
Published: (2025)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
by: Garg, Chirag, et al.
Published: (2025)
by: Garg, Chirag, et al.
Published: (2025)
An MLIR Lowering Pipeline for Stencils at Wafer-Scale
by: Stawinoga, Nicolai, et al.
Published: (2026)
by: Stawinoga, Nicolai, et al.
Published: (2026)
PASS: An Asynchronous Probabilistic Processor for Next Generation Intelligence
by: Patel, Saavan, et al.
Published: (2024)
by: Patel, Saavan, et al.
Published: (2024)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Transforming the Hybrid Cloud for Emerging AI Workloads
by: Chen, Deming, et al.
Published: (2024)
by: Chen, Deming, et al.
Published: (2024)
Experience Deploying Containerized GenAI Services at an HPC Center
by: Beltre, Angel M., et al.
Published: (2025)
by: Beltre, Angel M., et al.
Published: (2025)
DarwinWafer: A Wafer-Scale Neuromorphic Chip
by: Zhu, Xiaolei, et al.
Published: (2025)
by: Zhu, Xiaolei, et al.
Published: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
by: Mo, Zhiwen, et al.
Published: (2026)
by: Mo, Zhiwen, et al.
Published: (2026)
Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System
by: Santos, Kylee, et al.
Published: (2024)
by: Santos, Kylee, et al.
Published: (2024)
Stencil Computations on Cerebras Wafer-Scale Engine
by: Belli, Elia, et al.
Published: (2026)
by: Belli, Elia, et al.
Published: (2026)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
by: Meng, William, et al.
Published: (2025)
by: Meng, William, et al.
Published: (2025)
Highly Versatile FPGA-Implemented Cyber Coherent Ising Machine
by: Aonishi, Toru, et al.
Published: (2024)
by: Aonishi, Toru, et al.
Published: (2024)
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
by: Liu, Yiqi, et al.
Published: (2025)
by: Liu, Yiqi, et al.
Published: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
by: Shi, Tianyao, et al.
Published: (2024)
by: Shi, Tianyao, et al.
Published: (2024)
Flex-PE: Flexible and SIMD Multi-Precision Processing Element for AI Workloads
by: Lokhande, Mukul, et al.
Published: (2024)
by: Lokhande, Mukul, et al.
Published: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
by: Li, Jiamin, et al.
Published: (2025)
by: Li, Jiamin, et al.
Published: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
by: Kreß, Fabian, et al.
Published: (2024)
by: Kreß, Fabian, et al.
Published: (2024)
Pooling Engram Conditional Memory in Large Language Models using CXL
by: Ma, Ruiyang, et al.
Published: (2026)
by: Ma, Ruiyang, et al.
Published: (2026)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
by: Qu, Huanyu, et al.
Published: (2025)
by: Qu, Huanyu, et al.
Published: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
Similar Items
-
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
by: Vo, Huynh Q. N., et al.
Published: (2025) -
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
by: Zhang, Xu, et al.
Published: (2024) -
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
by: Ortega, Cristobal, et al.
Published: (2024) -
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
by: Call, Aaron, et al.
Published: (2025) -
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)