Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zuoning, Parikh, Dhruv, Zhang, Youning, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stencil Computations on Cerebras Wafer-Scale Engine
by: Belli, Elia, et al.
Published: (2026)
by: Belli, Elia, et al.
Published: (2026)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)
by: Ramachandran, Anvitha, et al.
Published: (2025)
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
by: Gupta, Neelesh, et al.
Published: (2025)
by: Gupta, Neelesh, et al.
Published: (2025)
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026)
by: Ramachandran, Anvitha, et al.
Published: (2026)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Beyond Exascale: Dataflow Domain Translation on a Cerebras Cluster
by: Oppelstrup, Tomas, et al.
Published: (2025)
by: Oppelstrup, Tomas, et al.
Published: (2025)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024)
by: Parikh, Dhruv, et al.
Published: (2024)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
by: Cheng, Ke, et al.
Published: (2024)
by: Cheng, Ke, et al.
Published: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
by: Wickramasinghe, Sachini, et al.
Published: (2024)
by: Wickramasinghe, Sachini, et al.
Published: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
by: Tang, Xinru, et al.
Published: (2025)
by: Tang, Xinru, et al.
Published: (2025)
Scaling Performance of Large Language Model Pretraining
by: Interrante-Grant, Alexander, et al.
Published: (2025)
by: Interrante-Grant, Alexander, et al.
Published: (2025)
Trackable Agent-based Evolution Models at Wafer Scale
by: Moreno, Matthew Andres, et al.
Published: (2024)
by: Moreno, Matthew Andres, et al.
Published: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
by: Ren, Feng, et al.
Published: (2026)
by: Ren, Feng, et al.
Published: (2026)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025)
by: Golden, Alicia, et al.
Published: (2025)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024)
by: Yan, Ran, et al.
Published: (2024)
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
by: Coles, Jonathan, et al.
Published: (2026)
by: Coles, Jonathan, et al.
Published: (2026)
Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
by: Fan, Shuyuan, et al.
Published: (2025)
by: Fan, Shuyuan, et al.
Published: (2025)
Optimizing the Longhorn Cloud-native Software Defined Storage Engine for High Performance
by: Kampadais, Konstantinos, et al.
Published: (2025)
by: Kampadais, Konstantinos, et al.
Published: (2025)
Multi-agent Reinforcement Learning-based In-place Scaling Engine for Edge-cloud Systems
by: Prodanov, Jovan, et al.
Published: (2025)
by: Prodanov, Jovan, et al.
Published: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
ARGO: An Auto-Tuning Runtime System for Scalable GNN Training on Multi-Core Processor
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Cooling Matters: Benchmarking Large Language Models and Vision-Language Models on Liquid-Cooled Versus Air-Cooled H100 GPU Systems
by: Latif, Imran, et al.
Published: (2025)
by: Latif, Imran, et al.
Published: (2025)
WaferLLM: Large Language Model Inference at Wafer Scale
by: He, Congjie, et al.
Published: (2025)
by: He, Congjie, et al.
Published: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
Trackable Island-model Genetic Algorithms at Wafer Scale
by: Moreno, Matthew Andres, et al.
Published: (2024)
by: Moreno, Matthew Andres, et al.
Published: (2024)
Split Fine-Tuning for Large Language Models in Wireless Networks
by: Zhang, Songge, et al.
Published: (2025)
by: Zhang, Songge, et al.
Published: (2025)
Model Input Verification of Large Scale Simulations
by: Neykova, Rumyana, et al.
Published: (2024)
by: Neykova, Rumyana, et al.
Published: (2024)
Cascadia: An Efficient Cascade Serving System for Large Language Models
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Similar Items
-
Stencil Computations on Cerebras Wafer-Scale Engine
by: Belli, Elia, et al.
Published: (2026) -
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025) -
ScalableHD: Scalable and High-Throughput Hyperdimensional Computing Inference on Multi-Core CPUs
by: Parikh, Dhruv, et al.
Published: (2025) -
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
by: Gupta, Neelesh, et al.
Published: (2025) -
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026)