A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kundu, Yudhishthira, Kaur, Manroop, Wig, Tripty, Kumar, Kriti, Kumari, Pushpanjali, Puri, Vivek, Arora, Manish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
von: Luo, Weile, et al.
Veröffentlicht: (2024)
von: Luo, Weile, et al.
Veröffentlicht: (2024)
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
von: Feng, Yinxiao, et al.
Veröffentlicht: (2024)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2024)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
von: Iff, Patrick, et al.
Veröffentlicht: (2026)
von: Iff, Patrick, et al.
Veröffentlicht: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
von: Zhu, Jingchen, et al.
Veröffentlicht: (2024)
von: Zhu, Jingchen, et al.
Veröffentlicht: (2024)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
DarwinWafer: A Wafer-Scale Neuromorphic Chip
von: Zhu, Xiaolei, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaolei, et al.
Veröffentlicht: (2025)
Technology solutions targeting the performance of gen-AI inference in resource constrained platforms
von: Kundu, Joyjit, et al.
Veröffentlicht: (2026)
von: Kundu, Joyjit, et al.
Veröffentlicht: (2026)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Record Acceleration of the Two-Dimensional Ising Model Using High-Performance Wafer Scale Engine
von: Van Essendelft, Dirk, et al.
Veröffentlicht: (2024)
von: Van Essendelft, Dirk, et al.
Veröffentlicht: (2024)
Analyzing Modern NVIDIA GPU cores
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
von: Rashidi, Saeed, et al.
Veröffentlicht: (2024)
von: Rashidi, Saeed, et al.
Veröffentlicht: (2024)
JExplore: Design Space Exploration Tool for Nvidia Jetson Boards
von: Kutukcu, Basar, et al.
Veröffentlicht: (2025)
von: Kutukcu, Basar, et al.
Veröffentlicht: (2025)
Fixed and Movable Antenna Technology for 6G Integrated Sensing and Communication
von: Zeng, Yong, et al.
Veröffentlicht: (2024)
von: Zeng, Yong, et al.
Veröffentlicht: (2024)
RoboGPU: Accelerating GPU Collision Detection for Robotics
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
Design of a GPU with Heterogeneous Cores for Graphics
von: Tomás, Aurora, et al.
Veröffentlicht: (2026)
von: Tomás, Aurora, et al.
Veröffentlicht: (2026)
COOK Access Control on an embedded Volta GPU
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
von: Lesage, Benjamin, et al.
Veröffentlicht: (2024)
Hardware Accelerators for Artificial Intelligence
von: Ahsan, S M Mojahidul, et al.
Veröffentlicht: (2024)
von: Ahsan, S M Mojahidul, et al.
Veröffentlicht: (2024)
Multiport Support for Vortex OpenGPU Memory Hierarchy
von: Shin, Injae, et al.
Veröffentlicht: (2025)
von: Shin, Injae, et al.
Veröffentlicht: (2025)
CuLifter: Lifting GPU Binaries to Typed IR
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions
von: Boutros, Andrew, et al.
Veröffentlicht: (2024)
von: Boutros, Andrew, et al.
Veröffentlicht: (2024)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2025)
von: Mhatre, Kaustubh, et al.
Veröffentlicht: (2025)
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification
von: Patel, Vihaan, et al.
Veröffentlicht: (2026)
von: Patel, Vihaan, et al.
Veröffentlicht: (2026)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
von: Nagendra, Savinay
Veröffentlicht: (2024)
von: Nagendra, Savinay
Veröffentlicht: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
Towards Efficient Design Verification -- Constrained Random Verification using PyUVM
von: Gadde, Deepak Narayan, et al.
Veröffentlicht: (2024)
von: Gadde, Deepak Narayan, et al.
Veröffentlicht: (2024)
CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems
von: Wu, Qihang, et al.
Veröffentlicht: (2026)
von: Wu, Qihang, et al.
Veröffentlicht: (2026)
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
von: Parvathy, Aradhana Mohan, et al.
Veröffentlicht: (2026)
von: Parvathy, Aradhana Mohan, et al.
Veröffentlicht: (2026)
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
von: Latif, Imran, et al.
Veröffentlicht: (2024)
von: Latif, Imran, et al.
Veröffentlicht: (2024)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
von: Rahoof, Abdul, et al.
Veröffentlicht: (2025)
Neuromorphic Computing for Low-Power Artificial Intelligence
von: Katti, Keshava, et al.
Veröffentlicht: (2026)
von: Katti, Keshava, et al.
Veröffentlicht: (2026)
GPU-Accelerated Optimization Solver for Unit Commitment in Large-Scale Power Grids
von: Sharadga, Hussein, et al.
Veröffentlicht: (2025)
von: Sharadga, Hussein, et al.
Veröffentlicht: (2025)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2023)
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2023)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
von: Guan, Yue, et al.
Veröffentlicht: (2026)
von: Guan, Yue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
von: Luo, Weile, et al.
Veröffentlicht: (2024) -
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
von: Feng, Yinxiao, et al.
Veröffentlicht: (2024) -
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
von: Iff, Patrick, et al.
Veröffentlicht: (2026) -
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025) -
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
von: Zhu, Jingchen, et al.
Veröffentlicht: (2024)