Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Jingchen, Xue, Chenhao, Chen, Yiqi, Wang, Zhao, Zhang, Chen, Shen, Yu, Chen, Yifan, Cheng, Zekang, Jiang, Yu, Wang, Tianqi, Lin, Yibo, Hu, Wei, Cui, Bin, Wang, Runsheng, Liang, Yun, Sun, Guangyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
di: Wang, Zhao, et al.
Pubblicazione: (2021)
di: Wang, Zhao, et al.
Pubblicazione: (2021)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
di: Fu, Zizhuo, et al.
Pubblicazione: (2026)
di: Fu, Zizhuo, et al.
Pubblicazione: (2026)
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
di: Ren, Yi, et al.
Pubblicazione: (2025)
di: Ren, Yi, et al.
Pubblicazione: (2025)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
di: Iff, Patrick, et al.
Pubblicazione: (2026)
di: Iff, Patrick, et al.
Pubblicazione: (2026)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
DarwinWafer: A Wafer-Scale Neuromorphic Chip
di: Zhu, Xiaolei, et al.
Pubblicazione: (2025)
di: Zhu, Xiaolei, et al.
Pubblicazione: (2025)
CellE: Automated Standard Cell Library Extension via Equality Saturation
di: Ren, Yi, et al.
Pubblicazione: (2026)
di: Ren, Yi, et al.
Pubblicazione: (2026)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
di: Chen, Yiqi, et al.
Pubblicazione: (2025)
di: Chen, Yiqi, et al.
Pubblicazione: (2025)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
di: Ren, Yi, et al.
Pubblicazione: (2025)
di: Ren, Yi, et al.
Pubblicazione: (2025)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
di: Wang, Huizheng, et al.
Pubblicazione: (2025)
SEGA-DCIM: Design Space Exploration-Guided Automatic Digital CIM Compiler with Multiple Precision Support
di: Diao, Haikang, et al.
Pubblicazione: (2025)
di: Diao, Haikang, et al.
Pubblicazione: (2025)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
di: Luo, Shuqing, et al.
Pubblicazione: (2026)
di: Luo, Shuqing, et al.
Pubblicazione: (2026)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
LayoutCopilot: An LLM-powered Multi-agent Collaborative Framework for Interactive Analog Layout Design
di: Liu, Bingyang, et al.
Pubblicazione: (2024)
di: Liu, Bingyang, et al.
Pubblicazione: (2024)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
di: Li, Meng, et al.
Pubblicazione: (2026)
di: Li, Meng, et al.
Pubblicazione: (2026)
Optimizing and Exploring System Performance in Compact Processing-in-Memory-based Chips
di: Chen, Peilin, et al.
Pubblicazione: (2025)
di: Chen, Peilin, et al.
Pubblicazione: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
di: Wang, Zhao, et al.
Pubblicazione: (2025)
di: Wang, Zhao, et al.
Pubblicazione: (2025)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
di: Xie, Tong, et al.
Pubblicazione: (2026)
di: Xie, Tong, et al.
Pubblicazione: (2026)
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
di: Feng, Yinxiao, et al.
Pubblicazione: (2024)
di: Feng, Yinxiao, et al.
Pubblicazione: (2024)
E-morphic: Scalable Equality Saturation for Structural Exploration in Logic Synthesis
di: Chen, Chen, et al.
Pubblicazione: (2025)
di: Chen, Chen, et al.
Pubblicazione: (2025)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
Large Processor Chip Model
di: Chang, Kaiyan, et al.
Pubblicazione: (2025)
di: Chang, Kaiyan, et al.
Pubblicazione: (2025)
GSIM: Accelerating RTL Simulation for Large-Scale Designs
di: Chen, Lu, et al.
Pubblicazione: (2025)
di: Chen, Lu, et al.
Pubblicazione: (2025)
Towards Efficient and Accurate Detection of On-Chip Fail-Slow Failures for Many-Core Accelerators
di: Wu, Junchi, et al.
Pubblicazione: (2025)
di: Wu, Junchi, et al.
Pubblicazione: (2025)
A Systematic Approach for Multi-objective Double-side Clock Tree Synthesis
di: Jiang, Xun, et al.
Pubblicazione: (2025)
di: Jiang, Xun, et al.
Pubblicazione: (2025)
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
di: Bai, Kangbo, et al.
Pubblicazione: (2026)
di: Bai, Kangbo, et al.
Pubblicazione: (2026)
AnalogXpert: Automating Analog Topology Synthesis by Incorporating Circuit Design Expertise into Large Language Models
di: Zhang, Haoyi, et al.
Pubblicazione: (2024)
di: Zhang, Haoyi, et al.
Pubblicazione: (2024)
CircuitFusion: Multimodal Circuit Representation Learning for Agile Chip Design
di: Fang, Wenji, et al.
Pubblicazione: (2025)
di: Fang, Wenji, et al.
Pubblicazione: (2025)
AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
di: Xue, Chenhao, et al.
Pubblicazione: (2025)
Record Acceleration of the Two-Dimensional Ising Model Using High-Performance Wafer Scale Engine
di: Van Essendelft, Dirk, et al.
Pubblicazione: (2024)
di: Van Essendelft, Dirk, et al.
Pubblicazione: (2024)
Efficient yet Accurate End-to-End SC Accelerator Design
di: Li, Meng, et al.
Pubblicazione: (2024)
di: Li, Meng, et al.
Pubblicazione: (2024)
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
di: Wen, Jinqi, et al.
Pubblicazione: (2026)
di: Wen, Jinqi, et al.
Pubblicazione: (2026)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
di: Li, Cong, et al.
Pubblicazione: (2026)
di: Li, Cong, et al.
Pubblicazione: (2026)
Non-Overlapping Placement of Macro Cells based on Reinforcement Learning in Chip Design
di: Yu, Tao, et al.
Pubblicazione: (2024)
di: Yu, Tao, et al.
Pubblicazione: (2024)
Search-in-Memory (SiM): Reliable, Versatile, and Efficient Data Matching in SSD's NAND Flash Memory Chip for Data Indexing Acceleration
di: Chen, Yun-Chih, et al.
Pubblicazione: (2024)
di: Chen, Yun-Chih, et al.
Pubblicazione: (2024)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
di: Li, Huize, et al.
Pubblicazione: (2026)
di: Li, Huize, et al.
Pubblicazione: (2026)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
di: Lin, Chenqi, et al.
Pubblicazione: (2025)
di: Lin, Chenqi, et al.
Pubblicazione: (2025)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
di: Zhang, Rui, et al.
Pubblicazione: (2025)
di: Zhang, Rui, et al.
Pubblicazione: (2025)
EasyACIM: An End-to-End Automated Analog CIM with Synthesizable Architecture and Agile Design Space Exploration
di: Zhang, Haoyi, et al.
Pubblicazione: (2024)
di: Zhang, Haoyi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
di: Wang, Zhao, et al.
Pubblicazione: (2021) -
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
di: Fu, Zizhuo, et al.
Pubblicazione: (2026) -
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
di: Ren, Yi, et al.
Pubblicazione: (2025) -
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
di: Iff, Patrick, et al.
Pubblicazione: (2026) -
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
di: Liu, Yiqi, et al.
Pubblicazione: (2026)