Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Zhongkai, Liang, Shengwen, Ma, Tianyun, Cai, Yunke, Nan, Ziyuan, Huang, Di, Song, Xinkai, Hao, Yifan, Zhang, Jie, Zhi, Tian, Zhao, Yongwei, Du, Zidong, Hu, Xing, Guo, Qi, Chen, Tianshi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
by: Bai, Kangbo, et al.
Published: (2026)
by: Bai, Kangbo, et al.
Published: (2026)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
by: Chong, Yue Jiet, et al.
Published: (2025)
by: Chong, Yue Jiet, et al.
Published: (2025)
Chiplet Actuary: A Quantitative Cost Model and Multi-Chiplet Architecture Exploration
by: Feng, Yinxiao, et al.
Published: (2022)
by: Feng, Yinxiao, et al.
Published: (2022)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
by: Iff, Patrick, et al.
Published: (2023)
by: Iff, Patrick, et al.
Published: (2023)
Sangam: Chiplet-Based DRAM-PIM Accelerator with CXL Integration for LLM Inferencing
by: Kiyawat, Khyati, et al.
Published: (2025)
by: Kiyawat, Khyati, et al.
Published: (2025)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
by: Hao, Mingbo, et al.
Published: (2026)
by: Hao, Mingbo, et al.
Published: (2026)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
by: Jiang, Aojie, et al.
Published: (2026)
by: Jiang, Aojie, et al.
Published: (2026)
The Survey of Chiplet-based Integrated Architecture: An EDA perspective
by: Chen, Shixin, et al.
Published: (2024)
by: Chen, Shixin, et al.
Published: (2024)
AGON: Automated Design Framework for Customizing Processors from ISA Documents
by: Li, Chongxiao, et al.
Published: (2024)
by: Li, Chongxiao, et al.
Published: (2024)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
by: Xu, Weikai, et al.
Published: (2026)
by: Xu, Weikai, et al.
Published: (2026)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
by: Lin, Zi-Wei, et al.
Published: (2026)
by: Lin, Zi-Wei, et al.
Published: (2026)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
by: Liu, Zizhen, et al.
Published: (2026)
by: Liu, Zizhen, et al.
Published: (2026)
Chiplets on Wheels: Review Paper on Holistic Chiplet Solutions for Autonomous Vehicles
by: Narashiman, Swathi, et al.
Published: (2024)
by: Narashiman, Swathi, et al.
Published: (2024)
ECO-CHIP: Estimation of Carbon Footprint of Chiplet-based Architectures for Sustainable VLSI
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
by: Sudarshan, Chetan Choppali, et al.
Published: (2023)
Chiplet-Gym: Optimizing Chiplet-based AI Accelerator Design with Reinforcement Learning
by: Mishty, Kaniz, et al.
Published: (2024)
by: Mishty, Kaniz, et al.
Published: (2024)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
by: Zheng, Jianing, et al.
Published: (2025)
by: Zheng, Jianing, et al.
Published: (2025)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
by: Ganti, Ravindra, et al.
Published: (2026)
by: Ganti, Ravindra, et al.
Published: (2026)
LEXI: Lossless Exponent Coding for Efficient Inter-Chiplet Communication in Hybrid LLMs
by: Sun, Miao, et al.
Published: (2026)
by: Sun, Miao, et al.
Published: (2026)
CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems
by: Wu, Qihang, et al.
Published: (2026)
by: Wu, Qihang, et al.
Published: (2026)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
by: Yu, Feng, et al.
Published: (2026)
by: Yu, Feng, et al.
Published: (2026)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
by: Wang, Cong, et al.
Published: (2025)
by: Wang, Cong, et al.
Published: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
by: Ali, Wajid, et al.
Published: (2025)
by: Ali, Wajid, et al.
Published: (2025)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
by: Iff, Patrick, et al.
Published: (2025)
by: Iff, Patrick, et al.
Published: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
by: Kanani, Alish, et al.
Published: (2025)
by: Kanani, Alish, et al.
Published: (2025)
MFIT: Multi-Fidelity Thermal Modeling for 2.5D and 3D Multi-Chiplet Architectures
by: Pfromm, Lukas, et al.
Published: (2024)
by: Pfromm, Lukas, et al.
Published: (2024)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
by: Luo, Shuqing, et al.
Published: (2026)
by: Luo, Shuqing, et al.
Published: (2026)
Designing High-Performance and Thermally Feasible Multi-Chiplet Architectures enabled by Non-bendable Glass Interposer
by: Sharma, Harsh, et al.
Published: (2025)
by: Sharma, Harsh, et al.
Published: (2025)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
by: He, Siyuan, et al.
Published: (2025)
by: He, Siyuan, et al.
Published: (2025)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
by: Pan, Xiurui, et al.
Published: (2024)
by: Pan, Xiurui, et al.
Published: (2024)
Link Quality Aware Pathfinding for Chiplet Interconnects
by: Yen, Aaron, et al.
Published: (2026)
by: Yen, Aaron, et al.
Published: (2026)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
by: Wang, Yimin, et al.
Published: (2025)
by: Wang, Yimin, et al.
Published: (2025)
FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic Encryption
by: Shi, Shangyi, et al.
Published: (2025)
by: Shi, Shangyi, et al.
Published: (2025)
DS2SC-Agent: A Multi-Agent Automated Pipeline for Rapid Chiplet Model Generation
by: Wu, Yiwei, et al.
Published: (2026)
by: Wu, Yiwei, et al.
Published: (2026)
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
by: Iff, Patrick, et al.
Published: (2025)
by: Iff, Patrick, et al.
Published: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
by: He, Siyuan, et al.
Published: (2025)
by: He, Siyuan, et al.
Published: (2025)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Similar Items
-
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
by: Bai, Kangbo, et al.
Published: (2026) -
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
by: Li, Boyu, et al.
Published: (2025) -
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
by: Chong, Yue Jiet, et al.
Published: (2025) -
Chiplet Actuary: A Quantitative Cost Model and Multi-Chiplet Architecture Exploration
by: Feng, Yinxiao, et al.
Published: (2022) -
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
by: Iff, Patrick, et al.
Published: (2023)