What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Renjie, Wei, Wenjie, Xin, Qi, Liu, Xiaoli, Mao, Sixuan, Ma, Erik, Chen, Zijian, Zhang, Malu, Li, Haizhou, Zhang, Zhaoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916728702763008
author Li, Renjie
Wei, Wenjie
Xin, Qi
Liu, Xiaoli
Mao, Sixuan
Ma, Erik
Chen, Zijian
Zhang, Malu
Li, Haizhou
Zhang, Zhaoyu
author_facet Li, Renjie
Wei, Wenjie
Xin, Qi
Liu, Xiaoli
Mao, Sixuan
Ma, Erik
Chen, Zijian
Zhang, Malu
Li, Haizhou
Zhang, Zhaoyu
contents Large language models (LLMs) are rapidly pushing the limits of contemporary computing hardware. For example, training GPT-3 has been estimated to consume around 1300 MWh of electricity, and projections suggest future models may require city-scale (gigawatt) power budgets. These demands motivate exploration of computing paradigms beyond conventional von Neumann architectures. This review surveys emerging photonic hardware optimized for next-generation generative AI computing. We discuss integrated photonic neural network architectures (e.g., Mach-Zehnder interferometer meshes, lasers, wavelength-multiplexed microring resonators) that perform ultrafast matrix operations. We also examine promising alternative neuromorphic devices, including spiking neural network circuits and hybrid spintronic-photonic synapses, which combine memory and processing. The integration of two-dimensional materials (graphene, TMDCs) into silicon photonic platforms is reviewed for tunable modulators and on-chip synaptic elements. Transformer-based LLM architectures (self-attention and feed-forward layers) are analyzed in this context, identifying strategies and challenges for mapping dynamic matrix multiplications onto these novel hardware substrates. We then dissect the mechanisms of mainstream LLMs, such as ChatGPT, DeepSeek, and LLaMA, highlighting their architectural similarities and differences. We synthesize state-of-the-art components, algorithms, and integration methods, highlighting key advances and open issues in scaling such systems to mega-sized LLM models. We find that photonic computing systems could potentially surpass electronic processors by orders of magnitude in throughput and energy efficiency, but require breakthroughs in memory, especially for long-context windows and long token sequences, and in storage of ultra-large datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
Li, Renjie
Wei, Wenjie
Xin, Qi
Liu, Xiaoli
Mao, Sixuan
Ma, Erik
Chen, Zijian
Zhang, Malu
Li, Haizhou
Zhang, Zhaoyu
Hardware Architecture
Artificial Intelligence
Neural and Evolutionary Computing
Large language models (LLMs) are rapidly pushing the limits of contemporary computing hardware. For example, training GPT-3 has been estimated to consume around 1300 MWh of electricity, and projections suggest future models may require city-scale (gigawatt) power budgets. These demands motivate exploration of computing paradigms beyond conventional von Neumann architectures. This review surveys emerging photonic hardware optimized for next-generation generative AI computing. We discuss integrated photonic neural network architectures (e.g., Mach-Zehnder interferometer meshes, lasers, wavelength-multiplexed microring resonators) that perform ultrafast matrix operations. We also examine promising alternative neuromorphic devices, including spiking neural network circuits and hybrid spintronic-photonic synapses, which combine memory and processing. The integration of two-dimensional materials (graphene, TMDCs) into silicon photonic platforms is reviewed for tunable modulators and on-chip synaptic elements. Transformer-based LLM architectures (self-attention and feed-forward layers) are analyzed in this context, identifying strategies and challenges for mapping dynamic matrix multiplications onto these novel hardware substrates. We then dissect the mechanisms of mainstream LLMs, such as ChatGPT, DeepSeek, and LLaMA, highlighting their architectural similarities and differences. We synthesize state-of-the-art components, algorithms, and integration methods, highlighting key advances and open issues in scaling such systems to mega-sized LLM models. We find that photonic computing systems could potentially surpass electronic processors by orders of magnitude in throughput and energy efficiency, but require breakthroughs in memory, especially for long-context windows and long token sequences, and in storage of ultra-large datasets.
title What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
topic Hardware Architecture
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2505.05794