PRISM: Breaking the O(n) Memory Wall in Long-Context LLM Inference via O(1) Photonic Block Selection
Fuente:
arXiv
Salvato in:
| Autori principali: | Park, Hyoseok, Park, Yeonsang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
di: Matsushima, Kosuke, et al.
Pubblicazione: (2026)
di: Matsushima, Kosuke, et al.
Pubblicazione: (2026)
Harnessing Photonics for Machine Intelligence
di: Zhu, Hanqing, et al.
Pubblicazione: (2026)
di: Zhu, Hanqing, et al.
Pubblicazione: (2026)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
di: Liang, Yanbiao, et al.
Pubblicazione: (2025)
di: Liang, Yanbiao, et al.
Pubblicazione: (2025)
SimPhony: A Device-Circuit-Architecture Cross-Layer Modeling and Simulation Framework for Heterogeneous Electronic-Photonic AI System
di: Yin, Ziang, et al.
Pubblicazione: (2024)
di: Yin, Ziang, et al.
Pubblicazione: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)
AI Agents for Photonic Integrated Circuit Design Automation
di: Sharma, Ankita, et al.
Pubblicazione: (2025)
di: Sharma, Ankita, et al.
Pubblicazione: (2025)
LLM Inference Acceleration via Efficient Operation Fusion
di: Salmani, Mahsa, et al.
Pubblicazione: (2025)
di: Salmani, Mahsa, et al.
Pubblicazione: (2025)
Multi-Dimensional Reconfigurable, Physically Composable Hybrid Diffractive Optical Neural Network
di: Yin, Ziang, et al.
Pubblicazione: (2024)
di: Yin, Ziang, et al.
Pubblicazione: (2024)
Toward Large-Scale Photonics-Empowered AI Systems: From Physical Design Automation to System-Algorithm Co-Exploration
di: Yin, Ziang, et al.
Pubblicazione: (2025)
di: Yin, Ziang, et al.
Pubblicazione: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
di: Gope, Dibakar, et al.
Pubblicazione: (2024)
di: Gope, Dibakar, et al.
Pubblicazione: (2024)
Democratizing Electronic-Photonic AI Systems: An Open-Source AI-Infused Cross-Layer Co-Design and Design Automation Toolflow
di: Zhou, Hongjian, et al.
Pubblicazione: (2025)
di: Zhou, Hongjian, et al.
Pubblicazione: (2025)
Photonic Exponential Approximation via Cascaded TFLN Microring Resonators toward Softmax
di: Park, Hyoseok, et al.
Pubblicazione: (2026)
di: Park, Hyoseok, et al.
Pubblicazione: (2026)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
di: Wu, Haoran, et al.
Pubblicazione: (2025)
di: Wu, Haoran, et al.
Pubblicazione: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
di: Kim, Dowon, et al.
Pubblicazione: (2025)
di: Kim, Dowon, et al.
Pubblicazione: (2025)
Enhancing Realism in Holographic Augmented Reality Displays through Occlusion Handling
di: Han, Woongseob, et al.
Pubblicazione: (2025)
di: Han, Woongseob, et al.
Pubblicazione: (2025)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
di: Chen, Hongzheng, et al.
Pubblicazione: (2023)
di: Chen, Hongzheng, et al.
Pubblicazione: (2023)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
di: He, Zifan, et al.
Pubblicazione: (2025)
di: He, Zifan, et al.
Pubblicazione: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
di: Cong, Rongqing, et al.
Pubblicazione: (2024)
di: Cong, Rongqing, et al.
Pubblicazione: (2024)
Scaling Photonic Tensor Cores with Unary and Homodyne Designs
di: Alo, Oluwaseun, et al.
Pubblicazione: (2026)
di: Alo, Oluwaseun, et al.
Pubblicazione: (2026)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
di: Yun, Sungmin, et al.
Pubblicazione: (2025)
di: Yun, Sungmin, et al.
Pubblicazione: (2025)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
di: Negi, Shubham, et al.
Pubblicazione: (2025)
di: Negi, Shubham, et al.
Pubblicazione: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
di: Patwari, Rajeev, et al.
Pubblicazione: (2025)
di: Patwari, Rajeev, et al.
Pubblicazione: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
di: Zhang, Yu, et al.
Pubblicazione: (2024)
di: Zhang, Yu, et al.
Pubblicazione: (2024)
The Graph's Apprentice: Teaching an LLM Low Level Knowledge for Circuit Quality Estimation
di: Moravej, Reza, et al.
Pubblicazione: (2024)
di: Moravej, Reza, et al.
Pubblicazione: (2024)
hdl2v: A Code Translation Dataset for Enhanced LLM Verilog Generation
di: Hong, Charles, et al.
Pubblicazione: (2025)
di: Hong, Charles, et al.
Pubblicazione: (2025)
VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric
di: Nadimi, Bardia, et al.
Pubblicazione: (2025)
di: Nadimi, Bardia, et al.
Pubblicazione: (2025)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
di: Park, Dahoon, et al.
Pubblicazione: (2026)
di: Park, Dahoon, et al.
Pubblicazione: (2026)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
di: Du, Dayou, et al.
Pubblicazione: (2025)
di: Du, Dayou, et al.
Pubblicazione: (2025)
Nonlinear Computation with Linear Optics via Source-Position Encoding
di: Richardson, N., et al.
Pubblicazione: (2025)
di: Richardson, N., et al.
Pubblicazione: (2025)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
di: Fang, Chao, et al.
Pubblicazione: (2024)
di: Fang, Chao, et al.
Pubblicazione: (2024)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
Experimental Demonstration of an Optical Neural PDE Solver via On-Chip PINN Training
di: Zhao, Yequan, et al.
Pubblicazione: (2025)
di: Zhao, Yequan, et al.
Pubblicazione: (2025)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
di: Zhu, Zhantong, et al.
Pubblicazione: (2025)
di: Zhu, Zhantong, et al.
Pubblicazione: (2025)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
di: Pan, Xiurui, et al.
Pubblicazione: (2024)
di: Pan, Xiurui, et al.
Pubblicazione: (2024)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
di: Wu, Yutong, et al.
Pubblicazione: (2026)
di: Wu, Yutong, et al.
Pubblicazione: (2026)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
di: Demirkiran, Cansu, et al.
Pubblicazione: (2023)
di: Demirkiran, Cansu, et al.
Pubblicazione: (2023)
Pushing the Limits of BFP on Narrow Precision LLM Inference
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
di: Matsushima, Kosuke, et al.
Pubblicazione: (2026) -
Harnessing Photonics for Machine Intelligence
di: Zhu, Hanqing, et al.
Pubblicazione: (2026) -
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
di: Liang, Yanbiao, et al.
Pubblicazione: (2025) -
SimPhony: A Device-Circuit-Architecture Cross-Layer Modeling and Simulation Framework for Heterogeneous Electronic-Photonic AI System
di: Yin, Ziang, et al.
Pubblicazione: (2024) -
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
di: Kwon, Hyucksung, et al.
Pubblicazione: (2024)