Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yun, Sungmin, Park, Seonyong, Nam, Hwayong, Lee, Younjoo, Lee, Gunjun, Kyung, Kwanhee, Kim, Sangpyo, Kim, Nam Sung, Kim, Jongmin, Kim, Hyungyo, Cho, Juhwan, Baek, Seungmin, Ahn, Jung Ho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
von: Yun, Sungmin, et al.
Veröffentlicht: (2024)
von: Yun, Sungmin, et al.
Veröffentlicht: (2024)
SoK: Systematizing a Decade of Architectural RowHammer Defenses Through the Lens of Streaming Algorithms
von: Kim, Michael Jaemin, et al.
Veröffentlicht: (2025)
von: Kim, Michael Jaemin, et al.
Veröffentlicht: (2025)
PVAC: A RowHammer Mitigation Architecture Exploiting Per-victim-row Counting
von: Kim, Jumin, et al.
Veröffentlicht: (2026)
von: Kim, Jumin, et al.
Veröffentlicht: (2026)
RoMe: Row Granularity Access Memory System for Large Language Models
von: Nam, Hwayong, et al.
Veröffentlicht: (2025)
von: Nam, Hwayong, et al.
Veröffentlicht: (2025)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands
von: Nam, Hwayong, et al.
Veröffentlicht: (2024)
von: Nam, Hwayong, et al.
Veröffentlicht: (2024)
Per-Row Activation Counting on Real Hardware: Demystifying Performance Overheads
von: Kim, Jumin, et al.
Veröffentlicht: (2025)
von: Kim, Jumin, et al.
Veröffentlicht: (2025)
CiFHER: A Chiplet-Based FHE Accelerator with a Resizable Structure
von: Kim, Sangpyo, et al.
Veröffentlicht: (2023)
von: Kim, Sangpyo, et al.
Veröffentlicht: (2023)
IVE: An Accelerator for Single-Server Private Information Retrieval Using Versatile Processing Elements
von: Kim, Sangpyo, et al.
Veröffentlicht: (2025)
von: Kim, Sangpyo, et al.
Veröffentlicht: (2025)
Sudoku: Decomposing DRAM Address Mapping into Component Functions
von: Wi, Minbok, et al.
Veröffentlicht: (2025)
von: Wi, Minbok, et al.
Veröffentlicht: (2025)
AnalogToBi: Device-Level Analog Circuit Topology Generation via Bipartite Graph and Grammar Guided Decoding
von: Kim, Seungmin, et al.
Veröffentlicht: (2026)
von: Kim, Seungmin, et al.
Veröffentlicht: (2026)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
von: Kim, Taehyun, et al.
Veröffentlicht: (2024)
von: Kim, Taehyun, et al.
Veröffentlicht: (2024)
Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE Serving
von: Pan, Yue, et al.
Veröffentlicht: (2025)
von: Pan, Yue, et al.
Veröffentlicht: (2025)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
von: Lee, Minjae, et al.
Veröffentlicht: (2023)
von: Lee, Minjae, et al.
Veröffentlicht: (2023)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
von: Lee, Gunjun, et al.
Veröffentlicht: (2025)
von: Lee, Gunjun, et al.
Veröffentlicht: (2025)
Low-Power Encoding for PAM-3 DRAM Bus
von: Nam, Jonghyeon, et al.
Veröffentlicht: (2024)
von: Nam, Jonghyeon, et al.
Veröffentlicht: (2024)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
von: Park, Haneul, et al.
Veröffentlicht: (2025)
von: Park, Haneul, et al.
Veröffentlicht: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
von: Bambhaniya, Abhimanyu, et al.
Veröffentlicht: (2026)
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision Quantization
von: Kim, Seeyeon, et al.
Veröffentlicht: (2026)
von: Kim, Seeyeon, et al.
Veröffentlicht: (2026)
Theodosian: A Deep Dive into Memory-Hierarchy-Centric FHE Acceleration
von: Choi, Wonseok, et al.
Veröffentlicht: (2025)
von: Choi, Wonseok, et al.
Veröffentlicht: (2025)
Optimized Memory System Architecture for VESA VDC-M Decoder with Multi-Slice Support
von: Yang, Hannah, et al.
Veröffentlicht: (2025)
von: Yang, Hannah, et al.
Veröffentlicht: (2025)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
von: Han, Wontak, et al.
Veröffentlicht: (2024)
von: Han, Wontak, et al.
Veröffentlicht: (2024)
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
von: Kuper, Reese, et al.
Veröffentlicht: (2023)
von: Kuper, Reese, et al.
Veröffentlicht: (2023)
STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
von: Kwon, Omin, et al.
Veröffentlicht: (2025)
von: Kwon, Omin, et al.
Veröffentlicht: (2025)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
von: Hwang, Soojin, et al.
Veröffentlicht: (2025)
von: Hwang, Soojin, et al.
Veröffentlicht: (2025)
GPIR: Enabling Practical Private Information Retrieval with GPUs
von: Ji, Hyesung, et al.
Veröffentlicht: (2026)
von: Ji, Hyesung, et al.
Veröffentlicht: (2026)
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDs
von: Cho, Sungjun, et al.
Veröffentlicht: (2024)
von: Cho, Sungjun, et al.
Veröffentlicht: (2024)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
STRAW: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs
von: Chun, Myoungjun, et al.
Veröffentlicht: (2025)
von: Chun, Myoungjun, et al.
Veröffentlicht: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
von: Hong, Junguk, et al.
Veröffentlicht: (2026)
Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection
von: Kim, Junhwan, et al.
Veröffentlicht: (2026)
von: Kim, Junhwan, et al.
Veröffentlicht: (2026)
HURRY: Highly Utilized, Reconfigurable ReRAM-based In-situ Accelerator with Multifunctionality
von: Shin, Hery, et al.
Veröffentlicht: (2024)
von: Shin, Hery, et al.
Veröffentlicht: (2024)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
von: Moon, Seungjae, et al.
Veröffentlicht: (2024)
von: Moon, Seungjae, et al.
Veröffentlicht: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
von: Yun, Sungmin, et al.
Veröffentlicht: (2024) -
SoK: Systematizing a Decade of Architectural RowHammer Defenses Through the Lens of Streaming Algorithms
von: Kim, Michael Jaemin, et al.
Veröffentlicht: (2025) -
PVAC: A RowHammer Mitigation Architecture Exploiting Per-victim-row Counting
von: Kim, Jumin, et al.
Veröffentlicht: (2026) -
RoMe: Row Granularity Access Memory System for Large Language Models
von: Nam, Hwayong, et al.
Veröffentlicht: (2025) -
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)