A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Hongsun, Song, Jaeyong, Shin, Changmin, Noh, Si Ung, Jung, Jaewon, Park, Jisung, Lee, Jinho |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
by: Jang, Hongsun, et al.
Published: (2024)
by: Jang, Hongsun, et al.
Published: (2024)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
by: Shin, Changmin, et al.
Published: (2025)
by: Shin, Changmin, et al.
Published: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
by: Hong, Junguk, et al.
Published: (2026)
by: Hong, Junguk, et al.
Published: (2026)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
by: Pan, Xiurui, et al.
Published: (2024)
by: Pan, Xiurui, et al.
Published: (2024)
PeerAiD: Improving Adversarial Distillation from a Specialized Peer Tutor
by: Jung, Jaewon, et al.
Published: (2024)
by: Jung, Jaewon, et al.
Published: (2024)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
by: Noh, Si Ung, et al.
Published: (2024)
by: Noh, Si Ung, et al.
Published: (2024)
STRAW: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs
by: Chun, Myoungjun, et al.
Published: (2025)
by: Chun, Myoungjun, et al.
Published: (2025)
Flexible In-NAND Cryptographic Processing for Secure Flash Storage
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
All-rounder: A Flexible AI Accelerator with Diverse Data Format Support and Morphable Structure for Multi-DNN Processing
by: Noh, Seock-Hwan, et al.
Published: (2023)
by: Noh, Seock-Hwan, et al.
Published: (2023)
GraNNDis: Efficient Unified Distributed Training Framework for Deep GNNs on Large Clusters
by: Song, Jaeyong, et al.
Published: (2023)
by: Song, Jaeyong, et al.
Published: (2023)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
by: Ghiasi, Nika Mansouri, et al.
Published: (2024)
by: Ghiasi, Nika Mansouri, et al.
Published: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025)
by: Lee, Joonhyung, et al.
Published: (2025)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
by: Wu, Haoran, et al.
Published: (2025)
by: Wu, Haoran, et al.
Published: (2025)
Containerized In-Storage Processing and Computing-Enabled SSD Disaggregation
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
by: Kim, Dowon, et al.
Published: (2025)
by: Kim, Dowon, et al.
Published: (2025)
AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDs
by: Cho, Sungjun, et al.
Published: (2024)
by: Cho, Sungjun, et al.
Published: (2024)
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models
by: Yang, Jinho, et al.
Published: (2025)
by: Yang, Jinho, et al.
Published: (2025)
Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
by: Kwon, Jaewon, et al.
Published: (2025)
by: Kwon, Jaewon, et al.
Published: (2025)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
by: Sun, He, et al.
Published: (2026)
by: Sun, He, et al.
Published: (2026)
GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
by: Pei, Minnan, et al.
Published: (2025)
by: Pei, Minnan, et al.
Published: (2025)
Pipette: Automatic Fine-grained Large Language Model Training Configurator for Real-World Clusters
by: Yim, Jinkyu, et al.
Published: (2024)
by: Yim, Jinkyu, et al.
Published: (2024)
Dissecting and Re-architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMs
by: Jang, Yongjoo, et al.
Published: (2025)
by: Jang, Yongjoo, et al.
Published: (2025)
STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
by: Kwon, Omin, et al.
Published: (2025)
by: Kwon, Omin, et al.
Published: (2025)
REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
by: Chen, Kangqi, et al.
Published: (2025)
by: Chen, Kangqi, et al.
Published: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
by: Liu, Zizhen, et al.
Published: (2026)
by: Liu, Zizhen, et al.
Published: (2026)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
by: Liu, Qingyuan, et al.
Published: (2025)
by: Liu, Qingyuan, et al.
Published: (2025)
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
PRISM: Breaking the O(n) Memory Wall in Long-Context LLM Inference via O(1) Photonic Block Selection
by: Park, Hyoseok, et al.
Published: (2026)
by: Park, Hyoseok, et al.
Published: (2026)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
by: Shin, Yongwon, et al.
Published: (2024)
by: Shin, Yongwon, et al.
Published: (2024)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Developing Cost-Effective Drones for 5G Non-Terrestrial Network Research and Experimentation
by: Cáceres, Carlos de Quinto, et al.
Published: (2024)
by: Cáceres, Carlos de Quinto, et al.
Published: (2024)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
by: Li, Huize, et al.
Published: (2026)
by: Li, Huize, et al.
Published: (2026)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) Flash
by: Pinge, Sumukh, et al.
Published: (2025)
by: Pinge, Sumukh, et al.
Published: (2025)
Similar Items
-
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
by: Jang, Hongsun, et al.
Published: (2024) -
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
by: Shin, Changmin, et al.
Published: (2025) -
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
by: Hong, Junguk, et al.
Published: (2026) -
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
by: Pan, Xiurui, et al.
Published: (2024) -
PeerAiD: Improving Adversarial Distillation from a Specialized Peer Tutor
by: Jung, Jaewon, et al.
Published: (2024)