Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
Fuente:
arXiv
Saved in:
| Main Authors: | Sung, Mingyu, Palakonda, Vikas, Im, Suhwan, Moon, Sunghwan, Kim, Il-Min, Yun, Sangseok, Kang, Jae-Mo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
Impact of Large Language Models of Code on Fault Localization
by: Ji, Suhwan, et al.
Published: (2024)
by: Ji, Suhwan, et al.
Published: (2024)
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
by: Jo, Dongwon, et al.
Published: (2024)
by: Jo, Dongwon, et al.
Published: (2024)
Secure Power Control for Downlink Cell-Free Massive MIMO With Passive Eavesdroppers
by: Park, Junguk, et al.
Published: (2022)
by: Park, Junguk, et al.
Published: (2022)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
by: Moon, Seungjae, et al.
Published: (2024)
by: Moon, Seungjae, et al.
Published: (2024)
Distillation of Large Language Models via Concrete Score Matching
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
Reconstruction of the initial data from the trace of the solutions on an infinite time cylinder of damped wave equations
by: Kim, Seongyeon, et al.
Published: (2023)
by: Kim, Seongyeon, et al.
Published: (2023)
Explicit inversion for variable-speed wave equations on bounded domains
by: Moon, Sunghwan, et al.
Published: (2023)
by: Moon, Sunghwan, et al.
Published: (2023)
Trimma: Trimming Metadata Storage and Latency for Hybrid Memory Systems
by: Li, Yiwei, et al.
Published: (2024)
by: Li, Yiwei, et al.
Published: (2024)
Multifunctional In‐Memory Analog‐to‐Digital Converter for Next‐Gen Compute‐in‐Memory Systems
by: Jiseong Im, et al.
Published: (2024)
by: Jiseong Im, et al.
Published: (2024)
Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
by: Kang, Dongjin, et al.
Published: (2024)
by: Kang, Dongjin, et al.
Published: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
by: Moon, Sehwan, et al.
Published: (2025)
by: Moon, Sehwan, et al.
Published: (2025)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
by: Kang, Beomseok, et al.
Published: (2025)
by: Kang, Beomseok, et al.
Published: (2025)
Large Language Models Facilitate Vision Reflection in Image Classification
by: An, Guoyuan, et al.
Published: (2025)
by: An, Guoyuan, et al.
Published: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
Adaptive Computational Methods for Robust Object Tracking
by: Ji-Hoon Kwon, et al.
Published: (2024)
by: Ji-Hoon Kwon, et al.
Published: (2024)
Association Between Intracochlear Electrode Design and Electrically‐Evoked Compound Action Potential Measures in Cochlear Implant Users
by: Jeong‐Seo Kim, et al.
Published: (2024)
by: Jeong‐Seo Kim, et al.
Published: (2024)
Front Cover: Large‐Scale and Highly Reliable Hopfield Neural Networks Using Vertical NAND Flash Memory for the In‐Memory Associative Computing (Adv. Intell. Syst. 4/2026)
by: Jin Ho Chang, et al.
Published: (2026)
by: Jin Ho Chang, et al.
Published: (2026)
Large Language Model Partitioning for Low-Latency Inference at the Edge
by: Kafetzis, Dimitrios, et al.
Published: (2025)
by: Kafetzis, Dimitrios, et al.
Published: (2025)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
by: Choi, Suhwan, et al.
Published: (2025)
by: Choi, Suhwan, et al.
Published: (2025)
A Formalin-Inactivated Vaccine Enhances Survival and Mitigates Horizontal Transmission of Red Sea Bream Iridovirus (RSIV) in Rock Bream (Oplegnathus fasciatus): Insights From Viability Quantitative PCR.
by: Moon, Sung-Bin, et al.
Published: (2026)
by: Moon, Sung-Bin, et al.
Published: (2026)
A Formalin‐Inactivated Vaccine Enhances Survival and Mitigates Horizontal Transmission of Red Sea Bream Iridovirus ( RSIV ) in Rock Bream ( Oplegnathus fasciatus ): Insights From Viability Quantitative PCR
by: Sung‐Bin Moon, et al.
Published: (2026)
by: Sung‐Bin Moon, et al.
Published: (2026)
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
by: Na, Byeonghu, et al.
Published: (2026)
by: Na, Byeonghu, et al.
Published: (2026)
A Machine Learning and Finite Element Framework for Inverse Elliptic PDEs via Dirichlet-to-Neumann Mapping
by: Park, Dabin, et al.
Published: (2025)
by: Park, Dabin, et al.
Published: (2025)
Photoacoustic tomography with time-dependent damping: Theoretical and a convolutional neural network-guided numerical inversion procedure
by: Moon, Sunghwan, et al.
Published: (2026)
by: Moon, Sunghwan, et al.
Published: (2026)
Removal of Acid Orange 7 dye using Makgeolli lees with ultrasonic assistance
by: Nguyen Van Kien, et al.
Published: (2024)
by: Nguyen Van Kien, et al.
Published: (2024)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
Memory-aware feedback enhances power in active information engines
by: Bahng, Sehoon, et al.
Published: (2025)
by: Bahng, Sehoon, et al.
Published: (2025)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
by: Hwang, Soojin, et al.
Published: (2025)
by: Hwang, Soojin, et al.
Published: (2025)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
by: Kim, Geon-Woo, et al.
Published: (2025)
by: Kim, Geon-Woo, et al.
Published: (2025)
Artificial Modulation of the Hydrogen Evolution Reaction Kinetics via Control of Grain Boundaries Density in Mo2C Through Laser Processing (Adv. Funct. Mater. 28/2025)
by: Seok‐Ki Hyeong, et al.
Published: (2025)
by: Seok‐Ki Hyeong, et al.
Published: (2025)
Performance of Remote Task in Multi‐Robot Construction: Focusing on Pick‐and‐Place
by: Sanggyu Lee, et al.
Published: (2026)
by: Sanggyu Lee, et al.
Published: (2026)
71‐3: High Resolution Pixel Circuit Using a Double‐Gate LTPS TFT for AMOLED Displays in AR and VR Applications
by: Jae-Young Heo, et al.
Published: (2024)
by: Jae-Young Heo, et al.
Published: (2024)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
Similar Items
-
Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
by: Sung, Mingyu, et al.
Published: (2025) -
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
by: Sung, Mingyu, et al.
Published: (2025) -
H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models
by: Sung, Mingyu, et al.
Published: (2025) -
No Pose Estimation? No Problem: Pose-Agnostic and Instance-Aware Test-Time Adaptation for Monocular Depth Estimation
by: Sung, Mingyu, et al.
Published: (2025) -
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
by: Sung, Mingyu, et al.
Published: (2025)