MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Das, Anurita |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Temporal Abstraction in Reinforcement Learning with Offline Data
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2024)
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2024)
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
von: Sander, Jacob, et al.
Veröffentlicht: (2025)
von: Sander, Jacob, et al.
Veröffentlicht: (2025)
Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents
von: Borro, Luiz C., et al.
Veröffentlicht: (2026)
von: Borro, Luiz C., et al.
Veröffentlicht: (2026)
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices
von: Wang, Zhaode, et al.
Veröffentlicht: (2025)
von: Wang, Zhaode, et al.
Veröffentlicht: (2025)
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
von: Zhou, Jinan, et al.
Veröffentlicht: (2025)
von: Zhou, Jinan, et al.
Veröffentlicht: (2025)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment
von: AI, Lumen, et al.
Veröffentlicht: (2025)
von: AI, Lumen, et al.
Veröffentlicht: (2025)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
A Study of Skews, Imbalances, and Pathological Conditions in LLM Inference Deployment on GPU Clusters detectable from DPU
von: Moye, Javed I. Khan an Henry Uwabor
Veröffentlicht: (2025)
von: Moye, Javed I. Khan an Henry Uwabor
Veröffentlicht: (2025)
SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
von: Zhuge, Xiangwen, et al.
Veröffentlicht: (2025)
von: Zhuge, Xiangwen, et al.
Veröffentlicht: (2025)
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling
von: Li, Jialong, et al.
Veröffentlicht: (2024)
von: Li, Jialong, et al.
Veröffentlicht: (2024)
Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
von: Sung, Mingyu, et al.
Veröffentlicht: (2025)
von: Sung, Mingyu, et al.
Veröffentlicht: (2025)
LIDS: LLM Summary Inference Under the Layered Lens
von: Park, Dylan, et al.
Veröffentlicht: (2026)
von: Park, Dylan, et al.
Veröffentlicht: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Energy-Efficient Vision Transformer Inference for Edge-AI Deployment
von: Amanzhol, Nursultan, et al.
Veröffentlicht: (2025)
von: Amanzhol, Nursultan, et al.
Veröffentlicht: (2025)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
FIT-GNN: Faster Inference Time for GNNs that 'FIT' in Memory Using Coarsening
von: Roy, Shubhajit, et al.
Veröffentlicht: (2024)
von: Roy, Shubhajit, et al.
Veröffentlicht: (2024)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
Memory-Efficient Partitioned DNN Inference on Resource-Constrained Android Crowds
von: Manamperi, Lakshani, et al.
Veröffentlicht: (2026)
von: Manamperi, Lakshani, et al.
Veröffentlicht: (2026)
Constrained Online Convex Optimization with Memory and Predictions
von: Abdullah, Mohammed, et al.
Veröffentlicht: (2026)
von: Abdullah, Mohammed, et al.
Veröffentlicht: (2026)
STEER: Inference-Time Risk Control via Constrained Quality-Diversity Search
von: Yang, Eric, et al.
Veröffentlicht: (2026)
von: Yang, Eric, et al.
Veröffentlicht: (2026)
C-Learner: Constrained Learning for Causal Inference
von: Cai, Tiffany Tianhui, et al.
Veröffentlicht: (2024)
von: Cai, Tiffany Tianhui, et al.
Veröffentlicht: (2024)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
von: Han, Yuning, et al.
Veröffentlicht: (2026)
von: Han, Yuning, et al.
Veröffentlicht: (2026)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
von: Li, Weizhuo, et al.
Veröffentlicht: (2024)
von: Li, Weizhuo, et al.
Veröffentlicht: (2024)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
Memory by Design: Probabilistic Sequence Layers
von: Dowling, Matthew, et al.
Veröffentlicht: (2026)
von: Dowling, Matthew, et al.
Veröffentlicht: (2026)
Feasible-First Exploration for Constrained ML Deployment Optimization in Crash-Prone Hierarchical Search Spaces
von: Lysenstøen, Christian
Veröffentlicht: (2026)
von: Lysenstøen, Christian
Veröffentlicht: (2026)
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
von: Baroni, Luca, et al.
Veröffentlicht: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
von: Knoop, Jonathan, et al.
Veröffentlicht: (2026)
MicroNAS: Memory and Latency Constrained Hardware-Aware Neural Architecture Search for Time Series Classification on Microcontrollers
von: King, Tobias, et al.
Veröffentlicht: (2023)
von: King, Tobias, et al.
Veröffentlicht: (2023)
Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning
von: Wang, Peihao, et al.
Veröffentlicht: (2026)
von: Wang, Peihao, et al.
Veröffentlicht: (2026)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation
von: Li, Jiayi, et al.
Veröffentlicht: (2026)
von: Li, Jiayi, et al.
Veröffentlicht: (2026)
LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
Fast Heterogeneous Serving: Scalable Mixed-Scale LLM Allocation for SLO-Constrained Inference
von: Cheng, Jiaming, et al.
Veröffentlicht: (2026)
von: Cheng, Jiaming, et al.
Veröffentlicht: (2026)
SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
von: Tang, Yinghao, et al.
Veröffentlicht: (2025)
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
von: Ye, Zicong, et al.
Veröffentlicht: (2025)
von: Ye, Zicong, et al.
Veröffentlicht: (2025)
LLM-Driven Treatment Effect Estimation Under Inference Time Text Confounding
von: Ma, Yuchen, et al.
Veröffentlicht: (2025)
von: Ma, Yuchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Temporal Abstraction in Reinforcement Learning with Offline Data
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2024) -
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
von: Sander, Jacob, et al.
Veröffentlicht: (2025) -
Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents
von: Borro, Luiz C., et al.
Veröffentlicht: (2026) -
MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices
von: Wang, Zhaode, et al.
Veröffentlicht: (2025) -
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
von: Zhou, Jinan, et al.
Veröffentlicht: (2025)