Characterize LSM-tree Compaction Performance via On-Device LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Ding, Jiabiao, Lv, Yina, Li, Qiao, Shen, Zhirong, Xue, Chun Jason |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking LSM-tree based Key-Value Stores: A Survey
di: Lv, Yina, et al.
Pubblicazione: (2025)
di: Lv, Yina, et al.
Pubblicazione: (2025)
Performance Characterization of Expert Router for Scalable LLM Inference
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs
di: Yu, Dingcui, et al.
Pubblicazione: (2025)
di: Yu, Dingcui, et al.
Pubblicazione: (2025)
Dissecting Embedding Bag Performance in DLRM Inference
di: Ambati, Chandrish, et al.
Pubblicazione: (2025)
di: Ambati, Chandrish, et al.
Pubblicazione: (2025)
Performance Characterization of Containers in Edge Computing
di: Gupta, Ragini, et al.
Pubblicazione: (2025)
di: Gupta, Ragini, et al.
Pubblicazione: (2025)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
di: Huang, Zhengxiang, et al.
Pubblicazione: (2025)
di: Huang, Zhengxiang, et al.
Pubblicazione: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
Performance Characterization and Optimizations of Traditional ML Applications
di: Kumar, Harsh, et al.
Pubblicazione: (2024)
di: Kumar, Harsh, et al.
Pubblicazione: (2024)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
di: Patwari, Rajeev, et al.
Pubblicazione: (2025)
di: Patwari, Rajeev, et al.
Pubblicazione: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
di: Salaria, Shweta, et al.
Pubblicazione: (2025)
di: Salaria, Shweta, et al.
Pubblicazione: (2025)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
di: Shin, Jiho, et al.
Pubblicazione: (2024)
di: Shin, Jiho, et al.
Pubblicazione: (2024)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
di: Shi, Tianyao, et al.
Pubblicazione: (2025)
di: Shi, Tianyao, et al.
Pubblicazione: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
di: Zhang, Niansong, et al.
Pubblicazione: (2025)
di: Zhang, Niansong, et al.
Pubblicazione: (2025)
Explainable Port Mapping Inference with Sparse Performance Counters for AMD's Zen Architectures
di: Ritter, Fabian, et al.
Pubblicazione: (2024)
di: Ritter, Fabian, et al.
Pubblicazione: (2024)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
di: Xu, Guanyu, et al.
Pubblicazione: (2025)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
di: Wang, Can, et al.
Pubblicazione: (2024)
di: Wang, Can, et al.
Pubblicazione: (2024)
Characterization of Photovoltaic Performance through Current - Voltage Analysis
di: Priyalatha Alexander, et al.
Pubblicazione: (2026)
di: Priyalatha Alexander, et al.
Pubblicazione: (2026)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
di: Javat, Abdurrahman, et al.
Pubblicazione: (2026)
di: Javat, Abdurrahman, et al.
Pubblicazione: (2026)
Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies
di: Kim, Kyoungmin, et al.
Pubblicazione: (2024)
di: Kim, Kyoungmin, et al.
Pubblicazione: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
di: Jiang, Jevin, et al.
Pubblicazione: (2026)
di: Jiang, Jevin, et al.
Pubblicazione: (2026)
Application Research On Real-Time Perception Of Device Performance Status
di: Wang, Zhe, et al.
Pubblicazione: (2024)
di: Wang, Zhe, et al.
Pubblicazione: (2024)
An Inquiry into Datacenter TCO for LLM Inference with FP8
di: Kim, Jiwoo, et al.
Pubblicazione: (2025)
di: Kim, Jiwoo, et al.
Pubblicazione: (2025)
Forecasting GPU Performance for Deep Learning Training and Inference
di: Lee, Seonho, et al.
Pubblicazione: (2024)
di: Lee, Seonho, et al.
Pubblicazione: (2024)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
di: Chai, Yuji, et al.
Pubblicazione: (2025)
di: Chai, Yuji, et al.
Pubblicazione: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
di: Chrapek, Marcin, et al.
Pubblicazione: (2025)
di: Chrapek, Marcin, et al.
Pubblicazione: (2025)
A Study on Inference Latency for Vision Transformers on Mobile Devices
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
di: Li, Zhuojin, et al.
Pubblicazione: (2025)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
di: Cavagna, Hiari Pizzini, et al.
Pubblicazione: (2026)
di: Cavagna, Hiari Pizzini, et al.
Pubblicazione: (2026)
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
di: Wang, Yulei, et al.
Pubblicazione: (2025)
di: Wang, Yulei, et al.
Pubblicazione: (2025)
Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices
di: Yan, Xiao, et al.
Pubblicazione: (2025)
di: Yan, Xiao, et al.
Pubblicazione: (2025)
Performance Characterization of AutoNUMA Memory Tiering on Graph Analytics
di: Moura, Diego, et al.
Pubblicazione: (2022)
di: Moura, Diego, et al.
Pubblicazione: (2022)
Documenti analoghi
-
Rethinking LSM-tree based Key-Value Stores: A Survey
di: Lv, Yina, et al.
Pubblicazione: (2025) -
Performance Characterization of Expert Router for Scalable LLM Inference
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024) -
Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs
di: Yu, Dingcui, et al.
Pubblicazione: (2025) -
Dissecting Embedding Bag Performance in DLRM Inference
di: Ambati, Chandrish, et al.
Pubblicazione: (2025) -
Performance Characterization of Containers in Edge Computing
di: Gupta, Ragini, et al.
Pubblicazione: (2025)