On the Cost of Model-Serving Frameworks: An Experimental Evaluation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | De Rosa, Pasquale, Bromberg, Yérom-David, Felber, Pascal, Mvondo, Djob, Schiavoni, Valerio |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Practical Forecasting of Cryptocoins Timeseries using Correlation Patterns
par: De Rosa, Pasquale, et autres
Publié: (2024)
par: De Rosa, Pasquale, et autres
Publié: (2024)
PhishingHook: Catching Phishing Ethereum Smart Contracts leveraging EVM Opcodes
par: De Rosa, Pasquale, et autres
Publié: (2025)
par: De Rosa, Pasquale, et autres
Publié: (2025)
ScamDetect: Towards a Robust, Agnostic Framework to Uncover Threats in Smart Contracts
par: De Rosa, Pasquale, et autres
Publié: (2025)
par: De Rosa, Pasquale, et autres
Publié: (2025)
Plinius: Secure and Persistent Machine Learning Model Training
par: Yuhala, Peterson, et autres
Publié: (2021)
par: Yuhala, Peterson, et autres
Publié: (2021)
CryptoAnalytics: Cryptocoins Price Forecasting with Machine Learning Techniques
par: De Rosa, Pasquale, et autres
Publié: (2024)
par: De Rosa, Pasquale, et autres
Publié: (2024)
BlindexTEE: A Blind Index Approach towards TEE-supported End-to-end Encrypted DBMS
par: Vialar, Louis, et autres
Publié: (2024)
par: Vialar, Louis, et autres
Publié: (2024)
IM-PIR: In-Memory Private Information Retrieval
par: Mwaisela, Mpoki, et autres
Publié: (2025)
par: Mwaisela, Mpoki, et autres
Publié: (2025)
PIM-CACHE: High-Efficiency Content-Aware Copy for Processing-In-Memory
par: Yuhala, Peterson, et autres
Publié: (2026)
par: Yuhala, Peterson, et autres
Publié: (2026)
Evaluating the Potential of In-Memory Processing to Accelerate Homomorphic Encryption
par: Mwaisela, Mpoki, et autres
Publié: (2024)
par: Mwaisela, Mpoki, et autres
Publié: (2024)
Continuous Semantic Caching for Low-Cost LLM Serving
par: Atalar, Baran, et autres
Publié: (2026)
par: Atalar, Baran, et autres
Publié: (2026)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
par: Gao, Bin, et autres
Publié: (2024)
par: Gao, Bin, et autres
Publié: (2024)
Partition Detection in Byzantine Networks
par: Bromberg, Yérom-David, et autres
Publié: (2024)
par: Bromberg, Yérom-David, et autres
Publié: (2024)
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
par: Liu, Xutong, et autres
Publié: (2025)
par: Liu, Xutong, et autres
Publié: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
par: Griggs, Tyler, et autres
Publié: (2024)
par: Griggs, Tyler, et autres
Publié: (2024)
HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
par: Xue, Leyang, et autres
Publié: (2025)
par: Xue, Leyang, et autres
Publié: (2025)
Cost-Optimal Active AI Model Evaluation
par: Angelopoulos, Anastasios N., et autres
Publié: (2025)
par: Angelopoulos, Anastasios N., et autres
Publié: (2025)
Model Equality Testing: Which Model Is This API Serving?
par: Gao, Irena, et autres
Publié: (2024)
par: Gao, Irena, et autres
Publié: (2024)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
par: Lee, Heejun, et autres
Publié: (2024)
par: Lee, Heejun, et autres
Publié: (2024)
TriHaRd: Higher Resilience for TEE Trusted Time
par: Bettinger, Matthieu, et autres
Publié: (2025)
par: Bettinger, Matthieu, et autres
Publié: (2025)
Practical Secure Aggregation by Combining Cryptography and Trusted Execution Environments
par: de Laage, Romain, et autres
Publié: (2025)
par: de Laage, Romain, et autres
Publié: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
par: Kossmann, Ferdi, et autres
Publié: (2024)
par: Kossmann, Ferdi, et autres
Publié: (2024)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
par: Zeng, Pai, et autres
Publié: (2024)
par: Zeng, Pai, et autres
Publié: (2024)
TorchAO: PyTorch-Native Training-to-Serving Model Optimization
par: Or, Andrew, et autres
Publié: (2025)
par: Or, Andrew, et autres
Publié: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
par: Kim, Kihyun, et autres
Publié: (2025)
par: Kim, Kihyun, et autres
Publié: (2025)
IC-Cache: Efficient Large Language Model Serving via In-context Caching
par: Yu, Yifan, et autres
Publié: (2025)
par: Yu, Yifan, et autres
Publié: (2025)
Defining 'Good': Evaluation Framework for Synthetic Smart Meter Data
par: Chai, Sheng, et autres
Publié: (2024)
par: Chai, Sheng, et autres
Publié: (2024)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
par: Gao, Lei, et autres
Publié: (2025)
par: Gao, Lei, et autres
Publié: (2025)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
par: Gao, Shihong, et autres
Publié: (2025)
par: Gao, Shihong, et autres
Publié: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
par: Zhao, Adrian, et autres
Publié: (2026)
par: Zhao, Adrian, et autres
Publié: (2026)
EdgeServe: A Streaming System for Decentralized Model Serving
par: Shaowang, Ted, et autres
Publié: (2023)
par: Shaowang, Ted, et autres
Publié: (2023)
iServe: An Intent-based Serving System for LLMs
par: Liakopoulos, Dimitrios, et autres
Publié: (2025)
par: Liakopoulos, Dimitrios, et autres
Publié: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
par: Kong, Linghao, et autres
Publié: (2026)
par: Kong, Linghao, et autres
Publié: (2026)
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving
par: Wang, Minghe, et autres
Publié: (2026)
par: Wang, Minghe, et autres
Publié: (2026)
Conditional computation in neural networks: principles and research trends
par: Scardapane, Simone, et autres
Publié: (2024)
par: Scardapane, Simone, et autres
Publié: (2024)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
par: Wang, Ganghua, et autres
Publié: (2025)
par: Wang, Ganghua, et autres
Publié: (2025)
Distributional Regression with Tabular Foundation Models: Evaluating Probabilistic Predictions via Proper Scoring Rules
par: Landsgesell, Jonas, et autres
Publié: (2026)
par: Landsgesell, Jonas, et autres
Publié: (2026)
Fairness in Serving Large Language Models
par: Sheng, Ying, et autres
Publié: (2023)
par: Sheng, Ying, et autres
Publié: (2023)
A Framework for Bounding Deterministic Risk with PAC-Bayes: Applications to Majority Votes
par: Leblanc, Benjamin, et autres
Publié: (2025)
par: Leblanc, Benjamin, et autres
Publié: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
par: Yu, Shan, et autres
Publié: (2025)
par: Yu, Shan, et autres
Publié: (2025)
Instance-Level Costs for Nuanced Classifier Evaluation
par: Kang, Kabir, et autres
Publié: (2026)
par: Kang, Kabir, et autres
Publié: (2026)
Documents similaires
-
Practical Forecasting of Cryptocoins Timeseries using Correlation Patterns
par: De Rosa, Pasquale, et autres
Publié: (2024) -
PhishingHook: Catching Phishing Ethereum Smart Contracts leveraging EVM Opcodes
par: De Rosa, Pasquale, et autres
Publié: (2025) -
ScamDetect: Towards a Robust, Agnostic Framework to Uncover Threats in Smart Contracts
par: De Rosa, Pasquale, et autres
Publié: (2025) -
Plinius: Secure and Persistent Machine Learning Model Training
par: Yuhala, Peterson, et autres
Publié: (2021) -
CryptoAnalytics: Cryptocoins Price Forecasting with Machine Learning Techniques
par: De Rosa, Pasquale, et autres
Publié: (2024)