Towards Pareto Optimal Throughput in Small Language Model Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Recasens, Pol G., Zhu, Yue, Wang, Chen, Lee, Eun Kyung, Tardieu, Olivier, Youssef, Alaa, Torres, Jordi, Berral, Josep Ll. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2025)
von: Agullo, Ferran, et al.
Veröffentlicht: (2025)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
von: Agullo, Ferran, et al.
Veröffentlicht: (2026)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
FRIDA: Free-Rider Detection using Privacy Attacks
von: Recasens, Pol G., et al.
Veröffentlicht: (2024)
von: Recasens, Pol G., et al.
Veröffentlicht: (2024)
In-Context Bias Propagation in LLM-Based Tabular Data Generation
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
Enabling an OpenStack-based cloud on top of RISC-V hardware
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
von: Marrón, Diego, et al.
Veröffentlicht: (2024)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
ProST: Progressive Sub-task Training for Pareto-Optimal Multi-agent Systems Using Small Language Models
von: Bijoy, Biddut Sarker, et al.
Veröffentlicht: (2025)
von: Bijoy, Biddut Sarker, et al.
Veröffentlicht: (2025)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
Pareto Optimal Code Generation
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2025)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2025)
Pareto Optimal Learning for Estimating Large Language Model Errors
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
von: Li, Moxin, et al.
Veröffentlicht: (2025)
von: Li, Moxin, et al.
Veröffentlicht: (2025)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
von: Shen, Haiying, et al.
Veröffentlicht: (2025)
SOMA: Efficient Multi-turn LLM Serving via Small Language Model
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
RULE: Reinforcement UnLEarning Achieves Forget-Retain Pareto Optimality
von: Zhang, Chenlong, et al.
Veröffentlicht: (2025)
von: Zhang, Chenlong, et al.
Veröffentlicht: (2025)
LegiGPT: Party Politics and Transport Policy with Large Language Model
von: Yun, Hyunsoo, et al.
Veröffentlicht: (2025)
von: Yun, Hyunsoo, et al.
Veröffentlicht: (2025)
Towards Optimal Learning of Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
Small Language Models are Equation Reasoners
von: Kim, Bumjun, et al.
Veröffentlicht: (2024)
von: Kim, Bumjun, et al.
Veröffentlicht: (2024)
Towards Resiliency in Large Language Model Serving with KevlarFlow
von: Qian, Shangshu, et al.
Veröffentlicht: (2026)
von: Qian, Shangshu, et al.
Veröffentlicht: (2026)
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
von: Wang, Ying, et al.
Veröffentlicht: (2025)
von: Wang, Ying, et al.
Veröffentlicht: (2025)
UC-MOA: Utility-Conditioned Multi-Objective Alignment for Distributional Pareto-Optimality
von: Cheng, Zelei, et al.
Veröffentlicht: (2025)
von: Cheng, Zelei, et al.
Veröffentlicht: (2025)
ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving
von: Taberner-Miller, Annette
Veröffentlicht: (2026)
von: Taberner-Miller, Annette
Veröffentlicht: (2026)
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
ParetoHqD: Fast Offline Multiobjective Alignment of Large Language Models using Pareto High-quality Data
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
Pareto Multi-Objective Alignment for Language Models
von: He, Qiang, et al.
Veröffentlicht: (2025)
von: He, Qiang, et al.
Veröffentlicht: (2025)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
von: Elobaid, Alaa
Veröffentlicht: (2026)
von: Elobaid, Alaa
Veröffentlicht: (2026)
Towards Optimal Statistical Watermarking
von: Huang, Baihe, et al.
Veröffentlicht: (2023)
von: Huang, Baihe, et al.
Veröffentlicht: (2023)
1+1>2: Can Large Language Models Serve as Cross-Lingual Knowledge Aggregators?
von: Huang, Yue, et al.
Veröffentlicht: (2024)
von: Huang, Yue, et al.
Veröffentlicht: (2024)
Towards a World-English Language Model for On-Device Virtual Assistants
von: Jalota, Rricha, et al.
Veröffentlicht: (2024)
von: Jalota, Rricha, et al.
Veröffentlicht: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
Accelerating Retrieval-Augmented Language Model Serving with Speculation
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
Tending Towards Stability: Convergence Challenges in Small Language Models
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2024)
von: Martinez, Richard Diehl, et al.
Veröffentlicht: (2024)
Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models
von: Prucs, Ákos, et al.
Veröffentlicht: (2025)
von: Prucs, Ákos, et al.
Veröffentlicht: (2025)
PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models
von: Han, Ji-Eun, et al.
Veröffentlicht: (2024)
von: Han, Ji-Eun, et al.
Veröffentlicht: (2024)
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
von: Lee, Jonghyun, et al.
Veröffentlicht: (2025)
von: Lee, Jonghyun, et al.
Veröffentlicht: (2025)
SLaDe: A Portable Small Language Model Decompiler for Optimized Assembly
von: Armengol-Estapé, Jordi, et al.
Veröffentlicht: (2023)
von: Armengol-Estapé, Jordi, et al.
Veröffentlicht: (2023)
Pareto-optimal Non-uniform Language Generation
von: Charikar, Moses, et al.
Veröffentlicht: (2025)
von: Charikar, Moses, et al.
Veröffentlicht: (2025)
A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles
von: Lee, Eun-Kyoung Rosa, et al.
Veröffentlicht: (2024)
von: Lee, Eun-Kyoung Rosa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2025) -
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
von: Agullo, Ferran, et al.
Veröffentlicht: (2026) -
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025) -
FRIDA: Free-Rider Detection using Privacy Attacks
von: Recasens, Pol G., et al.
Veröffentlicht: (2024) -
In-Context Bias Propagation in LLM-Based Tabular Data Generation
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)