Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Huamin, Liu, Xunzhuo, Jiang, Junchen, He, Bowei, Liu, Xue |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
par: Lyu, Hongtao, et autres
Publié: (2025)
par: Lyu, Hongtao, et autres
Publié: (2025)
SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
par: Shen, Yuanzhe, et autres
Publié: (2025)
par: Shen, Yuanzhe, et autres
Publié: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
par: Liu, Wentao, et autres
Publié: (2025)
par: Liu, Wentao, et autres
Publié: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
par: Liu, Xing, et autres
Publié: (2025)
par: Liu, Xing, et autres
Publié: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
par: Liu, Dong, et autres
Publié: (2025)
par: Liu, Dong, et autres
Publié: (2025)
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
par: Jiang, Bo, et autres
Publié: (2025)
par: Jiang, Bo, et autres
Publié: (2025)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
par: Jiang, Bo, et autres
Publié: (2025)
par: Jiang, Bo, et autres
Publié: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
par: The AIBrix Team, et autres
Publié: (2025)
par: The AIBrix Team, et autres
Publié: (2025)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
par: Kim, Joon Ha, et autres
Publié: (2026)
par: Kim, Joon Ha, et autres
Publié: (2026)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
par: Stojkovic, Jovan, et autres
Publié: (2025)
par: Stojkovic, Jovan, et autres
Publié: (2025)
LAPS: A Length-Aware-Prefill LLM Serving System
par: She, Jianshu, et autres
Publié: (2026)
par: She, Jianshu, et autres
Publié: (2026)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
par: Yuan, Yichao, et autres
Publié: (2026)
par: Yuan, Yichao, et autres
Publié: (2026)
Delay-Aware Multi-Stage Edge Server Upgrade with Budget Constraint
par: Wihidayat, Endar Suprih, et autres
Publié: (2025)
par: Wihidayat, Endar Suprih, et autres
Publié: (2025)
MergePipe: A Budget-Aware Parameter Management System for Scalable LLM Merging
par: Wang, Yuanyi, et autres
Publié: (2026)
par: Wang, Yuanyi, et autres
Publié: (2026)
LLM Inference Serving: Survey of Recent Advances and Opportunities
par: Li, Baolin, et autres
Publié: (2024)
par: Li, Baolin, et autres
Publié: (2024)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
par: Argerich, Mauricio Fadel, et autres
Publié: (2026)
par: Argerich, Mauricio Fadel, et autres
Publié: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
par: Li, Rongzhi, et autres
Publié: (2025)
par: Li, Rongzhi, et autres
Publié: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
par: Xie, Jincheng, et autres
Publié: (2026)
par: Xie, Jincheng, et autres
Publié: (2026)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
par: Yu, Donglin
Publié: (2026)
par: Yu, Donglin
Publié: (2026)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
par: Wu, Linyu, et autres
Publié: (2025)
par: Wu, Linyu, et autres
Publié: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
par: Zheng, Wanyi, et autres
Publié: (2025)
par: Zheng, Wanyi, et autres
Publié: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
par: Wang, Haodong, et autres
Publié: (2025)
par: Wang, Haodong, et autres
Publié: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
par: Wu, Panlong, et autres
Publié: (2025)
par: Wu, Panlong, et autres
Publié: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
par: Huang, Tao, et autres
Publié: (2024)
par: Huang, Tao, et autres
Publié: (2024)
Seesaw: High-throughput LLM Inference via Model Re-sharding
par: Su, Qidong, et autres
Publié: (2025)
par: Su, Qidong, et autres
Publié: (2025)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
par: Luo, Xinhao, et autres
Publié: (2025)
par: Luo, Xinhao, et autres
Publié: (2025)
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
par: Zhu, Yue, et autres
Publié: (2025)
par: Zhu, Yue, et autres
Publié: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
par: Li, Xiangyu, et autres
Publié: (2025)
par: Li, Xiangyu, et autres
Publié: (2025)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
par: Song, Jingwei, et autres
Publié: (2025)
par: Song, Jingwei, et autres
Publié: (2025)
Accelerating LLM Inference with Precomputed Query Storage
par: Park, Jay H., et autres
Publié: (2025)
par: Park, Jay H., et autres
Publié: (2025)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
par: Behera, Adarsh Prasad, et autres
Publié: (2025)
par: Behera, Adarsh Prasad, et autres
Publié: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
par: Ramachandran, Arun, et autres
Publié: (2025)
par: Ramachandran, Arun, et autres
Publié: (2025)
Artificial Intelligence for Cost-Aware Resource Prediction in Big Data Pipelines
par: Goyal, Harshit
Publié: (2025)
par: Goyal, Harshit
Publié: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
par: Xue, Zhenliang, et autres
Publié: (2025)
par: Xue, Zhenliang, et autres
Publié: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
par: Olshansky, Daniel, et autres
Publié: (2024)
par: Olshansky, Daniel, et autres
Publié: (2024)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
par: Liu, Man, et autres
Publié: (2026)
par: Liu, Man, et autres
Publié: (2026)
Documents similaires
-
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
par: Chen, Huamin, et autres
Publié: (2026) -
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
par: Chen, Huamin, et autres
Publié: (2026) -
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
par: Chen, Huamin, et autres
Publié: (2026) -
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
par: Chen, Huamin, et autres
Publié: (2026) -
FairBatching: Fairness-Aware Batch Formation for LLM Inference
par: Lyu, Hongtao, et autres
Publié: (2025)