Reasoning Language Model Inference Serving Unveiled: An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qi, Wu, Junpan, Liu, Xiang, Wang, Yuxin, Li, Zeyu, Tang, Zhenheng, Chen, Yuhan, Shi, Shaohuai, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
Should We Really Edit Language Models? On the Evaluation of Edited Language Models
by: Li, Qi, et al.
Published: (2024)
by: Li, Qi, et al.
Published: (2024)
FedImpro: Measuring and Improving Client Update in Federated Learning
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
by: Tang, Zhenheng, et al.
Published: (2026)
by: Tang, Zhenheng, et al.
Published: (2026)
Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems
by: Lin, Wenxiang, et al.
Published: (2024)
by: Lin, Wenxiang, et al.
Published: (2024)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025)
by: Dong, Peijie, et al.
Published: (2025)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
by: Tang, Zichen, et al.
Published: (2024)
by: Tang, Zichen, et al.
Published: (2024)
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
OmniReview: A Large-scale Benchmark and LLM-enhanced Framework for Realistic Reviewer Recommendation
by: Huang, Yehua, et al.
Published: (2026)
by: Huang, Yehua, et al.
Published: (2026)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
by: Zhang, Longteng, et al.
Published: (2026)
by: Zhang, Longteng, et al.
Published: (2026)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
by: Lai, Kunfeng, et al.
Published: (2025)
by: Lai, Kunfeng, et al.
Published: (2025)
LPZero: Language Model Zero-cost Proxy Search from Zero
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
by: Li, Zeyu, et al.
Published: (2025)
by: Li, Zeyu, et al.
Published: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
by: He, Xin, et al.
Published: (2025)
by: He, Xin, et al.
Published: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
by: Bai, Fan, et al.
Published: (2026)
by: Bai, Fan, et al.
Published: (2026)
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
by: Yu, Sheldon, et al.
Published: (2025)
by: Yu, Sheldon, et al.
Published: (2025)
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
by: Liu, Xiang, et al.
Published: (2026)
by: Liu, Xiang, et al.
Published: (2026)
LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning
by: Zhang, Longteng, et al.
Published: (2023)
by: Zhang, Longteng, et al.
Published: (2023)
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
by: Wu, Yangzhen, et al.
Published: (2024)
by: Wu, Yangzhen, et al.
Published: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
An Empirical Study on Prompt Compression for Large Language Models
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
by: Pan, Xinglin, et al.
Published: (2024)
by: Pan, Xinglin, et al.
Published: (2024)
PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model Decoupling
by: Wang, Sijie, et al.
Published: (2025)
by: Wang, Sijie, et al.
Published: (2025)
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Reason Analogically via Cross-domain Prior Knowledge: An Empirical Study of Cross-domain Knowledge Transfer for In-Context Learning
by: Liu, Le, et al.
Published: (2026)
by: Liu, Le, et al.
Published: (2026)
Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
by: Hu, Zhengyu, et al.
Published: (2025)
by: Hu, Zhengyu, et al.
Published: (2025)
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification
by: Sanyal, Soumya, et al.
Published: (2024)
by: Sanyal, Soumya, et al.
Published: (2024)
Unveiling Hidden Collaboration within Mixture-of-Experts in Large Language Models
by: Tang, Yuanbo, et al.
Published: (2025)
by: Tang, Yuanbo, et al.
Published: (2025)
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
by: Chi, Haoang, et al.
Published: (2025)
by: Chi, Haoang, et al.
Published: (2025)
Unveiling Confirmation Bias in Chain-of-Thought Reasoning
by: Wan, Yue, et al.
Published: (2025)
by: Wan, Yue, et al.
Published: (2025)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
by: Chen, Kaiwen, et al.
Published: (2025)
by: Chen, Kaiwen, et al.
Published: (2025)
Similar Items
-
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025) -
Should We Really Edit Language Models? On the Evaluation of Edited Language Models
by: Li, Qi, et al.
Published: (2024) -
FedImpro: Measuring and Improving Client Update in Federated Learning
by: Tang, Zhenheng, et al.
Published: (2024) -
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
by: Liu, Xiang, et al.
Published: (2025) -
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
by: He, Xin, et al.
Published: (2024)