PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ning, Rui, Zhang, Wei, Lai, Fan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
von: Li, Siyuan, et al.
Veröffentlicht: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
von: Recasens, Pol G., et al.
Veröffentlicht: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
von: Tian, Jian, et al.
Veröffentlicht: (2025)
von: Tian, Jian, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
InferCept: Efficient Intercept Support for Augmented Large Language Model Inference
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024)
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
von: Butler, Branden, et al.
Veröffentlicht: (2024)
von: Butler, Branden, et al.
Veröffentlicht: (2024)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
von: Luan, Frank Sifei, et al.
Veröffentlicht: (2025)
von: Luan, Frank Sifei, et al.
Veröffentlicht: (2025)
DiSCo: Device-Server Collaborative LLM-Based Text Streaming Services
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
von: Yu, Jiahuan, et al.
Veröffentlicht: (2026)
von: Yu, Jiahuan, et al.
Veröffentlicht: (2026)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
von: Zheng, Zhen, et al.
Veröffentlicht: (2024)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
von: Chen, Hongyao, et al.
Veröffentlicht: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
von: Wei, Xinming, et al.
Veröffentlicht: (2025)
MoE-Gen: High-Throughput MoE Inference on a Single GPU with Module-Based Batching
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
von: Xu, Tairan, et al.
Veröffentlicht: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
von: Gao, Lei, et al.
Veröffentlicht: (2024)
von: Gao, Lei, et al.
Veröffentlicht: (2024)
Towards Integrated Fine-tuning and Inference when Generative AI meets Edge Intelligence
von: Chen, Ning, et al.
Veröffentlicht: (2024)
von: Chen, Ning, et al.
Veröffentlicht: (2024)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
Multi-Bin Batching for Increasing LLM Inference Throughput
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
Making MoE-based LLM Inference Resilient with Tarragon
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
von: Zhang, Songyu, et al.
Veröffentlicht: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
von: Prabhakar, Rohan Baskar, et al.
Veröffentlicht: (2024)
von: Prabhakar, Rohan Baskar, et al.
Veröffentlicht: (2024)
Pie: Pooling CPU Memory for LLM Inference
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
STAR: Decode-Phase Rescheduling for LLM Inference
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services
von: Liu, Jiachen, et al.
Veröffentlicht: (2024)
von: Liu, Jiachen, et al.
Veröffentlicht: (2024)
Collaborative Batch Size Optimization for Federated Learning
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
von: Li, Siyuan, et al.
Veröffentlicht: (2024) -
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025) -
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
von: Recasens, Pol G., et al.
Veröffentlicht: (2025) -
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
von: Barrak, Amine, et al.
Veröffentlicht: (2025) -
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)