FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zihao, Chen, Lequn, Lai, Ruihang, Lin, Wuwei, Zhang, Yineng, Wang, Stephanie, Chen, Tianqi, Kasikci, Baris, Grover, Vinod, Krishnamurthy, Arvind, Ceze, Luis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
von: Xing, Shanli, et al.
Veröffentlicht: (2026)
von: Xing, Shanli, et al.
Veröffentlicht: (2026)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
von: Zhao, Yilong, et al.
Veröffentlicht: (2023)
von: Zhao, Yilong, et al.
Veröffentlicht: (2023)
VoxServe: Streaming-Centric Serving System for Speech Language Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
von: Chen, Lequn, et al.
Veröffentlicht: (2023)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
von: Zhu, Kan, et al.
Veröffentlicht: (2024)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
von: Ning, Rui, et al.
Veröffentlicht: (2026)
von: Ning, Rui, et al.
Veröffentlicht: (2026)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
von: Ren, Feng, et al.
Veröffentlicht: (2026)
von: Ren, Feng, et al.
Veröffentlicht: (2026)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
von: Ye, Fanjiang, et al.
Veröffentlicht: (2026)
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
von: Licker, Nandor, et al.
Veröffentlicht: (2025)
von: Licker, Nandor, et al.
Veröffentlicht: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
von: Zarkadas, Ioannis, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Is Flash Attention Stable?
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
von: Golden, Alicia, et al.
Veröffentlicht: (2024)
A System for Microserving of LLMs
von: Jin, Hongyi, et al.
Veröffentlicht: (2024)
von: Jin, Hongyi, et al.
Veröffentlicht: (2024)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
von: Zhao, Yilong, et al.
Veröffentlicht: (2026)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
von: Pan, Yi, et al.
Veröffentlicht: (2025)
von: Pan, Yi, et al.
Veröffentlicht: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
Scaling Deep Learning Training with MPMD Pipeline Parallelism
von: Xhebraj, Anxhelo, et al.
Veröffentlicht: (2024)
von: Xhebraj, Anxhelo, et al.
Veröffentlicht: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
von: Pan, Yi, et al.
Veröffentlicht: (2026)
von: Pan, Yi, et al.
Veröffentlicht: (2026)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
Autellix: An Efficient Serving Engine for LLM Agents as General Programs
von: Luo, Michael, et al.
Veröffentlicht: (2025)
von: Luo, Michael, et al.
Veröffentlicht: (2025)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
von: Jin, Hongyi, et al.
Veröffentlicht: (2026)
von: Jin, Hongyi, et al.
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
von: Xing, Shanli, et al.
Veröffentlicht: (2026) -
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025) -
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
von: Zhao, Yilong, et al.
Veröffentlicht: (2023) -
VoxServe: Streaming-Centric Serving System for Speech Language Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026) -
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)