SGLang: Efficient Execution of Structured Language Model Programs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Lianmin, Yin, Liangsheng, Xie, Zhiqiang, Sun, Chuyue, Huang, Jeff, Yu, Cody Hao, Cao, Shiyi, Kozyrakis, Christos, Stoica, Ion, Gonzalez, Joseph E., Barrett, Clark, Sheng, Ying |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
por: Xie, Zhiqiang, et al.
Publicado: (2024)
por: Xie, Zhiqiang, et al.
Publicado: (2024)
Post-Training Sparse Attention with Double Sparsity
por: Yang, Shuo, et al.
Publicado: (2024)
por: Yang, Shuo, et al.
Publicado: (2024)
Clover: Closed-Loop Verifiable Code Generation
por: Sun, Chuyue, et al.
Publicado: (2023)
por: Sun, Chuyue, et al.
Publicado: (2023)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
por: Cao, Shiyi, et al.
Publicado: (2026)
por: Cao, Shiyi, et al.
Publicado: (2026)
Teaching Cloud Infrastructure and Scalable Application Deployment in an Undergraduate Computer Science Program
por: Saligrama, Aditya, et al.
Publicado: (2024)
por: Saligrama, Aditya, et al.
Publicado: (2024)
FailSafe: High-performance Resilient Serving
por: Xu, Ziyi, et al.
Publicado: (2025)
por: Xu, Ziyi, et al.
Publicado: (2025)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
por: Sheng, Ying, et al.
Publicado: (2023)
por: Sheng, Ying, et al.
Publicado: (2023)
Locality-aware Fair Scheduling in LLM Serving
por: Cao, Shiyi, et al.
Publicado: (2025)
por: Cao, Shiyi, et al.
Publicado: (2025)
Fairness in Serving Large Language Models
por: Sheng, Ying, et al.
Publicado: (2023)
por: Sheng, Ying, et al.
Publicado: (2023)
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
por: Xie, Zhiqiang, et al.
Publicado: (2024)
por: Xie, Zhiqiang, et al.
Publicado: (2024)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
por: Balaji, Vignesh, et al.
Publicado: (2025)
por: Balaji, Vignesh, et al.
Publicado: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
por: Jiang, Xuanlin, et al.
Publicado: (2024)
por: Jiang, Xuanlin, et al.
Publicado: (2024)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
por: Yu, Shan, et al.
Publicado: (2025)
por: Yu, Shan, et al.
Publicado: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
por: Gandhi, Swapnil, et al.
Publicado: (2024)
por: Gandhi, Swapnil, et al.
Publicado: (2024)
Some Present-Day Problems of Romanian Library Science
por: Stoica, Ion
Publicado: (1973)
por: Stoica, Ion
Publicado: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
por: Stoica, Ion
Publicado: (1972)
por: Stoica, Ion
Publicado: (1972)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
por: Park, Jongseok, et al.
Publicado: (2026)
por: Park, Jongseok, et al.
Publicado: (2026)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
por: Cao, Shiyi, et al.
Publicado: (2024)
por: Cao, Shiyi, et al.
Publicado: (2024)
Executive individualism and the tone of firms' annual reports
por: Wei Jiang, et al.
Publicado: (2024)
por: Wei Jiang, et al.
Publicado: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
por: Chiang, Wei-Lin, et al.
Publicado: (2024)
por: Chiang, Wei-Lin, et al.
Publicado: (2024)
On Optimizing the Communication of Model Parallelism
por: Zhuang, Yonghao, et al.
Publicado: (2022)
por: Zhuang, Yonghao, et al.
Publicado: (2022)
Pantograph: A Machine-to-Machine Interaction Interface for Advanced Theorem Proving, High Level Reasoning, and Data Extraction in Lean 4
por: Aniva, Leni, et al.
Publicado: (2024)
por: Aniva, Leni, et al.
Publicado: (2024)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
por: Zhao, Mark, et al.
Publicado: (2024)
por: Zhao, Mark, et al.
Publicado: (2024)
Hunting CUDA Bugs at Scale with cuFuzz
por: ziad, Mohamed Tarek Ibn, et al.
Publicado: (2026)
por: ziad, Mohamed Tarek Ibn, et al.
Publicado: (2026)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
por: Zhao, Yilong, et al.
Publicado: (2024)
por: Zhao, Yilong, et al.
Publicado: (2024)
S*: Test Time Scaling for Code Generation
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
por: Luan, Frank Sifei, et al.
Publicado: (2025)
por: Luan, Frank Sifei, et al.
Publicado: (2025)
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
por: Zheng, Lianmin, et al.
Publicado: (2023)
por: Zheng, Lianmin, et al.
Publicado: (2023)
VeriStruct: AI-assisted Automated Verification of Data-Structure Modules in Verus
por: Sun, Chuyue, et al.
Publicado: (2025)
por: Sun, Chuyue, et al.
Publicado: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
por: Xie, Zhiqiang, et al.
Publicado: (2025)
por: Xie, Zhiqiang, et al.
Publicado: (2025)
Efficient LLM Scheduling by Learning to Rank
por: Fu, Yichao, et al.
Publicado: (2024)
por: Fu, Yichao, et al.
Publicado: (2024)
How Fast Can I Run My VLA? Demystifying VLA Inference Performance with VLA-Perf
por: Jiang, Wenqi, et al.
Publicado: (2026)
por: Jiang, Wenqi, et al.
Publicado: (2026)
LIMINAL: Exploring The Frontiers of LLM Decode Performance
por: Davies, Michael, et al.
Publicado: (2025)
por: Davies, Michael, et al.
Publicado: (2025)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
por: Gandhi, Swapnil, et al.
Publicado: (2024)
por: Gandhi, Swapnil, et al.
Publicado: (2024)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
por: Fu, Yichao, et al.
Publicado: (2024)
por: Fu, Yichao, et al.
Publicado: (2024)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
por: Skiadopoulos, Athinagoras, et al.
Publicado: (2025)
por: Skiadopoulos, Athinagoras, et al.
Publicado: (2025)
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
por: Li, Dacheng, et al.
Publicado: (2025)
por: Li, Dacheng, et al.
Publicado: (2025)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
por: Ding, Hangliang, et al.
Publicado: (2025)
por: Ding, Hangliang, et al.
Publicado: (2025)
Autellix: An Efficient Serving Engine for LLM Agents as General Programs
por: Luo, Michael, et al.
Publicado: (2025)
por: Luo, Michael, et al.
Publicado: (2025)
Regulating Branch Parallelism in LLM Serving
por: Gandhi, Swapnil, et al.
Publicado: (2026)
por: Gandhi, Swapnil, et al.
Publicado: (2026)
Ejemplares similares
-
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
por: Xie, Zhiqiang, et al.
Publicado: (2024) -
Post-Training Sparse Attention with Double Sparsity
por: Yang, Shuo, et al.
Publicado: (2024) -
Clover: Closed-Loop Verifiable Code Generation
por: Sun, Chuyue, et al.
Publicado: (2023) -
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
por: Cao, Shiyi, et al.
Publicado: (2026) -
Teaching Cloud Infrastructure and Scalable Application Deployment in an Undergraduate Computer Science Program
por: Saligrama, Aditya, et al.
Publicado: (2024)