VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Zihan, Luo, Xinhao, Guo, Junxian, Ni, Wentao, Zhou, Yangjie, Guan, Yue, Guo, Cong, Cui, Weihao, Feng, Yu, Guo, Minyi, Zhu, Yuhao, Zhang, Minjia, Leng, Jingwen, Jin, Chen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
por: Luo, Xinhao, et al.
Publicado: (2025)
por: Luo, Xinhao, et al.
Publicado: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
por: Guo, Cong, et al.
Publicado: (2024)
por: Guo, Cong, et al.
Publicado: (2024)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2025)
por: Xu, Jiale, et al.
Publicado: (2025)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
por: Huang, Ziyu, et al.
Publicado: (2025)
por: Huang, Ziyu, et al.
Publicado: (2025)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
por: Qiang, Xinwei, et al.
Publicado: (2026)
por: Qiang, Xinwei, et al.
Publicado: (2026)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2024)
por: Xu, Jiale, et al.
Publicado: (2024)
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
por: Zhou, Yangjie, et al.
Publicado: (2024)
por: Zhou, Yangjie, et al.
Publicado: (2024)
MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
por: Xue, Chunyu, et al.
Publicado: (2026)
por: Xue, Chunyu, et al.
Publicado: (2026)
Towards Fast Setup and High Throughput of GPU Serverless Computing
por: Zhao, Han, et al.
Publicado: (2024)
por: Zhao, Han, et al.
Publicado: (2024)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
por: Xu, Chuhao, et al.
Publicado: (2025)
por: Xu, Chuhao, et al.
Publicado: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
por: Guo, Cong, et al.
Publicado: (2024)
por: Guo, Cong, et al.
Publicado: (2024)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
por: Xu, Ao, et al.
Publicado: (2025)
por: Xu, Ao, et al.
Publicado: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
por: Wang, Weiye, et al.
Publicado: (2026)
por: Wang, Weiye, et al.
Publicado: (2026)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026)
por: Du, Boxiao, et al.
Publicado: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
por: Gan, Zhenghao, et al.
Publicado: (2026)
por: Gan, Zhenghao, et al.
Publicado: (2026)
Enabling Dynamic Sparsity in Quantized LLM Inference
por: Wang, Rongxiang, et al.
Publicado: (2025)
por: Wang, Rongxiang, et al.
Publicado: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
por: Chen, Hongyu, et al.
Publicado: (2026)
por: Chen, Hongyu, et al.
Publicado: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
por: Liu, Di, et al.
Publicado: (2026)
por: Liu, Di, et al.
Publicado: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
por: Lin, Shouxu, et al.
Publicado: (2026)
por: Lin, Shouxu, et al.
Publicado: (2026)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
por: Yu, Jiahuan, et al.
Publicado: (2026)
por: Yu, Jiahuan, et al.
Publicado: (2026)
Federated Inference for Heterogeneous LLM Communication and Collaboration
por: Chen, Zihan, et al.
Publicado: (2026)
por: Chen, Zihan, et al.
Publicado: (2026)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
por: Yang, Zheming, et al.
Publicado: (2024)
por: Yang, Zheming, et al.
Publicado: (2024)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
por: Chen, Jinyuan, et al.
Publicado: (2025)
por: Chen, Jinyuan, et al.
Publicado: (2025)
Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
por: Chen, Ping, et al.
Publicado: (2025)
por: Chen, Ping, et al.
Publicado: (2025)
VoltanaLLM: Feedback-Driven Frequency Control and State-Space Routing for Energy-Efficient LLM Serving
por: Yu, Jiahuan, et al.
Publicado: (2025)
por: Yu, Jiahuan, et al.
Publicado: (2025)
MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
por: Wang, Zongwu, et al.
Publicado: (2025)
por: Wang, Zongwu, et al.
Publicado: (2025)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
por: Lian, Xinyu, et al.
Publicado: (2025)
por: Lian, Xinyu, et al.
Publicado: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
por: Zhang, Yaozheng, et al.
Publicado: (2025)
por: Zhang, Yaozheng, et al.
Publicado: (2025)
Cicero: Addressing Algorithmic and Architectural Bottlenecks in Neural Rendering by Radiance Warping and Memory Optimizations
por: Feng, Yu, et al.
Publicado: (2024)
por: Feng, Yu, et al.
Publicado: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
por: Chen, Haoyu, et al.
Publicado: (2025)
por: Chen, Haoyu, et al.
Publicado: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2026)
por: Yang, Zheming, et al.
Publicado: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
por: Zhang, Mingjin, et al.
Publicado: (2024)
por: Zhang, Mingjin, et al.
Publicado: (2024)
Justitia: Fair and Efficient Scheduling of Task-parallel LLM Agents with Selective Pampering
por: Yang, Mingyan, et al.
Publicado: (2025)
por: Yang, Mingyan, et al.
Publicado: (2025)
FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
por: Mei, Junyi, et al.
Publicado: (2024)
por: Mei, Junyi, et al.
Publicado: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
por: Oliaro, Gabriele, et al.
Publicado: (2024)
por: Oliaro, Gabriele, et al.
Publicado: (2024)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
por: Liu, Di, et al.
Publicado: (2026)
por: Liu, Di, et al.
Publicado: (2026)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
por: Guo, Tianyu, et al.
Publicado: (2025)
por: Guo, Tianyu, et al.
Publicado: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
por: He, Wenhao, et al.
Publicado: (2026)
por: He, Wenhao, et al.
Publicado: (2026)
Ejemplares similares
-
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
por: Luo, Xinhao, et al.
Publicado: (2025) -
Accelerating Sparse DNNs Based on Tiled GEMM
por: Guo, Cong, et al.
Publicado: (2024) -
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2025) -
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
por: Huang, Ziyu, et al.
Publicado: (2025) -
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
por: Qiang, Xinwei, et al.
Publicado: (2026)