A Survey on Efficient Inference for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Zixuan, Ning, Xuefei, Hong, Ke, Fu, Tianyu, Xu, Jiaming, Li, Shiyao, Lou, Yuming, Wang, Luning, Yuan, Zhihang, Li, Xiuhong, Yan, Shengen, Dai, Guohao, Zhang, Xiao-Ping, Dong, Yuhan, Wang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
von: Wang, Luning, et al.
Veröffentlicht: (2024)
von: Wang, Luning, et al.
Veröffentlicht: (2024)
Evaluating Quantized Large Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
FlashDecoding++: Faster Large Language Model Inference on GPUs
von: Hong, Ke, et al.
Veröffentlicht: (2023)
von: Hong, Ke, et al.
Veröffentlicht: (2023)
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
von: Liu, Enshu, et al.
Veröffentlicht: (2024)
von: Liu, Enshu, et al.
Veröffentlicht: (2024)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
von: Li, Jinhao, et al.
Veröffentlicht: (2023)
von: Li, Jinhao, et al.
Veröffentlicht: (2023)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
von: Zeng, Shulin, et al.
Veröffentlicht: (2024)
Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
von: Liu, Enshu, et al.
Veröffentlicht: (2025)
von: Liu, Enshu, et al.
Veröffentlicht: (2025)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
von: Zhao, Tianchen, et al.
Veröffentlicht: (2024)
von: Zhao, Tianchen, et al.
Veröffentlicht: (2024)
TASP: Topology-aware Sequence Parallelism
von: Wang, Yida, et al.
Veröffentlicht: (2025)
von: Wang, Yida, et al.
Veröffentlicht: (2025)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
von: Ning, Xuefei, et al.
Veröffentlicht: (2024)
von: Ning, Xuefei, et al.
Veröffentlicht: (2024)
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K
von: Yuan, Tao, et al.
Veröffentlicht: (2024)
von: Yuan, Tao, et al.
Veröffentlicht: (2024)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
von: Zhao, Tianchen, et al.
Veröffentlicht: (2024)
von: Zhao, Tianchen, et al.
Veröffentlicht: (2024)
VGDFR: Diffusion-based Video Generation with Dynamic Latent Frame Rate
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
E-CAR: Efficient Continuous Autoregressive Image Generation via Multistage Modeling
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
von: Liu, Tengxuan, et al.
Veröffentlicht: (2025)
von: Liu, Tengxuan, et al.
Veröffentlicht: (2025)
SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
von: Zhang, Hanling, et al.
Veröffentlicht: (2025)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)
von: Peng, Yanxin, et al.
Veröffentlicht: (2025)
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
von: Hong, Ke, et al.
Veröffentlicht: (2025)
von: Hong, Ke, et al.
Veröffentlicht: (2025)
DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis
von: Teng, Yao, et al.
Veröffentlicht: (2024)
von: Teng, Yao, et al.
Veröffentlicht: (2024)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
von: Hong, Ke, et al.
Veröffentlicht: (2025)
von: Hong, Ke, et al.
Veröffentlicht: (2025)
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
von: Liu, Enshu, et al.
Veröffentlicht: (2024)
von: Liu, Enshu, et al.
Veröffentlicht: (2024)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
von: Ning, Xuefei, et al.
Veröffentlicht: (2023)
von: Ning, Xuefei, et al.
Veröffentlicht: (2023)
SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding
von: Teng, Yao, et al.
Veröffentlicht: (2024)
von: Teng, Yao, et al.
Veröffentlicht: (2024)
BalanceGS: Algorithm-System Co-design for Efficient 3D Gaussian Splatting Training on GPU
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
The association between Internet use and intrinsic capacity among older adults in China: The mediating role of social participation
von: Jianying Fu, et al.
Veröffentlicht: (2024)
von: Jianying Fu, et al.
Veröffentlicht: (2024)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization
von: Xie, Rui, et al.
Veröffentlicht: (2024)
von: Xie, Rui, et al.
Veröffentlicht: (2024)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
von: Duanmu, Haojie, et al.
Veröffentlicht: (2024)
von: Duanmu, Haojie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
von: Wang, Luning, et al.
Veröffentlicht: (2024) -
Evaluating Quantized Large Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024) -
FlashDecoding++: Faster Large Language Model Inference on GPUs
von: Hong, Ke, et al.
Veröffentlicht: (2023) -
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024) -
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)