Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Tianyu, Huang, Haofeng, Ning, Xuefei, Zhang, Genghan, Chen, Boju, Wu, Tianqi, Wang, Hongyi, Huang, Zixiao, Li, Shiyao, Yan, Shengen, Dai, Guohao, Yang, Huazhong, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
by: Liu, Enshu, et al.
Published: (2024)
by: Liu, Enshu, et al.
Published: (2024)
Evaluating Quantized Large Language Models
by: Li, Shiyao, et al.
Published: (2024)
by: Li, Shiyao, et al.
Published: (2024)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
by: Wang, Luning, et al.
Published: (2024)
by: Wang, Luning, et al.
Published: (2024)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
by: Zeng, Shulin, et al.
Published: (2024)
by: Zeng, Shulin, et al.
Published: (2024)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
by: Ning, Xuefei, et al.
Published: (2024)
by: Ning, Xuefei, et al.
Published: (2024)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
by: Xu, Si, et al.
Published: (2024)
by: Xu, Si, et al.
Published: (2024)
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
by: Li, Shiyao, et al.
Published: (2024)
by: Li, Shiyao, et al.
Published: (2024)
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K
by: Yuan, Tao, et al.
Published: (2024)
by: Yuan, Tao, et al.
Published: (2024)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
by: Liu, Tengxuan, et al.
Published: (2025)
by: Liu, Tengxuan, et al.
Published: (2025)
A Survey on Efficient Inference for Large Language Models
by: Zhou, Zixuan, et al.
Published: (2024)
by: Zhou, Zixuan, et al.
Published: (2024)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
by: Liu, Enshu, et al.
Published: (2025)
by: Liu, Enshu, et al.
Published: (2025)
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
by: Liu, Enshu, et al.
Published: (2024)
by: Liu, Enshu, et al.
Published: (2024)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
by: Zhao, Lin, et al.
Published: (2024)
by: Zhao, Lin, et al.
Published: (2024)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
by: Li, Jinhao, et al.
Published: (2023)
by: Li, Jinhao, et al.
Published: (2023)
Spatial Heterogeneity of Dominant Controlling Factors for Seismic Landslide Susceptibility Zones: A Case Study of the Barkam Earthquake
by: Yutao Chen, et al.
Published: (2025)
by: Yutao Chen, et al.
Published: (2025)
Planar Length-Constrained Minimum Spanning Trees
by: Hershkowitz, D Ellis, et al.
Published: (2025)
by: Hershkowitz, D Ellis, et al.
Published: (2025)
Simple Length-Constrained Minimum Spanning Trees
by: Hershkowitz, D Ellis, et al.
Published: (2024)
by: Hershkowitz, D Ellis, et al.
Published: (2024)
Sliding Window Attention for Learned Video Compression
by: Kopte, Alexander, et al.
Published: (2025)
by: Kopte, Alexander, et al.
Published: (2025)
RATTENTION: Towards the Minimal Sliding Window Size in Local-Global Attention Models
by: Wang, Bailin, et al.
Published: (2025)
by: Wang, Bailin, et al.
Published: (2025)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Sliding Window Attention Training for Efficient Large Language Models
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
by: Ning, Xuefei, et al.
Published: (2023)
by: Ning, Xuefei, et al.
Published: (2023)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
On the skein polynomial for links
by: Jiang, Boju, et al.
Published: (2016)
by: Jiang, Boju, et al.
Published: (2016)
No embeddings of solenoids into surfaces
by: Jiang, Boju, et al.
Published: (2006)
by: Jiang, Boju, et al.
Published: (2006)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
by: Peng, Yanxin, et al.
Published: (2025)
by: Peng, Yanxin, et al.
Published: (2025)
Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits
by: Huang, Ziyi, et al.
Published: (2024)
by: Huang, Ziyi, et al.
Published: (2024)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
A Gaussian Sliding Windows Regression Model for Hydrological Inference
by: Schrunner, Stefan, et al.
Published: (2023)
by: Schrunner, Stefan, et al.
Published: (2023)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
by: Hong, Ke, et al.
Published: (2025)
by: Hong, Ke, et al.
Published: (2025)
Similar Items
-
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
by: Liu, Enshu, et al.
Published: (2024) -
Evaluating Quantized Large Language Models
by: Li, Shiyao, et al.
Published: (2024) -
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
by: Fu, Tianyu, et al.
Published: (2024) -
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024) -
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
by: Wang, Luning, et al.
Published: (2024)