Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Junhao, Li, Fangze, Xu, Mingtao, Meng, Feifan, Zhao, Shiju, Hu, Tiancheng, Peng, Ting, Liu, Anmin, Huang, Wenrui, Liu, Chenxu, Hua, Ziyue, Xie, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
di: Zhuang, Jiedong, et al.
Pubblicazione: (2026)
di: Zhuang, Jiedong, et al.
Pubblicazione: (2026)
You Only Need Less Attention at Each Stage in Vision Transformers
di: Zhang, Shuoxi, et al.
Pubblicazione: (2024)
di: Zhang, Shuoxi, et al.
Pubblicazione: (2024)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024)
di: Lou, Chao, et al.
Pubblicazione: (2024)
When Less Is More.
di: Lettis, Lucy
Pubblicazione: (1998)
di: Lettis, Lucy
Pubblicazione: (1998)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
di: Shen, Yuhao, et al.
Pubblicazione: (2026)
di: Shen, Yuhao, et al.
Pubblicazione: (2026)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
di: Liu, Anmin, et al.
Pubblicazione: (2026)
di: Liu, Anmin, et al.
Pubblicazione: (2026)
Antimicrobial Stewardship: When Less is More
di: R, Piscitelli, et al.
Pubblicazione: (2024)
di: R, Piscitelli, et al.
Pubblicazione: (2024)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
di: Yang, Lijie, et al.
Pubblicazione: (2025)
di: Yang, Lijie, et al.
Pubblicazione: (2025)
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
di: De Cristofaro, Domenico, et al.
Pubblicazione: (2026)
di: De Cristofaro, Domenico, et al.
Pubblicazione: (2026)
You Need an Encoder for Native Position-Independent Caching
di: Zhao, Shiju, et al.
Pubblicazione: (2026)
di: Zhao, Shiju, et al.
Pubblicazione: (2026)
When Less Is More: A Sparse Facial Motion Structure For Listening Motion Learning
di: Nguyen, Tri Tung Nguyen, et al.
Pubblicazione: (2025)
di: Nguyen, Tri Tung Nguyen, et al.
Pubblicazione: (2025)
Moirai 2.0: When Less Is More for Time Series Forecasting
di: Liu, Chenghao, et al.
Pubblicazione: (2025)
di: Liu, Chenghao, et al.
Pubblicazione: (2025)
SparseRL-Sync: Lossless Weight Synchronization with ~100x Less Communication
di: Hu, Lucas, et al.
Pubblicazione: (2026)
di: Hu, Lucas, et al.
Pubblicazione: (2026)
When Less Is More: Issues in Collection Development.
di: Cerny, Rosanne Cerny
Pubblicazione: (1991)
di: Cerny, Rosanne Cerny
Pubblicazione: (1991)
When Less is More: The LLM Scaling Paradox in Context Compression
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
di: Guo, Ruishan, et al.
Pubblicazione: (2026)
Less is More: Improving Motion Diffusion Models with Sparse Keyframes
di: Bae, Jinseok, et al.
Pubblicazione: (2025)
di: Bae, Jinseok, et al.
Pubblicazione: (2025)
Advancing Spherical Nucleic Acid Synthesis in Less‐Polar Solvents
di: Yichen Ye, et al.
Pubblicazione: (2025)
di: Yichen Ye, et al.
Pubblicazione: (2025)
When Less Is More: Cultivating a Healthy Collection.
di: Manning, Patricia
Pubblicazione: (1997)
di: Manning, Patricia
Pubblicazione: (1997)
LoRA Learns Less and Forgets Less
di: Biderman, Dan, et al.
Pubblicazione: (2024)
di: Biderman, Dan, et al.
Pubblicazione: (2024)
Getting More from Less: Transfer Learning Improves Sleep Stage Decoding Accuracy in Peripheral Wearable Devices
di: Coon, William G, et al.
Pubblicazione: (2025)
di: Coon, William G, et al.
Pubblicazione: (2025)
LIMI: Less is More for Agency
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
di: Goren, Shani, et al.
Pubblicazione: (2026)
di: Goren, Shani, et al.
Pubblicazione: (2026)
What Constitutes a Less Discriminatory Algorithm?
di: Laufer, Benjamin, et al.
Pubblicazione: (2024)
di: Laufer, Benjamin, et al.
Pubblicazione: (2024)
Statistical Guarantees in the Search for Less Discriminatory Algorithms
di: Hays, Chris, et al.
Pubblicazione: (2025)
di: Hays, Chris, et al.
Pubblicazione: (2025)
The Legal Duty to Search for Less Discriminatory Algorithms
di: Black, Emily, et al.
Pubblicazione: (2024)
di: Black, Emily, et al.
Pubblicazione: (2024)
MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving
di: Zhao, Shiju, et al.
Pubblicazione: (2025)
di: Zhao, Shiju, et al.
Pubblicazione: (2025)
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
di: Zhang, Chenyuan, et al.
Pubblicazione: (2026)
di: Zhang, Chenyuan, et al.
Pubblicazione: (2026)
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
di: Yi, Hao, et al.
Pubblicazione: (2026)
di: Yi, Hao, et al.
Pubblicazione: (2026)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
di: Zheng, Rui-Chen, et al.
Pubblicazione: (2025)
di: Zheng, Rui-Chen, et al.
Pubblicazione: (2025)
Less is More: Sparse Watermarking in LLMs with Enhanced Text Quality
di: Hoang, Duy C., et al.
Pubblicazione: (2024)
di: Hoang, Duy C., et al.
Pubblicazione: (2024)
Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks
di: Tang, Keke, et al.
Pubblicazione: (2025)
di: Tang, Keke, et al.
Pubblicazione: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
Rethinking Tokenization for Clinical Time Series: When Less is More
di: Attrach, Rafi Al, et al.
Pubblicazione: (2025)
di: Attrach, Rafi Al, et al.
Pubblicazione: (2025)
Lossy Compression of Network Feature Data: When Less Is Enough
di: Palmese, Fabio, et al.
Pubblicazione: (2026)
di: Palmese, Fabio, et al.
Pubblicazione: (2026)
When Less is Enough: Efficient Inference via Collaborative Reasoning
di: Chen, Yilei, et al.
Pubblicazione: (2026)
di: Chen, Yilei, et al.
Pubblicazione: (2026)
LIMO: Less is More for Reasoning
di: Ye, Yixin, et al.
Pubblicazione: (2025)
di: Ye, Yixin, et al.
Pubblicazione: (2025)
LIMR: Less is More for RL Scaling
di: Li, Xuefeng, et al.
Pubblicazione: (2025)
di: Li, Xuefeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model
di: Zhuang, Jiedong, et al.
Pubblicazione: (2026) -
You Only Need Less Attention at Each Stage in Vision Transformers
di: Zhang, Shuoxi, et al.
Pubblicazione: (2024) -
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024) -
When Less Is More.
di: Lettis, Lucy
Pubblicazione: (1998) -
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)