Efficient Sparse Attention needs Adaptive Token Release
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Chaoran, Zou, Lixin, Luo, Dan, Tang, Min, Luo, Xiangyang, Li, Zihao, Li, Chenliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
por: Sun, Yang, et al.
Publicado: (2025)
por: Sun, Yang, et al.
Publicado: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
por: Gao, Yizhao, et al.
Publicado: (2026)
por: Gao, Yizhao, et al.
Publicado: (2026)
STS: Efficient Sparse Attention with Speculative Token Sparsity
por: Xu, Ceyu, et al.
Publicado: (2026)
por: Xu, Ceyu, et al.
Publicado: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026)
por: Jo, Dongwon, et al.
Publicado: (2026)
Flow Matching based Sequential Recommender Model
por: Liu, Feng, et al.
Publicado: (2025)
por: Liu, Feng, et al.
Publicado: (2025)
AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation
por: Luo, Lvzhou, et al.
Publicado: (2025)
por: Luo, Lvzhou, et al.
Publicado: (2025)
SmartTrim: Adaptive Tokens and Attention Pruning for Efficient Vision-Language Models
por: Wang, Zekun, et al.
Publicado: (2023)
por: Wang, Zekun, et al.
Publicado: (2023)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
por: Tang, Zecheng, et al.
Publicado: (2026)
por: Tang, Zecheng, et al.
Publicado: (2026)
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
por: Xiao, Emily, et al.
Publicado: (2025)
por: Xiao, Emily, et al.
Publicado: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
por: Wang, Hanrui, et al.
Publicado: (2020)
por: Wang, Hanrui, et al.
Publicado: (2020)
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
por: Chen, Yu, et al.
Publicado: (2026)
por: Chen, Yu, et al.
Publicado: (2026)
Sparser Block-Sparse Attention via Token Permutation
por: Wang, Xinghao, et al.
Publicado: (2025)
por: Wang, Xinghao, et al.
Publicado: (2025)
AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
por: Luo, Shuqing, et al.
Publicado: (2025)
por: Luo, Shuqing, et al.
Publicado: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
por: Lai, Xunhao, et al.
Publicado: (2025)
por: Lai, Xunhao, et al.
Publicado: (2025)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
por: Shen, Zhenyi, et al.
Publicado: (2025)
por: Shen, Zhenyi, et al.
Publicado: (2025)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
por: Li, Shiwei, et al.
Publicado: (2025)
por: Li, Shiwei, et al.
Publicado: (2025)
BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
por: Hagos, Desta Haileselassie, et al.
Publicado: (2025)
por: Hagos, Desta Haileselassie, et al.
Publicado: (2025)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
por: Li, Zheng, et al.
Publicado: (2025)
por: Li, Zheng, et al.
Publicado: (2025)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
por: Wang, Xu, et al.
Publicado: (2025)
por: Wang, Xu, et al.
Publicado: (2025)
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
por: Li, Chen, et al.
Publicado: (2025)
por: Li, Chen, et al.
Publicado: (2025)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
por: Luo, Chuwei, et al.
Publicado: (2022)
por: Luo, Chuwei, et al.
Publicado: (2022)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
por: He, Mutian, et al.
Publicado: (2025)
por: He, Mutian, et al.
Publicado: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
por: Li, Wenxuan, et al.
Publicado: (2025)
por: Li, Wenxuan, et al.
Publicado: (2025)
Rectified Sparse Attention
por: Sun, Yutao, et al.
Publicado: (2025)
por: Sun, Yutao, et al.
Publicado: (2025)
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
por: Li, Ruosen, et al.
Publicado: (2025)
por: Li, Ruosen, et al.
Publicado: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
por: Yan, Siyuan, et al.
Publicado: (2025)
por: Yan, Siyuan, et al.
Publicado: (2025)
AdaSplash: Adaptive Sparse Flash Attention
por: Gonçalves, Nuno, et al.
Publicado: (2025)
por: Gonçalves, Nuno, et al.
Publicado: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
por: Yuan, Jingyang, et al.
Publicado: (2025)
por: Yuan, Jingyang, et al.
Publicado: (2025)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
por: Zhu, Mingcheng, et al.
Publicado: (2026)
por: Zhu, Mingcheng, et al.
Publicado: (2026)
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
por: Xu, Ruotao, et al.
Publicado: (2026)
por: Xu, Ruotao, et al.
Publicado: (2026)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
por: Zhu, Kan, et al.
Publicado: (2025)
por: Zhu, Kan, et al.
Publicado: (2025)
FoodSky: A Food-oriented Large Language Model that Passes the Chef and Dietetic Examination
por: Zhou, Pengfei, et al.
Publicado: (2024)
por: Zhou, Pengfei, et al.
Publicado: (2024)
Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving
por: Liu, Xunzhuo, et al.
Publicado: (2026)
por: Liu, Xunzhuo, et al.
Publicado: (2026)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
por: Li, Xue, et al.
Publicado: (2025)
por: Li, Xue, et al.
Publicado: (2025)
Token-weighted Direct Preference Optimization with Attention
por: Huang, Chengyu, et al.
Publicado: (2026)
por: Huang, Chengyu, et al.
Publicado: (2026)
Lag-Relative Sparse Attention In Long Context Training
por: Liang, Manlai, et al.
Publicado: (2025)
por: Liang, Manlai, et al.
Publicado: (2025)
SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
por: Wang, Bailin, et al.
Publicado: (2026)
por: Wang, Bailin, et al.
Publicado: (2026)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
por: Lu, Zhiyuan, et al.
Publicado: (2026)
por: Lu, Zhiyuan, et al.
Publicado: (2026)
Trainable Dynamic Mask Sparse Attention
por: Shi, Jingze, et al.
Publicado: (2025)
por: Shi, Jingze, et al.
Publicado: (2025)
Ejemplares similares
-
LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
por: Sun, Yang, et al.
Publicado: (2025) -
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
por: Gao, Yizhao, et al.
Publicado: (2026) -
STS: Efficient Sparse Attention with Speculative Token Sparsity
por: Xu, Ceyu, et al.
Publicado: (2026) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026) -
Flow Matching based Sequential Recommender Model
por: Liu, Feng, et al.
Publicado: (2025)