Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Chaofan, Tang, Jiaming, Yang, Shuo, Wang, Hanshuo, Tang, Tian, Tian, Boyu, Stoica, Ion, Han, Song, Gao, Mingyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
HashAttention: Semantic Sparsity for Faster Inference
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
von: Xu, Peng, et al.
Veröffentlicht: (2024)
von: Xu, Peng, et al.
Veröffentlicht: (2024)
Frustratingly Easy Task-aware Pruning for Large Language Models
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
von: Fu, Yichao, et al.
Veröffentlicht: (2024)
von: Fu, Yichao, et al.
Veröffentlicht: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
von: Xiu, Di, et al.
Veröffentlicht: (2025)
von: Xiu, Di, et al.
Veröffentlicht: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
von: Wang, Hanrui, et al.
Veröffentlicht: (2020)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models
von: Gao, Shangqian, et al.
Veröffentlicht: (2024)
von: Gao, Shangqian, et al.
Veröffentlicht: (2024)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
High-Layer Attention Pruning with Rescaling
von: Liu, Songtao, et al.
Veröffentlicht: (2025)
von: Liu, Songtao, et al.
Veröffentlicht: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
GRASS: Gradient-based Adaptive Layer-wise Importance Sampling for Memory-efficient Large Language Model Fine-tuning
von: Tian, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Tian, Kaiyuan, et al.
Veröffentlicht: (2026)
Prompt-to-Leaderboard
von: Frick, Evan, et al.
Veröffentlicht: (2025)
von: Frick, Evan, et al.
Veröffentlicht: (2025)
Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
Enhanced Structured State Space Models via Grouped FIR Filtering and Attention Sink Mechanisms
von: Meng, Tian, et al.
Veröffentlicht: (2024)
von: Meng, Tian, et al.
Veröffentlicht: (2024)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
von: Tian, Qiwei, et al.
Veröffentlicht: (2025)
von: Tian, Qiwei, et al.
Veröffentlicht: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
von: Cao, Mingyu, et al.
Veröffentlicht: (2024)
von: Cao, Mingyu, et al.
Veröffentlicht: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal Supervision
von: Zhang, Yunyi, et al.
Veröffentlicht: (2024)
von: Zhang, Yunyi, et al.
Veröffentlicht: (2024)
Earley-Driven Dynamic Pruning for Efficient Structured Decoding
von: Sun, Xintong, et al.
Veröffentlicht: (2025)
von: Sun, Xintong, et al.
Veröffentlicht: (2025)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
von: Jain, Naman, et al.
Veröffentlicht: (2025)
von: Jain, Naman, et al.
Veröffentlicht: (2025)
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
von: You, Kaichao, et al.
Veröffentlicht: (2024)
von: You, Kaichao, et al.
Veröffentlicht: (2024)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
von: Tang, Shengkun, et al.
Veröffentlicht: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
von: Lee, Tzu-Yun, et al.
Veröffentlicht: (2025)
von: Lee, Tzu-Yun, et al.
Veröffentlicht: (2025)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
von: Li, Xue, et al.
Veröffentlicht: (2025)
von: Li, Xue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024) -
HashAttention: Semantic Sparsity for Faster Inference
von: Desai, Aditya, et al.
Veröffentlicht: (2024) -
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024) -
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
von: Xu, Peng, et al.
Veröffentlicht: (2024) -
Frustratingly Easy Task-aware Pruning for Large Language Models
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)