Salvato in:
| Autori principali: | Jo, Hyun-rae, Shin, Dongkun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.20485 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Computation Pruning for the Forgetting Transformer
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
di: Lin, Zhixuan, et al.
Pubblicazione: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
di: He, Mutian, et al.
Pubblicazione: (2025)
di: He, Mutian, et al.
Pubblicazione: (2025)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
di: Zhong, Shuzhang, et al.
Pubblicazione: (2024)
di: Zhong, Shuzhang, et al.
Pubblicazione: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
di: Jo, Dongwon, et al.
Pubblicazione: (2026)
di: Jo, Dongwon, et al.
Pubblicazione: (2026)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
di: Shang, Bingqi, et al.
Pubblicazione: (2025)
di: Shang, Bingqi, et al.
Pubblicazione: (2025)
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
di: Lee, Hayun, et al.
Pubblicazione: (2024)
di: Lee, Hayun, et al.
Pubblicazione: (2024)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
di: Yang, Songlin, et al.
Pubblicazione: (2025)
di: Yang, Songlin, et al.
Pubblicazione: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Think Clearly: Improving Reasoning via Redundant Token Pruning
di: Choi, Daewon, et al.
Pubblicazione: (2025)
di: Choi, Daewon, et al.
Pubblicazione: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
di: Mihaila, George
Pubblicazione: (2026)
di: Mihaila, George
Pubblicazione: (2026)
High-Layer Attention Pruning with Rescaling
di: Liu, Songtao, et al.
Pubblicazione: (2025)
di: Liu, Songtao, et al.
Pubblicazione: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
di: Guo, Zhiyu, et al.
Pubblicazione: (2024)
di: Guo, Zhiyu, et al.
Pubblicazione: (2024)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
di: Choi, Minsik, et al.
Pubblicazione: (2025)
di: Choi, Minsik, et al.
Pubblicazione: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
di: Koh, Seunghee, et al.
Pubblicazione: (2026)
di: Koh, Seunghee, et al.
Pubblicazione: (2026)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
di: Lee, Chungpa, et al.
Pubblicazione: (2026)
di: Lee, Chungpa, et al.
Pubblicazione: (2026)
Hardware-Efficient Attention for Fast Decoding
di: Zadouri, Ted, et al.
Pubblicazione: (2025)
di: Zadouri, Ted, et al.
Pubblicazione: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
di: Lin, Chaofan, et al.
Pubblicazione: (2025)
di: Lin, Chaofan, et al.
Pubblicazione: (2025)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
di: Phan, Buu, et al.
Pubblicazione: (2025)
di: Phan, Buu, et al.
Pubblicazione: (2025)
Earley-Driven Dynamic Pruning for Efficient Structured Decoding
di: Sun, Xintong, et al.
Pubblicazione: (2025)
di: Sun, Xintong, et al.
Pubblicazione: (2025)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
di: Oba, Daisuke, et al.
Pubblicazione: (2026)
di: Oba, Daisuke, et al.
Pubblicazione: (2026)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
di: Ziashahabi, Amir, et al.
Pubblicazione: (2025)
di: Ziashahabi, Amir, et al.
Pubblicazione: (2025)
Softmax Attention with Constant Cost per Token
di: Heinsen, Franz A.
Pubblicazione: (2024)
di: Heinsen, Franz A.
Pubblicazione: (2024)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2023)
di: Anagnostidis, Sotiris, et al.
Pubblicazione: (2023)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
di: Shin, Sungbin, et al.
Pubblicazione: (2024)
di: Shin, Sungbin, et al.
Pubblicazione: (2024)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
di: Qin, Zongyue, et al.
Pubblicazione: (2024)
di: Qin, Zongyue, et al.
Pubblicazione: (2024)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
di: Li, Xue, et al.
Pubblicazione: (2025)
di: Li, Xue, et al.
Pubblicazione: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
di: Xu, Ceyu, et al.
Pubblicazione: (2026)
di: Xu, Ceyu, et al.
Pubblicazione: (2026)
Attention with Trained Embeddings Provably Selects Important Tokens
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
di: Deng, Difan, et al.
Pubblicazione: (2026)
di: Deng, Difan, et al.
Pubblicazione: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
di: Kwek, Eugene, et al.
Pubblicazione: (2025)
di: Kwek, Eugene, et al.
Pubblicazione: (2025)
Mixture of Attentions For Speculative Decoding
di: Zimmer, Matthieu, et al.
Pubblicazione: (2024)
di: Zimmer, Matthieu, et al.
Pubblicazione: (2024)
EntmaxKV: Support-Aware Decoding for Entmax Attention
di: Duarte, Gonçalo, et al.
Pubblicazione: (2026)
di: Duarte, Gonçalo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Adaptive Computation Pruning for the Forgetting Transformer
di: Lin, Zhixuan, et al.
Pubblicazione: (2025) -
Forgetting Transformer: Softmax Attention with a Forget Gate
di: Lin, Zhixuan, et al.
Pubblicazione: (2025) -
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
di: He, Mutian, et al.
Pubblicazione: (2025) -
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
di: Zhong, Shuzhang, et al.
Pubblicazione: (2024) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
di: Jo, Dongwon, et al.
Pubblicazione: (2026)