Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
Fuente:
arXiv
Salvato in:
| Autori principali: | Jo, Dongwon, Kang, Beomseok, Song, Jiwon, Kim, Jae-Joon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
di: Song, Jiwon, et al.
Pubblicazione: (2026)
di: Song, Jiwon, et al.
Pubblicazione: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
di: Kang, Beomseok, et al.
Pubblicazione: (2026)
di: Kang, Beomseok, et al.
Pubblicazione: (2026)
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
di: Jo, Dongwon, et al.
Pubblicazione: (2024)
di: Jo, Dongwon, et al.
Pubblicazione: (2024)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
di: Song, Jiwon, et al.
Pubblicazione: (2025)
di: Song, Jiwon, et al.
Pubblicazione: (2025)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
di: Kang, Beomseok, et al.
Pubblicazione: (2025)
di: Kang, Beomseok, et al.
Pubblicazione: (2025)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
di: Wu, Wei, et al.
Pubblicazione: (2024)
di: Wu, Wei, et al.
Pubblicazione: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
di: Xu, Ceyu, et al.
Pubblicazione: (2026)
di: Xu, Ceyu, et al.
Pubblicazione: (2026)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
di: Wu, Zimeng, et al.
Pubblicazione: (2026)
di: Wu, Zimeng, et al.
Pubblicazione: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
di: Fu, Qichen, et al.
Pubblicazione: (2024)
di: Fu, Qichen, et al.
Pubblicazione: (2024)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
di: Lai, Xunhao, et al.
Pubblicazione: (2025)
di: Lai, Xunhao, et al.
Pubblicazione: (2025)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
di: Song, Jiwon, et al.
Pubblicazione: (2024)
di: Song, Jiwon, et al.
Pubblicazione: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
di: He, Mutian, et al.
Pubblicazione: (2025)
di: He, Mutian, et al.
Pubblicazione: (2025)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
di: Kim, Jongsuk, et al.
Pubblicazione: (2024)
di: Kim, Jongsuk, et al.
Pubblicazione: (2024)
One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
di: Addanki, Raghav, et al.
Pubblicazione: (2023)
di: Addanki, Raghav, et al.
Pubblicazione: (2023)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
Token Distillation: Attention-aware Input Embeddings For New Tokens
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
TokenShapley: Token Level Context Attribution with Shapley Value
di: Xiao, Yingtai, et al.
Pubblicazione: (2025)
di: Xiao, Yingtai, et al.
Pubblicazione: (2025)
Multipole Attention for Efficient Long Context Reasoning
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
di: Simoulin, Antoine, et al.
Pubblicazione: (2025)
di: Simoulin, Antoine, et al.
Pubblicazione: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
di: Shin, Seungjun, et al.
Pubblicazione: (2025)
di: Shin, Seungjun, et al.
Pubblicazione: (2025)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
di: Jaber, Jaber, et al.
Pubblicazione: (2026)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
di: Nagaraj, Manish, et al.
Pubblicazione: (2025)
di: Nagaraj, Manish, et al.
Pubblicazione: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
di: Tang, Jiaming, et al.
Pubblicazione: (2024)
di: Tang, Jiaming, et al.
Pubblicazione: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
Softmax Attention with Constant Cost per Token
di: Heinsen, Franz A.
Pubblicazione: (2024)
di: Heinsen, Franz A.
Pubblicazione: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
di: Ma, Qingsen, et al.
Pubblicazione: (2026)
di: Ma, Qingsen, et al.
Pubblicazione: (2026)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
di: MiniCPM Team, et al.
Pubblicazione: (2026)
di: MiniCPM Team, et al.
Pubblicazione: (2026)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
di: Julistiono, Addison Kristanto, et al.
Pubblicazione: (2024)
di: Julistiono, Addison Kristanto, et al.
Pubblicazione: (2024)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
di: Goel, Raghavv, et al.
Pubblicazione: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
TIDE: Every Layer Knows the Token Beneath the Context
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
di: Du, Yufeng, et al.
Pubblicazione: (2026)
di: Du, Yufeng, et al.
Pubblicazione: (2026)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
di: Liu, Di, et al.
Pubblicazione: (2024)
di: Liu, Di, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
di: Song, Jiwon, et al.
Pubblicazione: (2026) -
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025) -
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
di: Kang, Beomseok, et al.
Pubblicazione: (2026) -
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
di: Jo, Dongwon, et al.
Pubblicazione: (2024) -
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
di: Song, Jiwon, et al.
Pubblicazione: (2025)