STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Meng, Weikang, Huo, Liangyu, Luo, Yadan, Guan, Jiawen, Zhang, Jingyi, Li, Yingjian, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MirrorLA: Reflecting Feature Map for Vision Linear Attention
von: Meng, Weikang, et al.
Veröffentlicht: (2026)
von: Meng, Weikang, et al.
Veröffentlicht: (2026)
Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
von: Deng, Difan, et al.
Veröffentlicht: (2026)
von: Deng, Difan, et al.
Veröffentlicht: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025)
von: He, Mutian, et al.
Veröffentlicht: (2025)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
von: Li, Meng, et al.
Veröffentlicht: (2026)
von: Li, Meng, et al.
Veröffentlicht: (2026)
Implicit Bias in Deep Linear Discriminant Analysis
von: Li, Jiawen
Veröffentlicht: (2026)
von: Li, Jiawen
Veröffentlicht: (2026)
A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation
von: Peng, Yang, et al.
Veröffentlicht: (2025)
von: Peng, Yang, et al.
Veröffentlicht: (2025)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025)
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025)
Benign Overfitting in Token Selection of Attention Mechanism
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2024)
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
von: Wu, Ziyang, et al.
Veröffentlicht: (2024)
von: Wu, Ziyang, et al.
Veröffentlicht: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
Mixture of Layers with Hybrid Attention
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2026)
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2026)
State Rank Dynamics in Linear Attention LLMs
von: Sun, Ao, et al.
Veröffentlicht: (2026)
von: Sun, Ao, et al.
Veröffentlicht: (2026)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
von: Noh, Kanghyun, et al.
Veröffentlicht: (2026)
von: Noh, Kanghyun, et al.
Veröffentlicht: (2026)
Between the Layers Lies the Truth: Uncertainty Estimation in LLMs Using Intra-Layer Local Information Scores
von: Badash, Zvi N., et al.
Veröffentlicht: (2026)
von: Badash, Zvi N., et al.
Veröffentlicht: (2026)
High-Dimensional Analysis of Single-Layer Attention for Sparse-Token Classification
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2025)
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2025)
ConjNorm: Tractable Density Estimation for Out-of-Distribution Detection
von: Peng, Bo, et al.
Veröffentlicht: (2024)
von: Peng, Bo, et al.
Veröffentlicht: (2024)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2026)
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2026)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
Robust Hallucination Detection in LLMs via Adaptive Token Selection
von: Niu, Mengjia, et al.
Veröffentlicht: (2025)
von: Niu, Mengjia, et al.
Veröffentlicht: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
Attention with Trained Embeddings Provably Selects Important Tokens
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation
von: Ding, Fei, et al.
Veröffentlicht: (2026)
von: Ding, Fei, et al.
Veröffentlicht: (2026)
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs
von: Guo, Tianyu, et al.
Veröffentlicht: (2024)
von: Guo, Tianyu, et al.
Veröffentlicht: (2024)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
von: Wang, Zixin, et al.
Veröffentlicht: (2024)
von: Wang, Zixin, et al.
Veröffentlicht: (2024)
Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs
von: Khosravi, Hamed, et al.
Veröffentlicht: (2026)
von: Khosravi, Hamed, et al.
Veröffentlicht: (2026)
Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces
von: Peng, Yang, et al.
Veröffentlicht: (2024)
von: Peng, Yang, et al.
Veröffentlicht: (2024)
Federated Reinforcement Learning with Constraint Heterogeneity
von: Jin, Hao, et al.
Veröffentlicht: (2024)
von: Jin, Hao, et al.
Veröffentlicht: (2024)
Adaptive Time Series Reasoning via Segment Selection
von: Messica, Shvat, et al.
Veröffentlicht: (2026)
von: Messica, Shvat, et al.
Veröffentlicht: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
Kimi Linear: An Expressive, Efficient Attention Architecture
von: Kimi Team, et al.
Veröffentlicht: (2025)
von: Kimi Team, et al.
Veröffentlicht: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
von: Saikumar, Dhananjay, et al.
Veröffentlicht: (2025)
von: Saikumar, Dhananjay, et al.
Veröffentlicht: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
von: Zhang, Yunzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yunzhe, et al.
Veröffentlicht: (2026)
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MirrorLA: Reflecting Feature Map for Vision Linear Attention
von: Meng, Weikang, et al.
Veröffentlicht: (2026) -
Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention
von: Meng, Weikang, et al.
Veröffentlicht: (2025) -
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
von: Meng, Weikang, et al.
Veröffentlicht: (2025) -
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
von: Deng, Difan, et al.
Veröffentlicht: (2026) -
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025)