Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bu, Rui, Zhong, Haofeng, Chen, Wenzheng, Li, Yangyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoVE: Mixture of Value Embeddings -- A New Axis for Scaling Parametric Memory in Autoregressive Models
von: Li, Yangyan
Veröffentlicht: (2026)
von: Li, Yangyan
Veröffentlicht: (2026)
ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecast
von: Xu, Wanghan, et al.
Veröffentlicht: (2024)
von: Xu, Wanghan, et al.
Veröffentlicht: (2024)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion
von: Huang, Haofeng, et al.
Veröffentlicht: (2025)
von: Huang, Haofeng, et al.
Veröffentlicht: (2025)
DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
von: Lin, Abigail
Veröffentlicht: (2025)
von: Lin, Abigail
Veröffentlicht: (2025)
Rethinking Tokenized Graph Transformers for Node Classification
von: Chen, Jinsong, et al.
Veröffentlicht: (2025)
von: Chen, Jinsong, et al.
Veröffentlicht: (2025)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
Patch-Level Tokenization with CNN Encoders and Attention for Improved Transformer Time-Series Forecasting
von: Nagrath, Saurish, et al.
Veröffentlicht: (2026)
von: Nagrath, Saurish, et al.
Veröffentlicht: (2026)
Extreme Region Policy Distillation
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
von: Chen, Changyu, et al.
Veröffentlicht: (2026)
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
von: Zixian, Wang
Veröffentlicht: (2025)
von: Zixian, Wang
Veröffentlicht: (2025)
Echo State Transformer: Attention Over Finite Memories
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2025)
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2025)
Twin Transformer using Gated Dynamic Learnable Attention mechanism for Fault Detection and Diagnosis in the Tennessee Eastman Process
von: Labbaf-Khaniki, Mohammad Ali, et al.
Veröffentlicht: (2024)
von: Labbaf-Khaniki, Mohammad Ali, et al.
Veröffentlicht: (2024)
Towards Mitigating Architecture Overfitting on Distilled Datasets
von: Zhong, Xuyang, et al.
Veröffentlicht: (2023)
von: Zhong, Xuyang, et al.
Veröffentlicht: (2023)
A Neural Network Architecture Based on Attention Gate Mechanism for 3D Magnetotelluric Forward Modeling
von: Zhong, Xin, et al.
Veröffentlicht: (2025)
von: Zhong, Xin, et al.
Veröffentlicht: (2025)
GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
von: Ou, Jie, et al.
Veröffentlicht: (2025)
von: Ou, Jie, et al.
Veröffentlicht: (2025)
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
von: Karbevski, Marko, et al.
Veröffentlicht: (2025)
von: Karbevski, Marko, et al.
Veröffentlicht: (2025)
Mitigating Overconfidence in Out-of-Distribution Detection by Capturing Extreme Activations
von: Azizmalayeri, Mohammad, et al.
Veröffentlicht: (2024)
von: Azizmalayeri, Mohammad, et al.
Veröffentlicht: (2024)
Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
von: Zhong, Qimin, et al.
Veröffentlicht: (2025)
von: Zhong, Qimin, et al.
Veröffentlicht: (2025)
Scaling Law Phenomena Across Regression Paradigms: Multiple and Kernel Approaches
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
SEA: State-Exchange Attention for High-Fidelity Physics Based Transformers
von: Esmati, Parsa, et al.
Veröffentlicht: (2024)
von: Esmati, Parsa, et al.
Veröffentlicht: (2024)
VSFormer: Value and Shape-Aware Transformer with Prior-Enhanced Self-Attention for Multivariate Time Series Classification
von: Xi, Wenjie, et al.
Veröffentlicht: (2024)
von: Xi, Wenjie, et al.
Veröffentlicht: (2024)
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
ToMA: Token Merge with Attention for Diffusion Models
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lu, Wenbo, et al.
Veröffentlicht: (2025)
Geometric Analysis of Token Selection in Multi-Head Attention
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
Graph Tokenization for Bridging Graphs and Transformers
von: Guo, Zeyuan, et al.
Veröffentlicht: (2026)
von: Guo, Zeyuan, et al.
Veröffentlicht: (2026)
MIGT: Memory Instance Gated Transformer Framework for Financial Portfolio Management
von: Gu, Fengchen, et al.
Veröffentlicht: (2025)
von: Gu, Fengchen, et al.
Veröffentlicht: (2025)
Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
von: Zhao, Weijie, et al.
Veröffentlicht: (2026)
von: Zhao, Weijie, et al.
Veröffentlicht: (2026)
MiMu: Mitigating Multiple Shortcut Learning Behavior of Transformers
von: Zhao, Lili, et al.
Veröffentlicht: (2025)
von: Zhao, Lili, et al.
Veröffentlicht: (2025)
Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
von: Xie, Xuanting, et al.
Veröffentlicht: (2025)
von: Xie, Xuanting, et al.
Veröffentlicht: (2025)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
Enhanced Graph Transformer with Serialized Graph Tokens
von: Wang, Ruixiang, et al.
Veröffentlicht: (2026)
von: Wang, Ruixiang, et al.
Veröffentlicht: (2026)
GIAT: A Geologically-Informed Attention Transformer for Lithology Identification
von: Li, Jie, et al.
Veröffentlicht: (2026)
von: Li, Jie, et al.
Veröffentlicht: (2026)
CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation
von: Minelli, Giovanni, et al.
Veröffentlicht: (2025)
von: Minelli, Giovanni, et al.
Veröffentlicht: (2025)
DeepGate3: Towards Scalable Circuit Representation Learning
von: Shi, Zhengyuan, et al.
Veröffentlicht: (2024)
von: Shi, Zhengyuan, et al.
Veröffentlicht: (2024)
Quantum Error Mitigation with Attention Graph Transformers for Burgers Equation Solvers on NISQ Hardware
von: Tousi, Seyed Mohamad Ali, et al.
Veröffentlicht: (2025)
von: Tousi, Seyed Mohamad Ali, et al.
Veröffentlicht: (2025)
The Bayesian Geometry of Transformer Attention
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
Physics-Driven Learning Framework for Tomographic Tactile Sensing
von: Yang, Xuanxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xuanxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MoVE: Mixture of Value Embeddings -- A New Axis for Scaling Parametric Memory in Autoregressive Models
von: Li, Yangyan
Veröffentlicht: (2026) -
ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecast
von: Xu, Wanghan, et al.
Veröffentlicht: (2024) -
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
von: Li, Yu, et al.
Veröffentlicht: (2026) -
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026) -
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)