CAST: Clustering Self-Attention using Surrogate Tokens for Efficient Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | van Engelenhoven, Adjorn, Strisciuglio, Nicola, Talavera, Estefanía |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Regressing Transformers for Data-efficient Visual Place Recognition
von: Leyva-Vallina, María, et al.
Veröffentlicht: (2024)
von: Leyva-Vallina, María, et al.
Veröffentlicht: (2024)
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
von: Kim, Minwook, et al.
Veröffentlicht: (2023)
von: Kim, Minwook, et al.
Veröffentlicht: (2023)
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
von: Vaish, Puru, et al.
Veröffentlicht: (2024)
von: Vaish, Puru, et al.
Veröffentlicht: (2024)
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
von: Ferrari, Alan
Veröffentlicht: (2026)
von: Ferrari, Alan
Veröffentlicht: (2026)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)
Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
von: Pham, Duy-Tung, et al.
Veröffentlicht: (2025)
von: Pham, Duy-Tung, et al.
Veröffentlicht: (2025)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
von: Berasi, Davide, et al.
Veröffentlicht: (2025)
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
von: Kinfu, Kaleab A., et al.
Veröffentlicht: (2025)
von: Kinfu, Kaleab A., et al.
Veröffentlicht: (2025)
CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction
von: Lee, Jaewan, et al.
Veröffentlicht: (2025)
von: Lee, Jaewan, et al.
Veröffentlicht: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
von: Huang, Siyuan, et al.
Veröffentlicht: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
von: Wu, Ziyang, et al.
Veröffentlicht: (2024)
von: Wu, Ziyang, et al.
Veröffentlicht: (2024)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
Geometry Aware Operator Transformer as an Efficient and Accurate Neural Surrogate for PDEs on Arbitrary Domains
von: Wen, Shizheng, et al.
Veröffentlicht: (2025)
von: Wen, Shizheng, et al.
Veröffentlicht: (2025)
Token Sample Complexity of Attention
von: Bohbot, Léa, et al.
Veröffentlicht: (2025)
von: Bohbot, Léa, et al.
Veröffentlicht: (2025)
Efficient Multivector Retrieval with Token-Aware Clustering and Hierarchical Indexing
von: Martinico, Silvio, et al.
Veröffentlicht: (2026)
von: Martinico, Silvio, et al.
Veröffentlicht: (2026)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
von: Bu, Rui, et al.
Veröffentlicht: (2025)
von: Bu, Rui, et al.
Veröffentlicht: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
von: Mihaila, George
Veröffentlicht: (2026)
von: Mihaila, George
Veröffentlicht: (2026)
Understanding Differential Transformer Unchains Pretrained Self-Attentions
von: Kong, Chaerin, et al.
Veröffentlicht: (2025)
von: Kong, Chaerin, et al.
Veröffentlicht: (2025)
CAST: Causal Anchored Simplex Transport for Distribution-Valued Time Series
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
Indoor scene recognition from images under visual corruptions
von: Costa, Willams de Lima, et al.
Veröffentlicht: (2024)
von: Costa, Willams de Lima, et al.
Veröffentlicht: (2024)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
von: Akulov, Dmitry, et al.
Veröffentlicht: (2025)
von: Akulov, Dmitry, et al.
Veröffentlicht: (2025)
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
von: Kim, Gyudong, et al.
Veröffentlicht: (2025)
von: Kim, Gyudong, et al.
Veröffentlicht: (2025)
Clustering by Attention: Leveraging Prior Fitted Transformers for Data Partitioning
von: Shokry, Ahmed, et al.
Veröffentlicht: (2025)
von: Shokry, Ahmed, et al.
Veröffentlicht: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
Multi-level Optimal Control with Neural Surrogate Models
von: Kalise, Dante, et al.
Veröffentlicht: (2024)
von: Kalise, Dante, et al.
Veröffentlicht: (2024)
Attention-Informed Surrogates for Navigating Power-Performance Trade-offs in HPC
von: Ahmed, Ashna Nawar, et al.
Veröffentlicht: (2026)
von: Ahmed, Ashna Nawar, et al.
Veröffentlicht: (2026)
Graph Convolutions Enrich the Self-Attention in Transformers!
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2023)
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2023)
The CAST package for training and assessment of spatial prediction models in R
von: Meyer, Hanna, et al.
Veröffentlicht: (2024)
von: Meyer, Hanna, et al.
Veröffentlicht: (2024)
Estimating Treatment Effects using Multiple Surrogates: The Role of the Surrogate Score and the Surrogate Index
von: Athey, Susan, et al.
Veröffentlicht: (2016)
von: Athey, Susan, et al.
Veröffentlicht: (2016)
Efficient Visual Transformer by Learnable Token Merging
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yancheng, et al.
Veröffentlicht: (2024)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
von: Simoulin, Antoine, et al.
Veröffentlicht: (2025)
von: Simoulin, Antoine, et al.
Veröffentlicht: (2025)
Patch-Level Tokenization with CNN Encoders and Attention for Improved Transformer Time-Series Forecasting
von: Nagrath, Saurish, et al.
Veröffentlicht: (2026)
von: Nagrath, Saurish, et al.
Veröffentlicht: (2026)
Self-Attention as Distributional Projection: A Unified Interpretation of Transformer Architecture
von: Mehta, Nihal
Veröffentlicht: (2025)
von: Mehta, Nihal
Veröffentlicht: (2025)
Multistability of Self-Attention Dynamics in Transformers
von: Altafini, Claudio
Veröffentlicht: (2025)
von: Altafini, Claudio
Veröffentlicht: (2025)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
von: Nagaraj, Manish, et al.
Veröffentlicht: (2025)
von: Nagaraj, Manish, et al.
Veröffentlicht: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
von: Zarch, Hossein Entezari, et al.
Veröffentlicht: (2025)
Benign Overfitting in Token Selection of Attention Mechanism
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2024)
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Regressing Transformers for Data-efficient Visual Place Recognition
von: Leyva-Vallina, María, et al.
Veröffentlicht: (2024) -
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
von: Kim, Minwook, et al.
Veröffentlicht: (2023) -
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
von: Vaish, Puru, et al.
Veröffentlicht: (2024) -
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
von: Ferrari, Alan
Veröffentlicht: (2026) -
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
von: Zhou, Jingbo, et al.
Veröffentlicht: (2026)