CAST: Clustering Self-Attention using Surrogate Tokens for Efficient Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | van Engelenhoven, Adjorn, Strisciuglio, Nicola, Talavera, Estefanía |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Regressing Transformers for Data-efficient Visual Place Recognition
di: Leyva-Vallina, María, et al.
Pubblicazione: (2024)
di: Leyva-Vallina, María, et al.
Pubblicazione: (2024)
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
di: Kim, Minwook, et al.
Pubblicazione: (2023)
di: Kim, Minwook, et al.
Pubblicazione: (2023)
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
di: Vaish, Puru, et al.
Pubblicazione: (2024)
di: Vaish, Puru, et al.
Pubblicazione: (2024)
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
di: Ferrari, Alan
Pubblicazione: (2026)
di: Ferrari, Alan
Pubblicazione: (2026)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
di: Zhou, Jingbo, et al.
Pubblicazione: (2026)
di: Zhou, Jingbo, et al.
Pubblicazione: (2026)
Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
di: Pham, Duy-Tung, et al.
Pubblicazione: (2025)
di: Pham, Duy-Tung, et al.
Pubblicazione: (2025)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
di: Berasi, Davide, et al.
Pubblicazione: (2025)
di: Berasi, Davide, et al.
Pubblicazione: (2025)
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
di: Kinfu, Kaleab A., et al.
Pubblicazione: (2025)
di: Kinfu, Kaleab A., et al.
Pubblicazione: (2025)
CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction
di: Lee, Jaewan, et al.
Pubblicazione: (2025)
di: Lee, Jaewan, et al.
Pubblicazione: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
di: Jo, Dongwon, et al.
Pubblicazione: (2026)
di: Jo, Dongwon, et al.
Pubblicazione: (2026)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
di: Fu, Zihao, et al.
Pubblicazione: (2025)
di: Fu, Zihao, et al.
Pubblicazione: (2025)
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
di: Huang, Siyuan, et al.
Pubblicazione: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
di: Wu, Ziyang, et al.
Pubblicazione: (2024)
di: Wu, Ziyang, et al.
Pubblicazione: (2024)
STS: Efficient Sparse Attention with Speculative Token Sparsity
di: Xu, Ceyu, et al.
Pubblicazione: (2026)
di: Xu, Ceyu, et al.
Pubblicazione: (2026)
Geometry Aware Operator Transformer as an Efficient and Accurate Neural Surrogate for PDEs on Arbitrary Domains
di: Wen, Shizheng, et al.
Pubblicazione: (2025)
di: Wen, Shizheng, et al.
Pubblicazione: (2025)
Token Sample Complexity of Attention
di: Bohbot, Léa, et al.
Pubblicazione: (2025)
di: Bohbot, Léa, et al.
Pubblicazione: (2025)
Efficient Multivector Retrieval with Token-Aware Clustering and Hierarchical Indexing
di: Martinico, Silvio, et al.
Pubblicazione: (2026)
di: Martinico, Silvio, et al.
Pubblicazione: (2026)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
di: Bu, Rui, et al.
Pubblicazione: (2025)
di: Bu, Rui, et al.
Pubblicazione: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
di: Mihaila, George
Pubblicazione: (2026)
di: Mihaila, George
Pubblicazione: (2026)
Understanding Differential Transformer Unchains Pretrained Self-Attentions
di: Kong, Chaerin, et al.
Pubblicazione: (2025)
di: Kong, Chaerin, et al.
Pubblicazione: (2025)
CAST: Causal Anchored Simplex Transport for Distribution-Valued Time Series
di: Lu, Jiecheng, et al.
Pubblicazione: (2026)
di: Lu, Jiecheng, et al.
Pubblicazione: (2026)
Indoor scene recognition from images under visual corruptions
di: Costa, Willams de Lima, et al.
Pubblicazione: (2024)
di: Costa, Willams de Lima, et al.
Pubblicazione: (2024)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
di: Akulov, Dmitry, et al.
Pubblicazione: (2025)
di: Akulov, Dmitry, et al.
Pubblicazione: (2025)
Mechanics of Next Token Prediction with Self-Attention
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
di: Kim, Gyudong, et al.
Pubblicazione: (2025)
di: Kim, Gyudong, et al.
Pubblicazione: (2025)
Clustering by Attention: Leveraging Prior Fitted Transformers for Data Partitioning
di: Shokry, Ahmed, et al.
Pubblicazione: (2025)
di: Shokry, Ahmed, et al.
Pubblicazione: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
Multi-level Optimal Control with Neural Surrogate Models
di: Kalise, Dante, et al.
Pubblicazione: (2024)
di: Kalise, Dante, et al.
Pubblicazione: (2024)
Attention-Informed Surrogates for Navigating Power-Performance Trade-offs in HPC
di: Ahmed, Ashna Nawar, et al.
Pubblicazione: (2026)
di: Ahmed, Ashna Nawar, et al.
Pubblicazione: (2026)
Graph Convolutions Enrich the Self-Attention in Transformers!
di: Choi, Jeongwhan, et al.
Pubblicazione: (2023)
di: Choi, Jeongwhan, et al.
Pubblicazione: (2023)
The CAST package for training and assessment of spatial prediction models in R
di: Meyer, Hanna, et al.
Pubblicazione: (2024)
di: Meyer, Hanna, et al.
Pubblicazione: (2024)
Estimating Treatment Effects using Multiple Surrogates: The Role of the Surrogate Score and the Surrogate Index
di: Athey, Susan, et al.
Pubblicazione: (2016)
di: Athey, Susan, et al.
Pubblicazione: (2016)
Efficient Visual Transformer by Learnable Token Merging
di: Wang, Yancheng, et al.
Pubblicazione: (2024)
di: Wang, Yancheng, et al.
Pubblicazione: (2024)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
di: Simoulin, Antoine, et al.
Pubblicazione: (2025)
di: Simoulin, Antoine, et al.
Pubblicazione: (2025)
Patch-Level Tokenization with CNN Encoders and Attention for Improved Transformer Time-Series Forecasting
di: Nagrath, Saurish, et al.
Pubblicazione: (2026)
di: Nagrath, Saurish, et al.
Pubblicazione: (2026)
Self-Attention as Distributional Projection: A Unified Interpretation of Transformer Architecture
di: Mehta, Nihal
Pubblicazione: (2025)
di: Mehta, Nihal
Pubblicazione: (2025)
Multistability of Self-Attention Dynamics in Transformers
di: Altafini, Claudio
Pubblicazione: (2025)
di: Altafini, Claudio
Pubblicazione: (2025)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
di: Nagaraj, Manish, et al.
Pubblicazione: (2025)
di: Nagaraj, Manish, et al.
Pubblicazione: (2025)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
Benign Overfitting in Token Selection of Attention Mechanism
di: Sakamoto, Keitaro, et al.
Pubblicazione: (2024)
di: Sakamoto, Keitaro, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Regressing Transformers for Data-efficient Visual Place Recognition
di: Leyva-Vallina, María, et al.
Pubblicazione: (2024) -
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
di: Kim, Minwook, et al.
Pubblicazione: (2023) -
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
di: Vaish, Puru, et al.
Pubblicazione: (2024) -
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
di: Ferrari, Alan
Pubblicazione: (2026) -
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
di: Zhou, Jingbo, et al.
Pubblicazione: (2026)