Guardado en:
| Autores principales: | van Engelenhoven, Adjorn, Strisciuglio, Nicola, Talavera, Estefanía |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2402.04239 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Regressing Transformers for Data-efficient Visual Place Recognition
por: Leyva-Vallina, María, et al.
Publicado: (2024)
por: Leyva-Vallina, María, et al.
Publicado: (2024)
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
por: Kim, Minwook, et al.
Publicado: (2023)
por: Kim, Minwook, et al.
Publicado: (2023)
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
por: Vaish, Puru, et al.
Publicado: (2024)
por: Vaish, Puru, et al.
Publicado: (2024)
Indoor scene recognition from images under visual corruptions
por: Costa, Willams de Lima, et al.
Publicado: (2024)
por: Costa, Willams de Lima, et al.
Publicado: (2024)
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
por: Ferrari, Alan
Publicado: (2026)
por: Ferrari, Alan
Publicado: (2026)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
por: Berasi, Davide, et al.
Publicado: (2025)
por: Berasi, Davide, et al.
Publicado: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
por: Zhou, Jingbo, et al.
Publicado: (2026)
por: Zhou, Jingbo, et al.
Publicado: (2026)
CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction
por: Lee, Jaewan, et al.
Publicado: (2025)
por: Lee, Jaewan, et al.
Publicado: (2025)
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
por: Kinfu, Kaleab A., et al.
Publicado: (2025)
por: Kinfu, Kaleab A., et al.
Publicado: (2025)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
por: Fu, Zihao, et al.
Publicado: (2025)
por: Fu, Zihao, et al.
Publicado: (2025)
Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
por: Pham, Duy-Tung, et al.
Publicado: (2025)
por: Pham, Duy-Tung, et al.
Publicado: (2025)
Multi-level Optimal Control with Neural Surrogate Models
por: Kalise, Dante, et al.
Publicado: (2024)
por: Kalise, Dante, et al.
Publicado: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026)
por: Jo, Dongwon, et al.
Publicado: (2026)
STS: Efficient Sparse Attention with Speculative Token Sparsity
por: Xu, Ceyu, et al.
Publicado: (2026)
por: Xu, Ceyu, et al.
Publicado: (2026)
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
por: Huang, Siyuan, et al.
Publicado: (2024)
por: Huang, Siyuan, et al.
Publicado: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
por: Wu, Ziyang, et al.
Publicado: (2024)
por: Wu, Ziyang, et al.
Publicado: (2024)
Geometry Aware Operator Transformer as an Efficient and Accurate Neural Surrogate for PDEs on Arbitrary Domains
por: Wen, Shizheng, et al.
Publicado: (2025)
por: Wen, Shizheng, et al.
Publicado: (2025)
Mechanics of Next Token Prediction with Self-Attention
por: Li, Yingcong, et al.
Publicado: (2024)
por: Li, Yingcong, et al.
Publicado: (2024)
CAST: Causal Anchored Simplex Transport for Distribution-Valued Time Series
por: Lu, Jiecheng, et al.
Publicado: (2026)
por: Lu, Jiecheng, et al.
Publicado: (2026)
Efficient Multivector Retrieval with Token-Aware Clustering and Hierarchical Indexing
por: Martinico, Silvio, et al.
Publicado: (2026)
por: Martinico, Silvio, et al.
Publicado: (2026)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
por: Bu, Rui, et al.
Publicado: (2025)
por: Bu, Rui, et al.
Publicado: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
por: Mihaila, George
Publicado: (2026)
por: Mihaila, George
Publicado: (2026)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
por: Akulov, Dmitry, et al.
Publicado: (2025)
por: Akulov, Dmitry, et al.
Publicado: (2025)
Token Sample Complexity of Attention
por: Bohbot, Léa, et al.
Publicado: (2025)
por: Bohbot, Léa, et al.
Publicado: (2025)
Postcolonial Memory in the Netherlands
por: van Engelenhoven, Gerlov
Publicado: (2022)
por: van Engelenhoven, Gerlov
Publicado: (2022)
The CAST package for training and assessment of spatial prediction models in R
por: Meyer, Hanna, et al.
Publicado: (2024)
por: Meyer, Hanna, et al.
Publicado: (2024)
Understanding Differential Transformer Unchains Pretrained Self-Attentions
por: Kong, Chaerin, et al.
Publicado: (2025)
por: Kong, Chaerin, et al.
Publicado: (2025)
Clustering by Attention: Leveraging Prior Fitted Transformers for Data Partitioning
por: Shokry, Ahmed, et al.
Publicado: (2025)
por: Shokry, Ahmed, et al.
Publicado: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
Estimating Treatment Effects using Multiple Surrogates: The Role of the Surrogate Score and the Surrogate Index
por: Athey, Susan, et al.
Publicado: (2016)
por: Athey, Susan, et al.
Publicado: (2016)
First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
por: Kim, Gyudong, et al.
Publicado: (2025)
por: Kim, Gyudong, et al.
Publicado: (2025)
Graph Convolutions Enrich the Self-Attention in Transformers!
por: Choi, Jeongwhan, et al.
Publicado: (2023)
por: Choi, Jeongwhan, et al.
Publicado: (2023)
Efficient Visual Transformer by Learnable Token Merging
por: Wang, Yancheng, et al.
Publicado: (2024)
por: Wang, Yancheng, et al.
Publicado: (2024)
Multistability of Self-Attention Dynamics in Transformers
por: Altafini, Claudio
Publicado: (2025)
por: Altafini, Claudio
Publicado: (2025)
Attention-Informed Surrogates for Navigating Power-Performance Trade-offs in HPC
por: Ahmed, Ashna Nawar, et al.
Publicado: (2026)
por: Ahmed, Ashna Nawar, et al.
Publicado: (2026)
CAST: Modeling Semantic-Level Transitions for Complementary-Aware Sequential Recommendation
por: Zhang, Qian, et al.
Publicado: (2026)
por: Zhang, Qian, et al.
Publicado: (2026)
Cascade Token Selection for Transformer Attention Acceleration
por: Thomas, Stephen J.
Publicado: (2026)
por: Thomas, Stephen J.
Publicado: (2026)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
por: Jia, Mumin, et al.
Publicado: (2025)
por: Jia, Mumin, et al.
Publicado: (2025)
Patch-Level Tokenization with CNN Encoders and Attention for Improved Transformer Time-Series Forecasting
por: Nagrath, Saurish, et al.
Publicado: (2026)
por: Nagrath, Saurish, et al.
Publicado: (2026)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
por: Simoulin, Antoine, et al.
Publicado: (2025)
por: Simoulin, Antoine, et al.
Publicado: (2025)
Ejemplares similares
-
Regressing Transformers for Data-efficient Visual Place Recognition
por: Leyva-Vallina, María, et al.
Publicado: (2024) -
CAST: Cluster-Aware Self-Training for Tabular Data via Reliable Confidence
por: Kim, Minwook, et al.
Publicado: (2023) -
Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification
por: Vaish, Puru, et al.
Publicado: (2024) -
Indoor scene recognition from images under visual corruptions
por: Costa, Willams de Lima, et al.
Publicado: (2024) -
Meta-Attention: Bayesian Per-Token Routing for Efficient Transformer Inference
por: Ferrari, Alan
Publicado: (2026)