Self-attention vector output similarities reveal how machines pay attention
Fuente:
arXiv
Saved in:
| Main Authors: | Halevi, Tal, Tzach, Yarden, Gross, Ronit D., Rosner, Shalom, Kanter, Ido |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
by: Tzach, Yarden, et al.
Published: (2025)
by: Tzach, Yarden, et al.
Published: (2025)
Low-latency vision transformers via large-scale multi-head attention
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
by: Koresh, Ella, et al.
Published: (2025)
by: Koresh, Ella, et al.
Published: (2025)
Tiny language models
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
Single-Nodal Spontaneous Symmetry Breaking in NLP Models
by: Rosner, Shalom, et al.
Published: (2026)
by: Rosner, Shalom, et al.
Published: (2026)
Translation Entropy: A Statistical Framework for Evaluating Translation Systems
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
Advanced deep architecture pruning using single filter performance
by: Tzach, Yarden, et al.
Published: (2025)
by: Tzach, Yarden, et al.
Published: (2025)
Towards a universal mechanism for successful deep learning
by: Meir, Yuval, et al.
Published: (2023)
by: Meir, Yuval, et al.
Published: (2023)
Role of Delay in Brain Dynamics
by: Meir, Yuval, et al.
Published: (2024)
by: Meir, Yuval, et al.
Published: (2024)
Visualizing attention zones in machine reading comprehension models
by: Cui, Yiming, et al.
Published: (2024)
by: Cui, Yiming, et al.
Published: (2024)
Deciphering public attention to geoengineering and climate issues using machine learning and dynamic analysis
by: Debnath, Ramit, et al.
Published: (2024)
by: Debnath, Ramit, et al.
Published: (2024)
Inference-time sparse attention with asymmetric indexing
by: Mazaré, Pierre-Emmanuel, et al.
Published: (2025)
by: Mazaré, Pierre-Emmanuel, et al.
Published: (2025)
Relational inductive biases on attention mechanisms
by: Mijangos, Víctor, et al.
Published: (2025)
by: Mijangos, Víctor, et al.
Published: (2025)
Translution: Unifying Self-attention and Convolution for Adaptive and Relative Modeling
by: Fan, Hehe, et al.
Published: (2025)
by: Fan, Hehe, et al.
Published: (2025)
Sentiment analysis with adaptive multi-head attention in Transformer
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
by: Ali, Wazir, et al.
Published: (2020)
by: Ali, Wazir, et al.
Published: (2020)
Cross-attention for State-based model RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
Evaluation of a semi-autonomous attentive listening system with takeover prompting
by: Kawai, Haruki, et al.
Published: (2024)
by: Kawai, Haruki, et al.
Published: (2024)
TransformerFAM: Feedback attention is working memory
by: Hwang, Dongseong, et al.
Published: (2024)
by: Hwang, Dongseong, et al.
Published: (2024)
Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
by: Mo, Shibing, et al.
Published: (2025)
by: Mo, Shibing, et al.
Published: (2025)
An alternative formulation of attention pooling function in translation
by: Conti, Eddie
Published: (2024)
by: Conti, Eddie
Published: (2024)
Simple linear attention language models balance the recall-throughput tradeoff
by: Arora, Simran, et al.
Published: (2024)
by: Arora, Simran, et al.
Published: (2024)
On the token distance modeling ability of higher RoPE attention dimension
by: Hong, Xiangyu, et al.
Published: (2024)
by: Hong, Xiangyu, et al.
Published: (2024)
Integrating a Heterogeneous Graph with Entity-aware Self-attention using Relative Position Labels for Reading Comprehension Model
by: Foolad, Shima, et al.
Published: (2023)
by: Foolad, Shima, et al.
Published: (2023)
Attribute First, then Generate: Locally-attributable Grounded Text Generation
by: Slobodkin, Aviv, et al.
Published: (2024)
by: Slobodkin, Aviv, et al.
Published: (2024)
Multi-head attention debiasing and contrastive learning for mitigating Dataset Artifacts in Natural Language Inference
by: Sivakoti, Karthik
Published: (2024)
by: Sivakoti, Karthik
Published: (2024)
What are you sinking? A geometric approach on attention sink
by: Ruscio, Valeria, et al.
Published: (2025)
by: Ruscio, Valeria, et al.
Published: (2025)
Comparison of different Unique hard attention transformer models by the formal languages they can recognize
by: Ryvkin, Leonid
Published: (2025)
by: Ryvkin, Leonid
Published: (2025)
Conversational salience and mutual attention
by: Christian De Leon
Published: (2025)
by: Christian De Leon
Published: (2025)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
by: Yao, Jinwei, et al.
Published: (2024)
by: Yao, Jinwei, et al.
Published: (2024)
Mapping of attention mechanisms to a generalized Potts model
by: Rende, Riccardo, et al.
Published: (2023)
by: Rende, Riccardo, et al.
Published: (2023)
Supervised learning pays attention
by: Craig, Erin, et al.
Published: (2025)
by: Craig, Erin, et al.
Published: (2025)
Intra-neuronal attention within language models Relationships between activation and semantics
by: Pichat, Michael, et al.
Published: (2025)
by: Pichat, Michael, et al.
Published: (2025)
Probing self-attention in self-supervised speech models for cross-linguistic differences
by: Gopinath, Sai, et al.
Published: (2024)
by: Gopinath, Sai, et al.
Published: (2024)
Institutional-Level Monitoring of Immune Checkpoint Inhibitor IrAEs Using a Novel Natural Language Processing Algorithmic Pipeline
by: Shapiro, Michael, et al.
Published: (2024)
by: Shapiro, Michael, et al.
Published: (2024)
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
by: Zhang, Zhisong, et al.
Published: (2024)
by: Zhang, Zhisong, et al.
Published: (2024)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
by: Shingi, Geet, et al.
Published: (2021)
by: Shingi, Geet, et al.
Published: (2021)
A hybrid transformer and attention based recurrent neural network for robust and interpretable sentiment analysis of tweets
by: Jahin, Md Abrar, et al.
Published: (2024)
by: Jahin, Md Abrar, et al.
Published: (2024)
Ensemble Self-Training for Unsupervised Machine Translation
by: Aharon, Ido, et al.
Published: (2026)
by: Aharon, Ido, et al.
Published: (2026)
Similar Items
-
Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
by: Tzach, Yarden, et al.
Published: (2025) -
Low-latency vision transformers via large-scale multi-head attention
by: Gross, Ronit D., et al.
Published: (2025) -
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
by: Koresh, Ella, et al.
Published: (2025) -
Tiny language models
by: Gross, Ronit D., et al.
Published: (2025) -
Single-Nodal Spontaneous Symmetry Breaking in NLP Models
by: Rosner, Shalom, et al.
Published: (2026)