On the Emergence of Position Bias in Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Xinyi, Wang, Yifei, Jegelka, Stefanie, Jadbabaie, Ali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On the Role of Attention Masks and LayerNorm in Transformers
di: Wu, Xinyi, et al.
Pubblicazione: (2024)
di: Wu, Xinyi, et al.
Pubblicazione: (2024)
Residual Connections and Normalization Can Provably Prevent Oversmoothing in GNNs
di: Scholkemper, Michael, et al.
Pubblicazione: (2024)
di: Scholkemper, Michael, et al.
Pubblicazione: (2024)
Demystifying Oversmoothing in Attention-Based Graph Neural Networks
di: Wu, Xinyi, et al.
Pubblicazione: (2023)
di: Wu, Xinyi, et al.
Pubblicazione: (2023)
A Theoretical Understanding of Self-Correction through In-context Alignment
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
How to Craft Backdoors with Unlabeled Data Alone?
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
A Canonicalization Perspective on Invariant and Equivariant Learning
di: Ma, George, et al.
Pubblicazione: (2024)
di: Ma, George, et al.
Pubblicazione: (2024)
Higher-Order Graphon Neural Networks: Approximation and Cut Distance
di: Herbst, Daniel, et al.
Pubblicazione: (2025)
di: Herbst, Daniel, et al.
Pubblicazione: (2025)
Sample Complexity Bounds for Estimating Probability Divergences under Invariances
di: Tahmasebi, Behrooz, et al.
Pubblicazione: (2023)
di: Tahmasebi, Behrooz, et al.
Pubblicazione: (2023)
The Exact Sample Complexity Gain from Invariances for Kernel Regression
di: Tahmasebi, Behrooz, et al.
Pubblicazione: (2023)
di: Tahmasebi, Behrooz, et al.
Pubblicazione: (2023)
Simplicity Bias via Global Convergence of Sharpness Minimization
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
di: Sun, Haoyuan, et al.
Pubblicazione: (2025)
di: Sun, Haoyuan, et al.
Pubblicazione: (2025)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
di: Guo, Xiaojun, et al.
Pubblicazione: (2025)
di: Guo, Xiaojun, et al.
Pubblicazione: (2025)
In-Context Symmetries: Self-Supervised Learning through Contextual World Models
di: Gupta, Sharut, et al.
Pubblicazione: (2024)
di: Gupta, Sharut, et al.
Pubblicazione: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
Generalization, Expressivity, and Universality of Graph Neural Networks on Attributed Graphs
di: Rauchwerger, Levi, et al.
Pubblicazione: (2024)
di: Rauchwerger, Levi, et al.
Pubblicazione: (2024)
Neural Networks With Dense Weights Are Not Universal Approximators
di: Rauchwerger, Levi, et al.
Pubblicazione: (2026)
di: Rauchwerger, Levi, et al.
Pubblicazione: (2026)
Counting Substructures with Higher-Order Graph Neural Networks: Possibility and Impossibility Results
di: Tahmasebi, Behrooz, et al.
Pubblicazione: (2020)
di: Tahmasebi, Behrooz, et al.
Pubblicazione: (2020)
A Poincaré Inequality and Consistency Results for Signal Sampling on Large Graphs
di: Le, Thien, et al.
Pubblicazione: (2023)
di: Le, Thien, et al.
Pubblicazione: (2023)
Online Learning for Supervisory Switching Control
di: Sun, Haoyuan, et al.
Pubblicazione: (2026)
di: Sun, Haoyuan, et al.
Pubblicazione: (2026)
A least-square method for non-asymptotic identification in linear switching control
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
di: Sun, Haoyuan, et al.
Pubblicazione: (2024)
On the Stability of Expressive Positional Encodings for Graphs
di: Huang, Yinan, et al.
Pubblicazione: (2023)
di: Huang, Yinan, et al.
Pubblicazione: (2023)
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
di: Zhang, Qi, et al.
Pubblicazione: (2024)
di: Zhang, Qi, et al.
Pubblicazione: (2024)
The Pitfalls of Imitation Learning when Actions are Continuous
di: Simchowitz, Max, et al.
Pubblicazione: (2025)
di: Simchowitz, Max, et al.
Pubblicazione: (2025)
A Test-Function Approach to Incremental Stability
di: Pfrommer, Daniel, et al.
Pubblicazione: (2025)
di: Pfrommer, Daniel, et al.
Pubblicazione: (2025)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
Scaling Attention via Feature Sparsity
di: Xie, Yan, et al.
Pubblicazione: (2026)
di: Xie, Yan, et al.
Pubblicazione: (2026)
A Structural Theory of Position Bias in Transformers
di: Herasimchyk, Hanna, et al.
Pubblicazione: (2026)
di: Herasimchyk, Hanna, et al.
Pubblicazione: (2026)
Understanding the Role of Equivariance in Self-supervised Learning
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
A projection-based framework for gradient-free and parallel learning
di: Bergmeister, Andreas, et al.
Pubblicazione: (2025)
di: Bergmeister, Andreas, et al.
Pubblicazione: (2025)
LieAugmenter: Equivariant Learning by Discovering Symmetries with Learnable Augmentations
di: Santos-Escriche, Eduardo, et al.
Pubblicazione: (2025)
di: Santos-Escriche, Eduardo, et al.
Pubblicazione: (2025)
Learning Efficient Positional Encodings with Graph Neural Networks
di: Kanatsoulis, Charilaos I., et al.
Pubblicazione: (2025)
di: Kanatsoulis, Charilaos I., et al.
Pubblicazione: (2025)
How to escape sharp minima with random perturbations
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
Federated Optimization of Smooth Loss Functions
di: Jadbabaie, Ali, et al.
Pubblicazione: (2022)
di: Jadbabaie, Ali, et al.
Pubblicazione: (2022)
What is Wrong with Perplexity for Long-context Language Modeling?
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
di: Fang, Lizhe, et al.
Pubblicazione: (2024)
Geometric Algorithms for Neural Combinatorial Optimization with Constraints
di: Karalias, Nikolaos, et al.
Pubblicazione: (2025)
di: Karalias, Nikolaos, et al.
Pubblicazione: (2025)
An Information Criterion for Controlled Disentanglement of Multimodal Data
di: Wang, Chenyu, et al.
Pubblicazione: (2024)
di: Wang, Chenyu, et al.
Pubblicazione: (2024)
Learning with Exact Invariances in Polynomial Time
di: Soleymani, Ashkan, et al.
Pubblicazione: (2025)
di: Soleymani, Ashkan, et al.
Pubblicazione: (2025)
Survey on Generalization Theory for Graph Neural Networks
di: Vasileiou, Antonis, et al.
Pubblicazione: (2025)
di: Vasileiou, Antonis, et al.
Pubblicazione: (2025)
Documenti analoghi
-
On the Role of Attention Masks and LayerNorm in Transformers
di: Wu, Xinyi, et al.
Pubblicazione: (2024) -
Residual Connections and Normalization Can Provably Prevent Oversmoothing in GNNs
di: Scholkemper, Michael, et al.
Pubblicazione: (2024) -
Demystifying Oversmoothing in Attention-Based Graph Neural Networks
di: Wu, Xinyi, et al.
Pubblicazione: (2023) -
A Theoretical Understanding of Self-Correction through In-context Alignment
di: Wang, Yifei, et al.
Pubblicazione: (2024) -
How to Craft Backdoors with Unlabeled Data Alone?
di: Wang, Yifei, et al.
Pubblicazione: (2024)