Learning interpretable positional encodings in transformers depends on initialization
Fuente:
arXiv
Saved in:
| Main Authors: | Ito, Takuya, Cocchi, Luca, Klinger, Tim, Ram, Parikshit, Campbell, Murray, Hearne, Luke |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying artificial intelligence through algorithmic generalization
by: Ito, Takuya, et al.
Published: (2024)
by: Ito, Takuya, et al.
Published: (2024)
What makes Models Compositional? A Theoretical View: With Supplement
by: Ram, Parikshit, et al.
Published: (2024)
by: Ram, Parikshit, et al.
Published: (2024)
Compositional Program Generation for Few-Shot Systematic Generalization
by: Klinger, Tim, et al.
Published: (2023)
by: Klinger, Tim, et al.
Published: (2023)
Relational Integration Demands Are Tracked by Temporally Delayed Neural Representations in Alpha and Beta Rhythms Within Higher‐Order Cortical Networks
by: Conor Robinson, et al.
Published: (2025)
by: Conor Robinson, et al.
Published: (2025)
Transformers Learn Faster with Semantic Focus
by: Ram, Parikshit, et al.
Published: (2025)
by: Ram, Parikshit, et al.
Published: (2025)
Finding Clustering Algorithms in the Transformer Architecture
by: Clarkson, Kenneth L., et al.
Published: (2025)
by: Clarkson, Kenneth L., et al.
Published: (2025)
On the generalization capacity of neural networks during generic multimodal reasoning
by: Ito, Takuya, et al.
Published: (2024)
by: Ito, Takuya, et al.
Published: (2024)
Alternative positional encoding functions for neural transformers
by: Lopez-Rubio, Ezequiel, et al.
Published: (2025)
by: Lopez-Rubio, Ezequiel, et al.
Published: (2025)
On Learning Representations for Tabular Data Distillation
by: Kang, Inwon, et al.
Published: (2025)
by: Kang, Inwon, et al.
Published: (2025)
Modern Methods in Associative Memory
by: Krotov, Dmitry, et al.
Published: (2025)
by: Krotov, Dmitry, et al.
Published: (2025)
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
by: Carrow, Stephen, et al.
Published: (2024)
by: Carrow, Stephen, et al.
Published: (2024)
Dense Associative Memory with Epanechnikov Energy
by: Hoover, Benjamin, et al.
Published: (2025)
by: Hoover, Benjamin, et al.
Published: (2025)
The role of positional encodings in the ARC benchmark
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
by: Costa, Guilherme H. Bandeira, et al.
Published: (2025)
Balancing Multi-modal Sensor Learning via Multi-objective Optimization
by: Fernando, Heshan, et al.
Published: (2025)
by: Fernando, Heshan, et al.
Published: (2025)
Dense Associative Memory Through the Lens of Random Features
by: Hoover, Benjamin, et al.
Published: (2024)
by: Hoover, Benjamin, et al.
Published: (2024)
Understanding Self-Supervised Learning via Gaussian Mixture Models
by: Bansal, Parikshit, et al.
Published: (2024)
by: Bansal, Parikshit, et al.
Published: (2024)
WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching
by: Helbling, Alec, et al.
Published: (2026)
by: Helbling, Alec, et al.
Published: (2026)
Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
by: Tasse, Geraud Nangue, et al.
Published: (2025)
by: Tasse, Geraud Nangue, et al.
Published: (2025)
Deep Clustering with Associative Memories
by: Saha, Bishwajit, et al.
Published: (2026)
by: Saha, Bishwajit, et al.
Published: (2026)
Do traveling waves make good positional encodings?
by: van de Geijn, Chase, et al.
Published: (2025)
by: van de Geijn, Chase, et al.
Published: (2025)
Context-Free Synthetic Data Mitigates Forgetting
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
Sparse encoding for more-interpretable feature-selecting representations in probabilistic matrix factorization
by: Chang, Joshua C., et al.
Published: (2020)
by: Chang, Joshua C., et al.
Published: (2020)
Weight-sparse transformers have interpretable circuits
by: Gao, Leo, et al.
Published: (2025)
by: Gao, Leo, et al.
Published: (2025)
Enhancing In-context Learning via Linear Probe Calibration
by: Abbas, Momin, et al.
Published: (2024)
by: Abbas, Momin, et al.
Published: (2024)
Learning Power Flow with Confidence: A Probabilistic Guarantee Framework for Voltage Risk
by: Pareek, Parikshit, et al.
Published: (2023)
by: Pareek, Parikshit, et al.
Published: (2023)
On the Utility of Domain-Adjacent Fine-Tuned Model Ensembles for Few-shot Problems
by: Alam, Md Ibrahim Ibne, et al.
Published: (2024)
by: Alam, Md Ibrahim Ibne, et al.
Published: (2024)
Enabling Approximate Joint Sampling in Diffusion LMs
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
Decoupled-Value Attention for Prior-Data Fitted Networks: GP Inference for Physical Equations
by: Sharma, Kaustubh, et al.
Published: (2025)
by: Sharma, Kaustubh, et al.
Published: (2025)
Model Sparsity Can Simplify Machine Unlearning
by: Jia, Jinghan, et al.
Published: (2023)
by: Jia, Jinghan, et al.
Published: (2023)
Targeted Time‐Varying Functional Connectivity
by: Sonsoles Alonso, et al.
Published: (2025)
by: Sonsoles Alonso, et al.
Published: (2025)
SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching
by: Kang, Inwon, et al.
Published: (2026)
by: Kang, Inwon, et al.
Published: (2026)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
Investigating Plausibility of Biologically Inspired Bayesian Learning in ANNs
by: Zaveri, Ram
Published: (2024)
by: Zaveri, Ram
Published: (2024)
Calibration through the Lens of Indistinguishability
by: Gopalan, Parikshit, et al.
Published: (2025)
by: Gopalan, Parikshit, et al.
Published: (2025)
Information-Theoretic Bayesian Optimization for Bilevel Optimization Problems
by: Kanayama, Takuya, et al.
Published: (2025)
by: Kanayama, Takuya, et al.
Published: (2025)
Analysis, Identification and Prediction of Parkinson Disease Sub-Types and Progression through Machine Learning
by: Ram, Ashwin
Published: (2023)
by: Ram, Ashwin
Published: (2023)
Rethinking the long-range dependency in Mamba/SSM and transformer models
by: Ma, Cong, et al.
Published: (2025)
by: Ma, Cong, et al.
Published: (2025)
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds
by: Tsuchiya, Taira, et al.
Published: (2023)
by: Tsuchiya, Taira, et al.
Published: (2023)
Similar Items
-
Quantifying artificial intelligence through algorithmic generalization
by: Ito, Takuya, et al.
Published: (2024) -
What makes Models Compositional? A Theoretical View: With Supplement
by: Ram, Parikshit, et al.
Published: (2024) -
Compositional Program Generation for Few-Shot Systematic Generalization
by: Klinger, Tim, et al.
Published: (2023) -
Relational Integration Demands Are Tracked by Temporally Delayed Neural Representations in Alpha and Beta Rhythms Within Higher‐Order Cortical Networks
by: Conor Robinson, et al.
Published: (2025) -
Transformers Learn Faster with Semantic Focus
by: Ram, Parikshit, et al.
Published: (2025)