A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Ustaomeroglu, Muhammed, Qu, Guannan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FSD-CAP: Fractional Subgraph Diffusion with Class-Aware Propagation for Graph Feature Imputation
by: Qiao, Xin, et al.
Published: (2026)
by: Qiao, Xin, et al.
Published: (2026)
Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing
by: Yi, Zeji, et al.
Published: (2026)
by: Yi, Zeji, et al.
Published: (2026)
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
by: Jeong, Halyun, et al.
Published: (2025)
by: Jeong, Halyun, et al.
Published: (2025)
Backpropagation Through Time For Networks With Long-Term Dependencies
by: Bird, George, et al.
Published: (2021)
by: Bird, George, et al.
Published: (2021)
Teaching and Learning under Deductive Errors
by: Telle, Jan Arne, et al.
Published: (2026)
by: Telle, Jan Arne, et al.
Published: (2026)
Representation and Regression Problems in Neural Networks: Relaxation, Generalization, and Numerics
by: Liu, Kang, et al.
Published: (2024)
by: Liu, Kang, et al.
Published: (2024)
Autoencoded UMAP-Enhanced Clustering for Unsupervised Learning
by: Chavooshi, Malihehsadat, et al.
Published: (2025)
by: Chavooshi, Malihehsadat, et al.
Published: (2025)
A polynomial-time algorithm for deciding the Hilbert Nullstellensatz over $\mathbb{Z}_2$. A proof of $\mathbf{P}=\mathbf{NP}$ hypothesis
by: Petrov, Petar P.
Published: (2022)
by: Petrov, Petar P.
Published: (2022)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025)
by: Li, Yin
Published: (2025)
Online Decision Making with Generative Action Sets
by: Xu, Jianyu, et al.
Published: (2025)
by: Xu, Jianyu, et al.
Published: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
Advances in Set Function Learning: A Survey of Techniques and Applications
by: Xie, Jiahao, et al.
Published: (2025)
by: Xie, Jiahao, et al.
Published: (2025)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
by: Xu, Zhi-Qin John, et al.
Published: (2019)
by: Xu, Zhi-Qin John, et al.
Published: (2019)
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
Different Statistical Perspectives for Understanding Generalisation in Graph Neural Networks
by: Ayday, Nil, et al.
Published: (2026)
by: Ayday, Nil, et al.
Published: (2026)
A Function-Space Stability Boundary for Generalization in Interpolating Learning Systems
by: Katende, Ronald
Published: (2026)
by: Katende, Ronald
Published: (2026)
Generalizing Adam to Manifolds for Efficiently Training Transformers
by: Brantner, Benedikt
Published: (2023)
by: Brantner, Benedikt
Published: (2023)
Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
by: Chen, Qiyu, et al.
Published: (2025)
by: Chen, Qiyu, et al.
Published: (2025)
The global convergence time of stochastic gradient descent in non-convex landscapes: Sharp estimates via large deviations
by: Azizian, Waïss, et al.
Published: (2025)
by: Azizian, Waïss, et al.
Published: (2025)
Consensus-based optimization for closed-box adversarial attacks and a connection to evolution strategies
by: Roith, Tim, et al.
Published: (2025)
by: Roith, Tim, et al.
Published: (2025)
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
by: Vary, Simon, et al.
Published: (2024)
by: Vary, Simon, et al.
Published: (2024)
What is the long-run distribution of stochastic gradient descent? A large deviations analysis
by: Azizian, Waïss, et al.
Published: (2024)
by: Azizian, Waïss, et al.
Published: (2024)
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
by: Sakabe, Eduardo Y., et al.
Published: (2025)
by: Sakabe, Eduardo Y., et al.
Published: (2025)
A Special Case of Quadratic Extrapolation Under the Neural Tangent Kernel
by: Kim, Abiel
Published: (2025)
by: Kim, Abiel
Published: (2025)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
by: Ilin, Ivan, et al.
Published: (2025)
by: Ilin, Ivan, et al.
Published: (2025)
Incremental Certificate Learning for Hybrid Neural Network Verification . A Solver Architecture for Piecewise-Linear Safety Queries
by: Gokavarapu, Chandrasekhar
Published: (2025)
by: Gokavarapu, Chandrasekhar
Published: (2025)
Ambiguous Online Learning
by: Kosoy, Vanessa
Published: (2025)
by: Kosoy, Vanessa
Published: (2025)
Adaptive Discretization in Online Reinforcement Learning
by: Sinclair, Sean R., et al.
Published: (2021)
by: Sinclair, Sean R., et al.
Published: (2021)
Regret Bounds for Robust Online Decision Making
by: Appel, Alexander, et al.
Published: (2025)
by: Appel, Alexander, et al.
Published: (2025)
Adaptive Risk Mitigation in Demand Learning
by: Pakiman, Parshan, et al.
Published: (2020)
by: Pakiman, Parshan, et al.
Published: (2020)
Greedy feature selection: Classifier-dependent feature selection via greedy methods
by: Camattari, Fabiana, et al.
Published: (2024)
by: Camattari, Fabiana, et al.
Published: (2024)
MMD-Balls as Credal Sets: A PAC-Bayesian Framework for Epistemic Uncertainty in Test-Time Adaptation
by: Ariq, Ahanaf Hasan
Published: (2026)
by: Ariq, Ahanaf Hasan
Published: (2026)
CRAFT: Conflict-Resolved Aggregation for Federated Training
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
Tracking the Median of Gradients with a Stochastic Proximal Point Method
by: Schaipp, Fabian, et al.
Published: (2024)
by: Schaipp, Fabian, et al.
Published: (2024)
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
by: Vary, Simon, et al.
Published: (2026)
by: Vary, Simon, et al.
Published: (2026)
The two clocks and the innovation window: When and how generative models learn rules
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
by: Lee, Wooin, et al.
Published: (2026)
by: Lee, Wooin, et al.
Published: (2026)
QGraphLIME - Explaining Quantum Graph Neural Networks
by: Jena, Haribandhu, et al.
Published: (2025)
by: Jena, Haribandhu, et al.
Published: (2025)
A Reduction from Delayed to Immediate Feedback for Online Convex Optimization with Improved Guarantees
by: Ryabchenko, Alexander, et al.
Published: (2026)
by: Ryabchenko, Alexander, et al.
Published: (2026)
The rate of convergence of Bregman proximal methods: Local geometry vs. regularity vs. sharpness
by: Azizian, Waïss, et al.
Published: (2022)
by: Azizian, Waïss, et al.
Published: (2022)
Similar Items
-
FSD-CAP: Fractional Subgraph Diffusion with Class-Aware Propagation for Graph Feature Imputation
by: Qiao, Xin, et al.
Published: (2026) -
Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing
by: Yi, Zeji, et al.
Published: (2026) -
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
by: Jeong, Halyun, et al.
Published: (2025) -
Backpropagation Through Time For Networks With Long-Term Dependencies
by: Bird, George, et al.
Published: (2021) -
Teaching and Learning under Deductive Errors
by: Telle, Jan Arne, et al.
Published: (2026)