Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
Fuente:
arXiv
Saved in:
| Main Authors: | Kajitsuka, Tokio, Sato, Issei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Optimal Memorization Capacity of Transformers
by: Kajitsuka, Tokio, et al.
Published: (2024)
by: Kajitsuka, Tokio, et al.
Published: (2024)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
The Curious Case of In-Training Compression of State Space Models
by: Chahine, Makram, et al.
Published: (2025)
by: Chahine, Makram, et al.
Published: (2025)
A ZeNN architecture to avoid the Gaussian trap
by: Carvalho, Luís, et al.
Published: (2025)
by: Carvalho, Luís, et al.
Published: (2025)
Interpretability Can Be Actionable
by: Orgad, Hadas, et al.
Published: (2026)
by: Orgad, Hadas, et al.
Published: (2026)
Improving Time Series Classification with Representation Soft Label Smoothing
by: Ma, Hengyi, et al.
Published: (2024)
by: Ma, Hengyi, et al.
Published: (2024)
Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
by: Spyra, Przemysław, et al.
Published: (2025)
by: Spyra, Przemysław, et al.
Published: (2025)
Graph Neural Networks Need Cluster-Normalize-Activate Modules
by: Skryagin, Arseny, et al.
Published: (2024)
by: Skryagin, Arseny, et al.
Published: (2024)
Adaptive Latent-Space Constraints in Personalized Federated Learning
by: Ayromlou, Sana, et al.
Published: (2025)
by: Ayromlou, Sana, et al.
Published: (2025)
Customizing Graph Neural Networks using Path Reweighting
by: Chen, Jianpeng, et al.
Published: (2021)
by: Chen, Jianpeng, et al.
Published: (2021)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
by: Edin, Joakim, et al.
Published: (2025)
by: Edin, Joakim, et al.
Published: (2025)
Continuous SUN (Stable, Unique, and Novel) Metric for Generative Modeling of Inorganic Crystals
by: Negishi, Masahiro, et al.
Published: (2025)
by: Negishi, Masahiro, et al.
Published: (2025)
Deceptive Diffusion: Generating Synthetic Adversarial Examples
by: Beerens, Lucas, et al.
Published: (2024)
by: Beerens, Lucas, et al.
Published: (2024)
TraXion: Rethinking Pre-training Frameworks for Mobility and Beyond
by: Hsu, Shang-Ling, et al.
Published: (2026)
by: Hsu, Shang-Ling, et al.
Published: (2026)
Union of Experts: Adapting Hierarchical Routing to Equivalently Decomposed Transformer
by: Yang, Yujiao, et al.
Published: (2025)
by: Yang, Yujiao, et al.
Published: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Convolutional Neural Networks Can (Meta-)Learn the Same-Different Relation
by: Gupta, Max, et al.
Published: (2025)
by: Gupta, Max, et al.
Published: (2025)
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
by: Lin, Zezheng, et al.
Published: (2026)
by: Lin, Zezheng, et al.
Published: (2026)
Learning to Decode the Surface Code with a Recurrent, Transformer-Based Neural Network
by: Bausch, Johannes, et al.
Published: (2023)
by: Bausch, Johannes, et al.
Published: (2023)
Safe Continual Reinforcement Learning Methods for Nonstationary Environments. Towards a Survey of the State of the Art
by: Tomashevskiy, Timofey
Published: (2026)
by: Tomashevskiy, Timofey
Published: (2026)
UR4NNV: Neural Network Verification, Under-approximation Reachability Works!
by: Liang, Zhen, et al.
Published: (2024)
by: Liang, Zhen, et al.
Published: (2024)
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference
by: Jørgensen, Tollef Emil
Published: (2025)
by: Jørgensen, Tollef Emil
Published: (2025)
The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning
by: Rajput, Vishal
Published: (2026)
by: Rajput, Vishal
Published: (2026)
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
by: Huang, Ningyuan, et al.
Published: (2025)
by: Huang, Ningyuan, et al.
Published: (2025)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
Inhibitor Transformers and Gated RNNs for Torus Efficient Fully Homomorphic Encryption
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
Stealth edits to large language models
by: Sutton, Oliver J., et al.
Published: (2024)
by: Sutton, Oliver J., et al.
Published: (2024)
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)
by: Berg, Cameron, et al.
Published: (2025)
Softly Constrained Denoisers for Diffusion Models Applied to Partial Differential Equations
by: Yeom-Song, Victor M., et al.
Published: (2025)
by: Yeom-Song, Victor M., et al.
Published: (2025)
Intelligent Routing for Sparse Demand Forecasting: A Comparative Evaluation of Selection Strategies
by: Zhang, Qiwen
Published: (2025)
by: Zhang, Qiwen
Published: (2025)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
by: Huang, Yingbing, et al.
Published: (2025)
by: Huang, Yingbing, et al.
Published: (2025)
Progressive Feedforward Collapse of ResNet Training
by: Wang, Sicong, et al.
Published: (2024)
by: Wang, Sicong, et al.
Published: (2024)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Beyond validation loss: Clinically-tailored optimization metrics improve a model's clinical performance
by: Delahunt, Charles B., et al.
Published: (2026)
by: Delahunt, Charles B., et al.
Published: (2026)
Knowledge Abstraction for Knowledge-based Semantic Communication: A Generative Causality Invariant Approach
by: Nguyen, Minh-Duong, et al.
Published: (2025)
by: Nguyen, Minh-Duong, et al.
Published: (2025)
Towards Interpretable Visual Decoding with Attention to Brain Representations
by: Feng, Pinyuan, et al.
Published: (2025)
by: Feng, Pinyuan, et al.
Published: (2025)
A scalable and real-time neural decoder for topological quantum codes
by: Senior, Andrew W., et al.
Published: (2025)
by: Senior, Andrew W., et al.
Published: (2025)
Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention
by: Pendharkar, Ayan
Published: (2026)
by: Pendharkar, Ayan
Published: (2026)
Similar Items
-
On the Optimal Memorization Capacity of Transformers
by: Kajitsuka, Tokio, et al.
Published: (2024) -
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026) -
The Curious Case of In-Training Compression of State Space Models
by: Chahine, Makram, et al.
Published: (2025) -
A ZeNN architecture to avoid the Gaussian trap
by: Carvalho, Luís, et al.
Published: (2025) -
Interpretability Can Be Actionable
by: Orgad, Hadas, et al.
Published: (2026)