Lower bounds for one-layer transformers that compute parity
Fuente:
arXiv
Salvato in:
| Autore principale: | Hsu, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
One-layer transformers fail to solve the induction heads task
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
Lower bounds on transformers with infinite precision
di: Kozachinskiy, Alexander
Pubblicazione: (2024)
di: Kozachinskiy, Alexander
Pubblicazione: (2024)
A completely uniform transformer for parity
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025)
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025)
Dimension lower bounds for linear approaches to function approximation
di: Hsu, Daniel
Pubblicazione: (2025)
di: Hsu, Daniel
Pubblicazione: (2025)
Transformers, parallel computation, and logarithmic depth
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
Expressivity of deterministic quantum computation with one qubit
di: Kim, Yujin, et al.
Pubblicazione: (2024)
di: Kim, Yujin, et al.
Pubblicazione: (2024)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
di: Smart, Matthew, et al.
Pubblicazione: (2025)
di: Smart, Matthew, et al.
Pubblicazione: (2025)
Myosotis: structured computation for attention like layer
di: Egorov, Evgenii, et al.
Pubblicazione: (2025)
di: Egorov, Evgenii, et al.
Pubblicazione: (2025)
Tabular data generation with tensor contraction layers and transformers
di: Silva, Aníbal, et al.
Pubblicazione: (2024)
di: Silva, Aníbal, et al.
Pubblicazione: (2024)
Optimal lower Lipschitz bounds for ReLU layers, saturation, and phase retrieval
di: Freeman, Daniel, et al.
Pubblicazione: (2025)
di: Freeman, Daniel, et al.
Pubblicazione: (2025)
Dynamic layer selection in decoder-only transformers
di: Glavas, Theodore, et al.
Pubblicazione: (2024)
di: Glavas, Theodore, et al.
Pubblicazione: (2024)
Distributed Online Convex Optimization with Efficient Communication: Improved Algorithm and Lower bounds
di: Yang, Sifan, et al.
Pubblicazione: (2026)
di: Yang, Sifan, et al.
Pubblicazione: (2026)
Scaling depth capacity via zero/one-layer model expansion
di: Bu, Zhiqi
Pubblicazione: (2025)
di: Bu, Zhiqi
Pubblicazione: (2025)
Low-degree Lower bounds for clustering in moderate dimension
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
di: Carpentier, Alexandra, et al.
Pubblicazione: (2026)
Explicit integral representations and quantitative bounds for two-layer ReLU networks
di: Lee, Anthony
Pubblicazione: (2026)
di: Lee, Anthony
Pubblicazione: (2026)
Fair regression under localized demographic parity constraints
di: Charpentier, Arthur, et al.
Pubblicazione: (2026)
di: Charpentier, Arthur, et al.
Pubblicazione: (2026)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
di: Agrawal, Shubhada, et al.
Pubblicazione: (2025)
di: Agrawal, Shubhada, et al.
Pubblicazione: (2025)
Optimal and computationally tractable lower bounds for logistic log-likelihoods
di: Anceschi, Niccolò, et al.
Pubblicazione: (2024)
di: Anceschi, Niccolò, et al.
Pubblicazione: (2024)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
di: Raposo, David, et al.
Pubblicazione: (2024)
di: Raposo, David, et al.
Pubblicazione: (2024)
Don't be lazy: CompleteP enables compute-efficient deep transformers
di: Dey, Nolan, et al.
Pubblicazione: (2025)
di: Dey, Nolan, et al.
Pubblicazione: (2025)
Estimation of Energy-dissipation Lower-bounds for Neuromorphic Learning-in-memory
di: Chen, Zihao, et al.
Pubblicazione: (2024)
di: Chen, Zihao, et al.
Pubblicazione: (2024)
GraphXForm: Graph transformer for computer-aided molecular design
di: Pirnay, Jonathan, et al.
Pubblicazione: (2024)
di: Pirnay, Jonathan, et al.
Pubblicazione: (2024)
Simple and near-optimal algorithms for hidden stratification and multi-group learning
di: Tosh, Christopher, et al.
Pubblicazione: (2021)
di: Tosh, Christopher, et al.
Pubblicazione: (2021)
Multi-group Learning for Hierarchical Groups
di: Deng, Samuel, et al.
Pubblicazione: (2024)
di: Deng, Samuel, et al.
Pubblicazione: (2024)
Efficient transformer adaptation for analog in-memory computing via low-rank adapters
di: Li, Chen, et al.
Pubblicazione: (2024)
di: Li, Chen, et al.
Pubblicazione: (2024)
Can AI-predicted complexes teach machine learning to compute drug binding affinity?
di: Hsu, Wei-Tse, et al.
Pubblicazione: (2025)
di: Hsu, Wei-Tse, et al.
Pubblicazione: (2025)
Demographic parity in regression and classification within the unawareness framework
di: Divol, Vincent, et al.
Pubblicazione: (2024)
di: Divol, Vincent, et al.
Pubblicazione: (2024)
Survey on Algorithms for multi-index models
di: Bruna, Joan, et al.
Pubblicazione: (2025)
di: Bruna, Joan, et al.
Pubblicazione: (2025)
Lipschitz-bounded 1D convolutional neural networks using the Cayley transform and the controllability Gramian
di: Pauli, Patricia, et al.
Pubblicazione: (2023)
di: Pauli, Patricia, et al.
Pubblicazione: (2023)
Infinite-dimensional generative diffusions via Doob's h-transform
di: Pieper-Sethmacher, Thorben, et al.
Pubblicazione: (2026)
di: Pieper-Sethmacher, Thorben, et al.
Pubblicazione: (2026)
How does promoting the minority fraction affect generalization? A theoretical study of the one-hidden-layer neural network on group imbalance
di: Li, Hongkang, et al.
Pubblicazione: (2024)
di: Li, Hongkang, et al.
Pubblicazione: (2024)
Asymptotics of feature learning in two-layer networks after one gradient-step
di: Cui, Hugo, et al.
Pubblicazione: (2024)
di: Cui, Hugo, et al.
Pubblicazione: (2024)
ShakyPrepend: A Multi-Group Learner with Improved Sample Complexity
di: Zhang, Lujing, et al.
Pubblicazione: (2026)
di: Zhang, Lujing, et al.
Pubblicazione: (2026)
Fixed Universal Transformers
di: Liu, Jingwen, et al.
Pubblicazione: (2026)
di: Liu, Jingwen, et al.
Pubblicazione: (2026)
A One-Inclusion Graph Approach to Multi-Group Learning
di: Bergam, Noah, et al.
Pubblicazione: (2026)
di: Bergam, Noah, et al.
Pubblicazione: (2026)
GeoAI Reproducibility and Replicability: a computational and spatial perspective
di: Li, Wenwen, et al.
Pubblicazione: (2024)
di: Li, Wenwen, et al.
Pubblicazione: (2024)
Injectivity of ReLU-layers: Tools from Frame Theory
di: Haider, Daniel, et al.
Pubblicazione: (2024)
di: Haider, Daniel, et al.
Pubblicazione: (2024)
Regression under demographic parity constraints via unlabeled post-processing
di: Chzhen, Evgenii, et al.
Pubblicazione: (2024)
di: Chzhen, Evgenii, et al.
Pubblicazione: (2024)
Group-blind optimal transport to group parity and its constrained variants
di: Zhou, Quan, et al.
Pubblicazione: (2023)
di: Zhou, Quan, et al.
Pubblicazione: (2023)
Non-omniscient backdoor injection with one poison sample: Proving the one-poison hypothesis for linear regression, linear classification, and 2-layer ReLU neural networks
di: Peinemann, Thorsten, et al.
Pubblicazione: (2025)
di: Peinemann, Thorsten, et al.
Pubblicazione: (2025)
Documenti analoghi
-
One-layer transformers fail to solve the induction heads task
di: Sanford, Clayton, et al.
Pubblicazione: (2024) -
Lower bounds on transformers with infinite precision
di: Kozachinskiy, Alexander
Pubblicazione: (2024) -
A completely uniform transformer for parity
di: Kozachinskiy, Alexander, et al.
Pubblicazione: (2025) -
Dimension lower bounds for linear approaches to function approximation
di: Hsu, Daniel
Pubblicazione: (2025) -
Transformers, parallel computation, and logarithmic depth
di: Sanford, Clayton, et al.
Pubblicazione: (2024)