Why are Sensitive Functions Hard for Transformers?
Fuente:
arXiv
Saved in:
| Main Authors: | Hahn, Michael, Rofin, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026)
by: Varre, Aditya, et al.
Published: (2026)
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
by: Ding, Yanna, et al.
Published: (2025)
by: Ding, Yanna, et al.
Published: (2025)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
by: Bhattamishra, Satwik, et al.
Published: (2024)
by: Bhattamishra, Satwik, et al.
Published: (2024)
On the Computational Hardness of Transformers
by: Saha, Barna, et al.
Published: (2026)
by: Saha, Barna, et al.
Published: (2026)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
by: Tiwari, Utkarsh, et al.
Published: (2025)
by: Tiwari, Utkarsh, et al.
Published: (2025)
The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works
by: Wang, Guanghui, et al.
Published: (2026)
by: Wang, Guanghui, et al.
Published: (2026)
Routing Absorption in Sparse Attention: Why Random Gates Are Hard to Beat
by: Aquino-Michaels, Keston
Published: (2026)
by: Aquino-Michaels, Keston
Published: (2026)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
The Hardness of Validating Observational Studies with Experimental Data
by: Fawkes, Jake, et al.
Published: (2025)
by: Fawkes, Jake, et al.
Published: (2025)
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
by: Cotnareanu, Joseph, et al.
Published: (2024)
by: Cotnareanu, Joseph, et al.
Published: (2024)
Parity, Sensitivity, and Transformers
by: Kozachinskiy, Alexander, et al.
Published: (2026)
by: Kozachinskiy, Alexander, et al.
Published: (2026)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
by: Kraus, Oliver, et al.
Published: (2026)
by: Kraus, Oliver, et al.
Published: (2026)
How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
by: Wang, Entang, et al.
Published: (2026)
by: Wang, Entang, et al.
Published: (2026)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
A Formal Framework for Understanding Length Generalization in Transformers
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
by: Liang, Zida, et al.
Published: (2025)
by: Liang, Zida, et al.
Published: (2025)
On the Approximation of Phylogenetic Distance Functions by Artificial Neural Networks
by: Rosenzweig, Benjamin K., et al.
Published: (2025)
by: Rosenzweig, Benjamin K., et al.
Published: (2025)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026)
by: Huang, Xinting, et al.
Published: (2026)
Why Transformers Need Adam: A Hessian Perspective
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
The Depth Delusion: Why Transformers Should Be Wider, Not Deeper
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
by: Fahim, Md Muhtasim Munif, et al.
Published: (2026)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
by: Lee, Nayoung, et al.
Published: (2025)
by: Lee, Nayoung, et al.
Published: (2025)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
Neural Proposals, Symbolic Guarantees: Neuro-Symbolic Graph Generation with Hard Constraints
by: Geng, Chuqin, et al.
Published: (2026)
by: Geng, Chuqin, et al.
Published: (2026)
Towards a More Complete Theory of Function Preserving Transforms
by: Painter, Michael
Published: (2024)
by: Painter, Michael
Published: (2024)
Exact Attention Sensitivity and the Geometry of Transformer Stability
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Softmax Transformers are Turing-Complete
by: Jiang, Hongjian, et al.
Published: (2025)
by: Jiang, Hongjian, et al.
Published: (2025)
Gradient-Free Generation for Hard-Constrained Systems
by: Cheng, Chaoran, et al.
Published: (2024)
by: Cheng, Chaoran, et al.
Published: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
On the Hardness of Junking LLMs
by: Rando, Marco, et al.
Published: (2026)
by: Rando, Marco, et al.
Published: (2026)
On the Hardness of Bandit Learning
by: Brukhim, Nataly, et al.
Published: (2025)
by: Brukhim, Nataly, et al.
Published: (2025)
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
by: Wang, Xiaoshuang, et al.
Published: (2025)
by: Wang, Xiaoshuang, et al.
Published: (2025)
AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
by: Keslaki, Ben
Published: (2025)
by: Keslaki, Ben
Published: (2025)
Spectral Analysis of Hard-Constraint PINNs: The Spatial Modulation Mechanism of Boundary Functions
by: Xie, Yuchen, et al.
Published: (2025)
by: Xie, Yuchen, et al.
Published: (2025)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Equivariant Neural Functional Networks for Transformers
by: Tran, Viet-Hoang, et al.
Published: (2024)
by: Tran, Viet-Hoang, et al.
Published: (2024)
Similar Items
-
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025) -
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026) -
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026) -
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026) -
Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
by: Ding, Yanna, et al.
Published: (2025)