Fixed Universal Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jingwen, Andoni, Alexandr, Hsu, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast attention mechanisms: a tale of parallelism
by: Liu, Jingwen, et al.
Published: (2025)
by: Liu, Jingwen, et al.
Published: (2025)
Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
by: Plashchinsky, Alexandr
Published: (2025)
by: Plashchinsky, Alexandr
Published: (2025)
Group-realizable multi-group learning by minimizing empirical risk
by: Ardeshir, Navid, et al.
Published: (2026)
by: Ardeshir, Navid, et al.
Published: (2026)
Group-wise oracle-efficient algorithms for online multi-group learning
by: Deng, Samuel, et al.
Published: (2024)
by: Deng, Samuel, et al.
Published: (2024)
Transformers, parallel computation, and logarithmic depth
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
A First Guess is Rarely the Final Answer: Learning to Search in the Traveling Salesperson Problem
by: Garmendia, Andoni Irazusta
Published: (2026)
by: Garmendia, Andoni Irazusta
Published: (2026)
In-Context Algorithm Emulation in Fixed-Weight Transformers
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Statistical-Computational Trade-offs for Density Estimation
by: Aamand, Anders, et al.
Published: (2024)
by: Aamand, Anders, et al.
Published: (2024)
Invertible Memory Flow Networks
by: Zerihun, Liyu, et al.
Published: (2026)
by: Zerihun, Liyu, et al.
Published: (2026)
Lower bounds for one-layer transformers that compute parity
by: Hsu, Daniel
Published: (2026)
by: Hsu, Daniel
Published: (2026)
Constraining the outputs of ReLU neural networks
by: Alexandr, Yulia, et al.
Published: (2025)
by: Alexandr, Yulia, et al.
Published: (2025)
Transformers Are Universally Consistent
by: Ghosh, Sagar, et al.
Published: (2025)
by: Ghosh, Sagar, et al.
Published: (2025)
Dimension lower bounds for linear approaches to function approximation
by: Hsu, Daniel
Published: (2025)
by: Hsu, Daniel
Published: (2025)
Simple and near-optimal algorithms for hidden stratification and multi-group learning
by: Tosh, Christopher, et al.
Published: (2021)
by: Tosh, Christopher, et al.
Published: (2021)
Multi-group Learning for Hierarchical Groups
by: Deng, Samuel, et al.
Published: (2024)
by: Deng, Samuel, et al.
Published: (2024)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
by: Zhang, Jinhao, et al.
Published: (2026)
by: Zhang, Jinhao, et al.
Published: (2026)
Enabling Population-Based Architectures for Neural Combinatorial Optimization
by: Garmendia, Andoni Irazusta, et al.
Published: (2026)
by: Garmendia, Andoni Irazusta, et al.
Published: (2026)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
Survey on Algorithms for multi-index models
by: Bruna, Joan, et al.
Published: (2025)
by: Bruna, Joan, et al.
Published: (2025)
Fixed-Mean Gaussian Processes for Post-hoc Bayesian Deep Learning
by: Ortega, Luis A., et al.
Published: (2024)
by: Ortega, Luis A., et al.
Published: (2024)
Risk-Averse Best Arm Set Identification with Fixed Budget and Fixed Confidence
by: Nonaga, Shunta, et al.
Published: (2025)
by: Nonaga, Shunta, et al.
Published: (2025)
Multi-Metric Adaptive Experimental Design under Fixed Budget with Validation
by: Zhang, Qining, et al.
Published: (2025)
by: Zhang, Qining, et al.
Published: (2025)
SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers
by: Liu, Joseph, et al.
Published: (2024)
by: Liu, Joseph, et al.
Published: (2024)
Algebraic Invariants of Lightning Self-Attention
by: Alexandr, Yulia, et al.
Published: (2026)
by: Alexandr, Yulia, et al.
Published: (2026)
Robustness Verification of Polynomial Neural Networks
by: Alexandr, Yulia, et al.
Published: (2026)
by: Alexandr, Yulia, et al.
Published: (2026)
Exact Fixed-Point Constraints in Neural-ODEs with Provable Universality
by: Pacifico, Feliciano Giuseppe, et al.
Published: (2026)
by: Pacifico, Feliciano Giuseppe, et al.
Published: (2026)
Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
by: Sandoval, Gustavo
Published: (2025)
by: Sandoval, Gustavo
Published: (2025)
Frozen in Time: Parameter-Efficient Time Series Transformers via Reservoir-Induced Feature Expansion and Fixed Random Dynamics
by: Singh, Pradeep, et al.
Published: (2025)
by: Singh, Pradeep, et al.
Published: (2025)
Enhancing the Cross-Size Generalization for Solving Vehicle Routing Problems via Continual Learning
by: Li, Jingwen, et al.
Published: (2025)
by: Li, Jingwen, et al.
Published: (2025)
Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence
by: Xiao, Liu
Published: (2026)
by: Xiao, Liu
Published: (2026)
ShakyPrepend: A Multi-Group Learner with Improved Sample Complexity
by: Zhang, Lujing, et al.
Published: (2026)
by: Zhang, Lujing, et al.
Published: (2026)
A One-Inclusion Graph Approach to Multi-Group Learning
by: Bergam, Noah, et al.
Published: (2026)
by: Bergam, Noah, et al.
Published: (2026)
One-layer transformers fail to solve the induction heads task
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
Verification-Guided Shielding for Deep Reinforcement Learning
by: Corsi, Davide, et al.
Published: (2024)
by: Corsi, Davide, et al.
Published: (2024)
Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware Minimization
by: Wang, Jinping, et al.
Published: (2026)
by: Wang, Jinping, et al.
Published: (2026)
Towards Understanding the Universality of Transformers for Next-Token Prediction
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
On the Universality of Transformer Architectures; How Much Attention Is Enough?
by: Abbasi, Amirreza, et al.
Published: (2025)
by: Abbasi, Amirreza, et al.
Published: (2025)
Unified Training of Universal Time Series Forecasting Transformers
by: Woo, Gerald, et al.
Published: (2024)
by: Woo, Gerald, et al.
Published: (2024)
Similar Items
-
Fast attention mechanisms: a tale of parallelism
by: Liu, Jingwen, et al.
Published: (2025) -
Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
by: Plashchinsky, Alexandr
Published: (2025) -
Group-realizable multi-group learning by minimizing empirical risk
by: Ardeshir, Navid, et al.
Published: (2026) -
Group-wise oracle-efficient algorithms for online multi-group learning
by: Deng, Samuel, et al.
Published: (2024) -
Transformers, parallel computation, and logarithmic depth
by: Sanford, Clayton, et al.
Published: (2024)