Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zixuan, Wei, Stanley, Hsu, Daniel, Lee, Jason D. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Transformers Provably Learn Directed Acyclic Graphs via Kernel-Guided Mutual Information
by: Cheng, Yuan, et al.
Published: (2025)
by: Cheng, Yuan, et al.
Published: (2025)
Matrix Completion via Nonsmooth Regularization of Fully Connected Neural Networks
by: Faramarzi, Sajad, et al.
Published: (2024)
by: Faramarzi, Sajad, et al.
Published: (2024)
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Sample-Optimal Locally Private Hypothesis Selection and the Provable Benefits of Interactivity
by: Pour, Alireza F., et al.
Published: (2023)
by: Pour, Alireza F., et al.
Published: (2023)
An Information-Theoretic Analysis of In-Context Learning
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
Sparse Max-Affine Regression
by: Kanj, Haitham, et al.
Published: (2024)
by: Kanj, Haitham, et al.
Published: (2024)
Statistical-Computational Trade-offs in Tensor PCA and Related Problems via Communication Complexity
by: Dudeja, Rishabh, et al.
Published: (2022)
by: Dudeja, Rishabh, et al.
Published: (2022)
On the sample complexity of parameter estimation in logistic regression with normal design
by: Hsu, Daniel, et al.
Published: (2023)
by: Hsu, Daniel, et al.
Published: (2023)
Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and Depth
by: Kögler, Kevin, et al.
Published: (2024)
by: Kögler, Kevin, et al.
Published: (2024)
Provably Efficient Information-Directed Sampling Algorithms for Multi-Agent Reinforcement Learning
by: Zhang, Qiaosheng, et al.
Published: (2024)
by: Zhang, Qiaosheng, et al.
Published: (2024)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
A Connection Between Learning to Reject and Bhattacharyya Divergences
by: Soen, Alexander
Published: (2025)
by: Soen, Alexander
Published: (2025)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Provable Privacy Advantages of Decentralized Federated Learning via Distributed Optimization
by: Yu, Wenrui, et al.
Published: (2024)
by: Yu, Wenrui, et al.
Published: (2024)
Breaking AR's Sampling Bottleneck: Provable Acceleration via Diffusion Language Models
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
On the Hardness of Unsupervised Domain Adaptation: Optimal Learners and Information-Theoretic Perspective
by: Dong, Zhiyi, et al.
Published: (2025)
by: Dong, Zhiyi, et al.
Published: (2025)
Sparse In-Network Learning via Shortest-Path Backpropagation and Finite-Rate Gating
by: Salehi, Mohammad Reza Deylam
Published: (2026)
by: Salehi, Mohammad Reza Deylam
Published: (2026)
LWM-Temporal: Sparse Spatio-Temporal Attention for Wireless Channel Representation Learning
by: Alikhani, Sadjad, et al.
Published: (2026)
by: Alikhani, Sadjad, et al.
Published: (2026)
A Unified Fractional Regularization Framework for Sparse Recovery
by: Zhao, Yinhao, et al.
Published: (2026)
by: Zhao, Yinhao, et al.
Published: (2026)
Estimating Conditional Mutual Information for Dynamic Feature Selection
by: Gadgil, Soham, et al.
Published: (2023)
by: Gadgil, Soham, et al.
Published: (2023)
Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
by: Dorent, Reuben, et al.
Published: (2025)
by: Dorent, Reuben, et al.
Published: (2025)
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
by: Chaabouni, Youssef, et al.
Published: (2025)
by: Chaabouni, Youssef, et al.
Published: (2025)
Byzantine-Resilient Over-the-Air Federated Learning under Zero-Trust Architecture
by: Yao, Jiacheng, et al.
Published: (2025)
by: Yao, Jiacheng, et al.
Published: (2025)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
by: Cai, Changxiao, et al.
Published: (2026)
by: Cai, Changxiao, et al.
Published: (2026)
Accelerating Convergence of Score-Based Diffusion Models, Provably
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
by: Shen, Wei, et al.
Published: (2024)
by: Shen, Wei, et al.
Published: (2024)
Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing
by: Soleymani, Mahdi, et al.
Published: (2025)
by: Soleymani, Mahdi, et al.
Published: (2025)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Route Experts by Sequence, not by Token
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
Computation-aware Energy-harvesting Federated Learning: Cyclic Scheduling with Selective Participation
by: Jeong, Eunjeong, et al.
Published: (2025)
by: Jeong, Eunjeong, et al.
Published: (2025)
VBO-MI: A Fully Gradient-Based Bayesian Optimization Framework Using Variational Mutual Information Estimation
by: Mirkarimi, Farhad
Published: (2026)
by: Mirkarimi, Farhad
Published: (2026)
Similar Items
-
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024) -
Transformers Provably Learn Directed Acyclic Graphs via Kernel-Guided Mutual Information
by: Cheng, Yuan, et al.
Published: (2025) -
Matrix Completion via Nonsmooth Regularization of Fully Connected Neural Networks
by: Faramarzi, Sajad, et al.
Published: (2024) -
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024) -
Sample-Optimal Locally Private Hypothesis Selection and the Provable Benefits of Interactivity
by: Pour, Alireza F., et al.
Published: (2023)