How Transformers Learn Causal Structure with Gradient Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nichani, Eshaan, Damian, Alex, Lee, Jason D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Factual Recall in Transformers via Associative Memories
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
von: Nichani, Eshaan, et al.
Veröffentlicht: (2023)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2023)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026)
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
von: Fu, Hengyu, et al.
Veröffentlicht: (2024)
von: Fu, Hengyu, et al.
Veröffentlicht: (2024)
Quantitative Bounds for Length Generalization in Transformers
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
von: Huang, Ruomin, et al.
Veröffentlicht: (2026)
von: Huang, Ruomin, et al.
Veröffentlicht: (2026)
On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach
von: Giapitzakis, George, et al.
Veröffentlicht: (2025)
von: Giapitzakis, George, et al.
Veröffentlicht: (2025)
On the Theory of Continual Learning with Gradient Descent for Neural Networks
von: Taheri, Hossein, et al.
Veröffentlicht: (2025)
von: Taheri, Hossein, et al.
Veröffentlicht: (2025)
Emergence and scaling laws in SGD learning of shallow neural networks
von: Ren, Yunwei, et al.
Veröffentlicht: (2025)
von: Ren, Yunwei, et al.
Veröffentlicht: (2025)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
von: Kim, Juno, et al.
Veröffentlicht: (2026)
von: Kim, Juno, et al.
Veröffentlicht: (2026)
Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
Fine-Tuning Language Models with Just Forward Passes
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)
Global Convergence Analysis of Vanilla Gradient Descent for Asymmetric Matrix Completion
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
Rank-1 Matrix Completion with Gradient Descent and Small Random Initialization
von: Kim, Daesung, et al.
Veröffentlicht: (2022)
von: Kim, Daesung, et al.
Veröffentlicht: (2022)
A Mirror Descent-Based Algorithm for Corruption-Tolerant Distributed Gradient Descent
von: Wang, Shuche, et al.
Veröffentlicht: (2024)
von: Wang, Shuche, et al.
Veröffentlicht: (2024)
Guaranteed Nonconvex Low-Rank Tensor Estimation via Scaled Gradient Descent
von: Wu, Tong
Veröffentlicht: (2025)
von: Wu, Tong
Veröffentlicht: (2025)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
von: Yang, Tong, et al.
Veröffentlicht: (2025)
von: Yang, Tong, et al.
Veröffentlicht: (2025)
Gradient Descent Finds Over-Parameterized Neural Networks with Sharp Generalization for Nonparametric Regression
von: Yang, Yingzhen, et al.
Veröffentlicht: (2024)
von: Yang, Yingzhen, et al.
Veröffentlicht: (2024)
Scaled Gradient Descent for Ill-Conditioned Low-Rank Matrix Recovery with Optimal Sampling Complexity
von: Li, Zhenxuan, et al.
Veröffentlicht: (2026)
von: Li, Zhenxuan, et al.
Veröffentlicht: (2026)
Non-Coherent Over-the-Air Decentralized Gradient Descent
von: Michelusi, Nicolo'
Veröffentlicht: (2022)
von: Michelusi, Nicolo'
Veröffentlicht: (2022)
Does Privacy Always Harm Fairness? Data-Dependent Trade-offs via Chernoff Information Neural Estimation
von: Nichani, Arjun, et al.
Veröffentlicht: (2026)
von: Nichani, Arjun, et al.
Veröffentlicht: (2026)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
von: Shen, Wei, et al.
Veröffentlicht: (2023)
von: Shen, Wei, et al.
Veröffentlicht: (2023)
An Information-Theoretic Analysis of In-Context Learning
von: Jeon, Hong Jun, et al.
Veröffentlicht: (2024)
von: Jeon, Hong Jun, et al.
Veröffentlicht: (2024)
Asymptotic Behavior of Multi--Task Learning: Implicit Regularization and Double Descent Effects
von: Alrashdi, Ayed M., et al.
Veröffentlicht: (2026)
von: Alrashdi, Ayed M., et al.
Veröffentlicht: (2026)
Learning Causality for Longitudinal Data
von: Bouchattaoui, Mouad EL
Veröffentlicht: (2025)
von: Bouchattaoui, Mouad EL
Veröffentlicht: (2025)
Joint Constellation Shaping Using Gradient Descent Approach for MU-MIMO Broadcast Channel
von: Vaillant, Maxime, et al.
Veröffentlicht: (2024)
von: Vaillant, Maxime, et al.
Veröffentlicht: (2024)
Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
Learning Unknown Intervention Targets in Structural Causal Models from Heterogeneous Data
von: Yang, Yuqin, et al.
Veröffentlicht: (2023)
von: Yang, Yuqin, et al.
Veröffentlicht: (2023)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
von: Taheri, Hossein, et al.
Veröffentlicht: (2024)
von: Taheri, Hossein, et al.
Veröffentlicht: (2024)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
Causal Entropy and Information Gain for Measuring Causal Control
von: Simoes, Francisco Nunes Ferreira Quialheiro, et al.
Veröffentlicht: (2023)
von: Simoes, Francisco Nunes Ferreira Quialheiro, et al.
Veröffentlicht: (2023)
Decoupled Descent: Exact Test Error Tracking Via Approximate Message Passing
von: Lovig, Max
Veröffentlicht: (2026)
von: Lovig, Max
Veröffentlicht: (2026)
On Causality in Domain Adaptation and Semi-Supervised Learning: an Information-Theoretic Analysis for Parametric Models
von: Wu, Xuetong, et al.
Veröffentlicht: (2022)
von: Wu, Xuetong, et al.
Veröffentlicht: (2022)
MASSFormer: Mobility-Aware Spectrum Sensing using Transformer-Driven Tiered Structure
von: Janu, Dimpal, et al.
Veröffentlicht: (2024)
von: Janu, Dimpal, et al.
Veröffentlicht: (2024)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
von: Huang, Jianhao, et al.
Veröffentlicht: (2025)
von: Huang, Jianhao, et al.
Veröffentlicht: (2025)
Structured IB: Improving Information Bottleneck with Structured Feature Learning
von: Yang, Hanzhe, et al.
Veröffentlicht: (2024)
von: Yang, Hanzhe, et al.
Veröffentlicht: (2024)
The Blind Normalized Stein Variational Gradient Descent-Based Detection for Intelligent Random Access in Cellular IoT
von: Zhu, Xin, et al.
Veröffentlicht: (2024)
von: Zhu, Xin, et al.
Veröffentlicht: (2024)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
Structure Learning via Mutual Information
von: Nixon, Jeremy
Veröffentlicht: (2024)
von: Nixon, Jeremy
Veröffentlicht: (2024)
Ähnliche Einträge
-
Understanding Factual Recall in Transformers via Associative Memories
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024) -
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
von: Nichani, Eshaan, et al.
Veröffentlicht: (2023) -
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026) -
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
von: Fu, Hengyu, et al.
Veröffentlicht: (2024) -
Quantitative Bounds for Length Generalization in Transformers
von: Izzo, Zachary, et al.
Veröffentlicht: (2025)