The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Edelman, Benjamin L., Edelman, Ezra, Goel, Surbhi, Malach, Eran, Tsilivis, Nikolaos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
von: Edelman, Ezra, et al.
Veröffentlicht: (2026)
von: Edelman, Ezra, et al.
Veröffentlicht: (2026)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023)
von: Malach, Eran
Veröffentlicht: (2023)
Transcendence: Generative Models Can Outperform The Experts That Train Them
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
Weight Clipping for Robust Conformal Inference under Unbounded Covariate Shifts
von: Wang, James, et al.
Veröffentlicht: (2026)
von: Wang, James, et al.
Veröffentlicht: (2026)
Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift
von: Goel, Surbhi, et al.
Veröffentlicht: (2026)
von: Goel, Surbhi, et al.
Veröffentlicht: (2026)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
The Price of Implicit Bias in Adversarially Robust Generalization
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
Semi-Explicit Neural DAEs: Learning Long-Horizon Dynamical Systems with Algebraic Constraints
von: Pal, Avik, et al.
Veröffentlicht: (2025)
von: Pal, Avik, et al.
Veröffentlicht: (2025)
LLM Priors for ERM over Programs
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)
Matrix Calculus (for Machine Learning and Beyond)
von: Bright, Paige, et al.
Veröffentlicht: (2025)
von: Bright, Paige, et al.
Veröffentlicht: (2025)
Complexity Matters: Dynamics of Feature Learning in the Presence of Spurious Correlations
von: Qiu, GuanWen, et al.
Veröffentlicht: (2024)
von: Qiu, GuanWen, et al.
Veröffentlicht: (2024)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
von: Su, Jingtong, et al.
Veröffentlicht: (2023)
von: Su, Jingtong, et al.
Veröffentlicht: (2023)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
Why Do Transformers Fail to Forecast Time Series In-Context?
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
Markov Chain Estimation with In-Context Learning
von: Lepage, Simon, et al.
Veröffentlicht: (2025)
von: Lepage, Simon, et al.
Veröffentlicht: (2025)
Learning to Think from Multiple Thinkers
von: Joshi, Nirmit, et al.
Veröffentlicht: (2026)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2026)
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
A Theory of Learning with Autoregressive Chain of Thought
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2025)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
Distinguishing the Knowable from the Unknowable with Language Models
von: Ahdritz, Gustaf, et al.
Veröffentlicht: (2024)
von: Ahdritz, Gustaf, et al.
Veröffentlicht: (2024)
On the Geometry of Regularization in Adversarial Training: High-Dimensional Asymptotics and Generalization Bounds
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2024)
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2024)
Tolerant Algorithms for Learning with Arbitrary Covariate Shift
von: Goel, Surbhi, et al.
Veröffentlicht: (2024)
von: Goel, Surbhi, et al.
Veröffentlicht: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
Selective Induction Heads: How Transformers Select Causal Structures In Context
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2025)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
von: Bondaschi, Marco, et al.
Veröffentlicht: (2025)
What are human values, and how do we align AI to them?
von: Klingefjord, Oliver, et al.
Veröffentlicht: (2024)
von: Klingefjord, Oliver, et al.
Veröffentlicht: (2024)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
Testing Noise Assumptions of Learning Algorithms
von: Goel, Surbhi, et al.
Veröffentlicht: (2025)
von: Goel, Surbhi, et al.
Veröffentlicht: (2025)
N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs
von: Zisman, Ilya, et al.
Veröffentlicht: (2024)
von: Zisman, Ilya, et al.
Veröffentlicht: (2024)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
Feature emergence via margin maximization: case studies in algebraic tasks
von: Morwani, Depen, et al.
Veröffentlicht: (2023)
von: Morwani, Depen, et al.
Veröffentlicht: (2023)
PlaceNav: Topological Navigation through Place Recognition
von: Suomela, Lauri, et al.
Veröffentlicht: (2023)
von: Suomela, Lauri, et al.
Veröffentlicht: (2023)
A New Perspective on Shampoo's Preconditioner
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
von: Sabbaghi, Mahdi, et al.
Veröffentlicht: (2024)
von: Sabbaghi, Mahdi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
von: Edelman, Ezra, et al.
Veröffentlicht: (2026) -
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025) -
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025) -
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023) -
Transcendence: Generative Models Can Outperform The Experts That Train Them
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)