What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
Fuente:
arXiv
Saved in:
| Main Authors: | Ekbote, Chanakya, Bondaschi, Marco, Rajaraman, Nived, Lee, Jason D., Gastpar, Michael, Makkuva, Ashok Vardhan, Liang, Paul Pu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
by: Nagle, Alliot, et al.
Published: (2024)
by: Nagle, Alliot, et al.
Published: (2024)
LASER: Linear Compression in Wireless Distributed Optimization
by: Makkuva, Ashok Vardhan, et al.
Published: (2023)
by: Makkuva, Ashok Vardhan, et al.
Published: (2023)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
by: Nagle, Alliot, et al.
Published: (2026)
by: Nagle, Alliot, et al.
Published: (2026)
Batch Universal Prediction
by: Bondaschi, Marco, et al.
Published: (2024)
by: Bondaschi, Marco, et al.
Published: (2024)
The Conditional Regret-Capacity Theorem for Batch Universal Prediction
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Alpha-NML Universal Predictors
by: Bondaschi, Marco, et al.
Published: (2022)
by: Bondaschi, Marco, et al.
Published: (2022)
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
Understanding the Emergence of Multimodal Representation Alignment
by: Tjandrasuwita, Megan, et al.
Published: (2025)
by: Tjandrasuwita, Megan, et al.
Published: (2025)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
by: Rajaraman, Nived, et al.
Published: (2026)
by: Rajaraman, Nived, et al.
Published: (2026)
Batch Normalization Decomposed
by: Nachum, Ido, et al.
Published: (2024)
by: Nachum, Ido, et al.
Published: (2024)
Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum
by: Rajaraman, Nived, et al.
Published: (2026)
by: Rajaraman, Nived, et al.
Published: (2026)
SCATR: Simple Calibrated Test-Time Ranking
by: Shyamal, Divya, et al.
Published: (2026)
by: Shyamal, Divya, et al.
Published: (2026)
SAMTok: Representing Any Mask with Two Words
by: Zhou, Yikang, et al.
Published: (2026)
by: Zhou, Yikang, et al.
Published: (2026)
Markov Chains Approximate Message Passing
by: Rajaraman, Amit, et al.
Published: (2025)
by: Rajaraman, Amit, et al.
Published: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024)
by: Edelman, Benjamin L., et al.
Published: (2024)
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Computational Intractability of Strategizing against Online Learners
by: Assos, Angelos, et al.
Published: (2025)
by: Assos, Angelos, et al.
Published: (2025)
Interleaved Head Attention
by: Duvvuri, Sai Surya, et al.
Published: (2026)
by: Duvvuri, Sai Surya, et al.
Published: (2026)
TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models
by: Lee, Jaewoo, et al.
Published: (2025)
by: Lee, Jaewoo, et al.
Published: (2025)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
by: Kidder, William, et al.
Published: (2024)
by: Kidder, William, et al.
Published: (2024)
Any Labor Union Can Represent Any Unit
Published: (2024)
Published: (2024)
Can a Higher Order Markov Chain Be Treated as a First Order Markov Chain?
by: Xu, Jianhong
Published: (2025)
by: Xu, Jianhong
Published: (2025)
Focus Managers on What They Can Say, Not What They Cannot
Published: (2024)
Published: (2024)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023)
by: Rajaraman, Nived, et al.
Published: (2023)
The Committee on Accreditation: What It Can and Cannot Do.
by: Kimmel, Margaret Mary
Published: (1987)
by: Kimmel, Margaret Mary
Published: (1987)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Probing Minimalist Phase Structure in LLMs: What Universal Dependencies Cannot Represent
by: Chen, Yuanhao, et al.
Published: (2026)
by: Chen, Yuanhao, et al.
Published: (2026)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025)
by: Flamich, Gergely, et al.
Published: (2025)
Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
by: Kumar, Dibyakanti, et al.
Published: (2025)
by: Kumar, Dibyakanti, et al.
Published: (2025)
Parental Presence at Induction—Are Two Parents Better Than One?
by: Nicole Almenrader, et al.
Published: (2025)
by: Nicole Almenrader, et al.
Published: (2025)
The Space Complexity of Learning-Unlearning Algorithms
by: Cherapanamjeri, Yeshwanth, et al.
Published: (2025)
by: Cherapanamjeri, Yeshwanth, et al.
Published: (2025)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
Locally Stationary Distributions: A Framework for Analyzing Slow-Mixing Markov Chains
by: Liu, Kuikui, et al.
Published: (2024)
by: Liu, Kuikui, et al.
Published: (2024)
Similar Items
-
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024) -
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025) -
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
by: Makkuva, Ashok Vardhan, et al.
Published: (2024) -
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024) -
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
by: Nagle, Alliot, et al.
Published: (2024)