Predicting the Emergence of Induction Heads in Language Model Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Aoyama, Tatsuya, Wilcox, Ethan Gotlieb, Schneider, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Models Grow Less Humanlike beyond Phase Transition
by: Aoyama, Tatsuya, et al.
Published: (2025)
by: Aoyama, Tatsuya, et al.
Published: (2025)
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
by: Scivetti, Wesley, et al.
Published: (2025)
by: Scivetti, Wesley, et al.
Published: (2025)
A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models
by: Yang, Xiulin, et al.
Published: (2026)
by: Yang, Xiulin, et al.
Published: (2026)
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
by: Yang, Xiulin, et al.
Published: (2025)
by: Yang, Xiulin, et al.
Published: (2025)
Looking forward: Linguistic theory and methods
by: Mansfield, John, et al.
Published: (2025)
by: Mansfield, John, et al.
Published: (2025)
Function Words as Statistical Cues for Language Learning
by: Yang, Xiulin, et al.
Published: (2026)
by: Yang, Xiulin, et al.
Published: (2026)
Dual Alignment Between Language Model Layers and Human Sentence Processing
by: Kuribayashi, Tatsuki, et al.
Published: (2026)
by: Kuribayashi, Tatsuki, et al.
Published: (2026)
Modeling Bottom-up Information Quality during Language Processing
by: Ding, Cui, et al.
Published: (2025)
by: Ding, Cui, et al.
Published: (2025)
Testing the Predictions of Surprisal Theory in 11 Languages
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)
by: Wilcox, Ethan Gotlieb, et al.
Published: (2023)
Information-Theoretic Storage Cost in Sentence Comprehension
by: Kajikawa, Kohei, et al.
Published: (2026)
by: Kajikawa, Kohei, et al.
Published: (2026)
On the Role of Context in Reading Time Prediction
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
by: Scivetti, Wesley, et al.
Published: (2026)
by: Scivetti, Wesley, et al.
Published: (2026)
On the Emergence of Induction Heads for In-Context Learning
by: Musat, Tiberiu, et al.
Published: (2025)
by: Musat, Tiberiu, et al.
Published: (2025)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Hu, Michael Y., et al.
Published: (2024)
by: Hu, Michael Y., et al.
Published: (2024)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
by: Doan, Nhi Hoai, et al.
Published: (2025)
by: Doan, Nhi Hoai, et al.
Published: (2025)
What Can String Probability Tell Us About Grammaticality?
by: Hu, Jennifer, et al.
Published: (2025)
by: Hu, Jennifer, et al.
Published: (2025)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
by: Wilcox, Ethan Gotlieb, et al.
Published: (2025)
by: Wilcox, Ethan Gotlieb, et al.
Published: (2025)
Reverse-Engineering the Reader
by: Kiegeland, Samuel, et al.
Published: (2024)
by: Kiegeland, Samuel, et al.
Published: (2024)
Cross-Domain Bilingual Lexicon Induction via Pretrained Language Models
by: Ding, Qiuyu, et al.
Published: (2025)
by: Ding, Qiuyu, et al.
Published: (2025)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Warstadt, Alex, et al.
Published: (2025)
by: Warstadt, Alex, et al.
Published: (2025)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
by: Choshen, Leshem, et al.
Published: (2026)
by: Choshen, Leshem, et al.
Published: (2026)
Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
by: Tsipidi, Eleftheria, et al.
Published: (2024)
by: Tsipidi, Eleftheria, et al.
Published: (2024)
Rethinking Associative Memory Mechanism in Induction Head
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
In-Context Learning in Speech Language Models: Analyzing the Role of Acoustic Features, Linguistic Structure, and Induction Heads
by: Pouw, Charlotte, et al.
Published: (2026)
by: Pouw, Charlotte, et al.
Published: (2026)
Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning
by: Crosbie, Joy, et al.
Published: (2024)
by: Crosbie, Joy, et al.
Published: (2024)
AI Managed Emergency Documentation with a Pretrained Model
by: Menzies, David, et al.
Published: (2024)
by: Menzies, David, et al.
Published: (2024)
Universal Response and Emergence of Induction in LLMs
by: Luick, Niclas
Published: (2024)
by: Luick, Niclas
Published: (2024)
Identifying Semantic Induction Heads to Understand In-Context Learning
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Temporal Dependencies in In-Context Learning: The Role of Induction Heads
by: Bajaj, Anooshka, et al.
Published: (2026)
by: Bajaj, Anooshka, et al.
Published: (2026)
Efficient and Flexible Topic Modeling using Pretrained Embeddings and Bag of Sentences
by: Schneider, Johannes
Published: (2023)
by: Schneider, Johannes
Published: (2023)
Pretraining Language Models Using Translationese
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
Geographic Adaptation of Pretrained Language Models
by: Hofmann, Valentin, et al.
Published: (2022)
by: Hofmann, Valentin, et al.
Published: (2022)
Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?
by: Sun, Kai, et al.
Published: (2023)
by: Sun, Kai, et al.
Published: (2023)
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
by: Prange, Jakob, et al.
Published: (2021)
by: Prange, Jakob, et al.
Published: (2021)
Speaking of Language: Reflections on Metalanguage Research in NLP
by: Schneider, Nathan, et al.
Published: (2026)
by: Schneider, Nathan, et al.
Published: (2026)
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Similar Items
-
Language Models Grow Less Humanlike beyond Phase Transition
by: Aoyama, Tatsuya, et al.
Published: (2025) -
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning
by: Scivetti, Wesley, et al.
Published: (2025) -
A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models
by: Yang, Xiulin, et al.
Published: (2026) -
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
by: Yang, Xiulin, et al.
Published: (2025) -
Looking forward: Linguistic theory and methods
by: Mansfield, John, et al.
Published: (2025)