Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Doan, Nhi Hoai, Hiraoka, Tatsuya, Inui, Kentaro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
by: Tsuji, Kohei, et al.
Published: (2025)
by: Tsuji, Kohei, et al.
Published: (2025)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
Identifying Semantic Induction Heads to Understand In-Context Learning
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Understanding In-Context Learning from Repetitions
by: Yan, Jianhao, et al.
Published: (2023)
by: Yan, Jianhao, et al.
Published: (2023)
Tokenization Preference for Human and Machine Learning Model: An Annotation Study
by: Hiraoka, Tatsuya, et al.
Published: (2023)
by: Hiraoka, Tatsuya, et al.
Published: (2023)
On the Emergence of Induction Heads for In-Context Learning
by: Musat, Tiberiu, et al.
Published: (2025)
by: Musat, Tiberiu, et al.
Published: (2025)
Predicting the Emergence of Induction Heads in Language Model Pretraining
by: Aoyama, Tatsuya, et al.
Published: (2025)
by: Aoyama, Tatsuya, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Monotonic Representation of Numeric Properties in Language Models
by: Heinzerling, Benjamin, et al.
Published: (2024)
by: Heinzerling, Benjamin, et al.
Published: (2024)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Temporal Dependencies in In-Context Learning: The Role of Induction Heads
by: Bajaj, Anooshka, et al.
Published: (2026)
by: Bajaj, Anooshka, et al.
Published: (2026)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025)
by: Nwadike, Munachiso, et al.
Published: (2025)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
SmallPlan: Leverage Small Language Models for Sequential Path Planning with Simulation-Powered, LLM-Guided Distillation
by: Pham, Quang P. M., et al.
Published: (2025)
by: Pham, Quang P. M., et al.
Published: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Do LLMs Need to Think in One Language? Correlation between Latent Language and Task Performance
by: Ozaki, Shintaro, et al.
Published: (2025)
by: Ozaki, Shintaro, et al.
Published: (2025)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
Cell-Based Representation of Relational Binding in Language Models
by: Dai, Qin, et al.
Published: (2026)
by: Dai, Qin, et al.
Published: (2026)
Representational Analysis of Binding in Language Models
by: Dai, Qin, et al.
Published: (2024)
by: Dai, Qin, et al.
Published: (2024)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference
by: Kamoda, Go, et al.
Published: (2025)
by: Kamoda, Go, et al.
Published: (2025)
Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted
by: Shimmei, Machi, et al.
Published: (2025)
by: Shimmei, Machi, et al.
Published: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
Syntactic Learnability of Echo State Neural Language Models at Scale
by: Ueda, Ryo, et al.
Published: (2025)
by: Ueda, Ryo, et al.
Published: (2025)
TopK Language Models
by: Takahashi, Ryosuke, et al.
Published: (2025)
by: Takahashi, Ryosuke, et al.
Published: (2025)
Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning
by: Crosbie, Joy, et al.
Published: (2024)
by: Crosbie, Joy, et al.
Published: (2024)
J-UniMorph: Japanese Morphological Annotation through the Universal Feature Schema
by: Matsuzaki, Kosuke, et al.
Published: (2024)
by: Matsuzaki, Kosuke, et al.
Published: (2024)
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems
by: Sato, Shiki, et al.
Published: (2024)
by: Sato, Shiki, et al.
Published: (2024)
Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
by: Kobayashi, Goro, et al.
Published: (2023)
by: Kobayashi, Goro, et al.
Published: (2023)
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization
by: Tsuji, Kohei, et al.
Published: (2024)
by: Tsuji, Kohei, et al.
Published: (2024)
In-Context Learning in Speech Language Models: Analyzing the Role of Acoustic Features, Linguistic Structure, and Induction Heads
by: Pouw, Charlotte, et al.
Published: (2026)
by: Pouw, Charlotte, et al.
Published: (2026)
FinchGPT: a Transformer based language model for birdsong analysis
by: Kobayashi, Kosei, et al.
Published: (2025)
by: Kobayashi, Kosei, et al.
Published: (2025)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
by: Chevi, Rendi, et al.
Published: (2025)
by: Chevi, Rendi, et al.
Published: (2025)
Neuron-based Personality Trait Induction in Large Language Models
by: Deng, Jia, et al.
Published: (2024)
by: Deng, Jia, et al.
Published: (2024)
Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)
by: Airlangga, Muhammad Cendekia, et al.
Published: (2025)
Similar Items
-
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024) -
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025) -
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026) -
Investigating Neurons and Heads in Transformer-based LLMs for Typographical Errors
by: Tsuji, Kohei, et al.
Published: (2025) -
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)