Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Crosbie, Joy, Shutova, Ekaterina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
A framework for annotating and modelling intentions behind metaphor use
di: Michelli, Gianluca, et al.
Pubblicazione: (2024)
di: Michelli, Gianluca, et al.
Pubblicazione: (2024)
Density Matrices for Metaphor Understanding
di: Owers, Jay, et al.
Pubblicazione: (2024)
di: Owers, Jay, et al.
Pubblicazione: (2024)
How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
di: Choenni, Rochelle, et al.
Pubblicazione: (2023)
di: Choenni, Rochelle, et al.
Pubblicazione: (2023)
Learning New Tasks from a Few Examples with Soft-Label Prototypes
di: Singh, Avyav Kumar, et al.
Pubblicazione: (2022)
di: Singh, Avyav Kumar, et al.
Pubblicazione: (2022)
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
di: Verhoeven, Ivo, et al.
Pubblicazione: (2024)
di: Verhoeven, Ivo, et al.
Pubblicazione: (2024)
Metaphor Understanding Challenge Dataset for LLMs
di: Tong, Xiaoyu, et al.
Pubblicazione: (2024)
di: Tong, Xiaoyu, et al.
Pubblicazione: (2024)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
di: Leidinger, Alina, et al.
Pubblicazione: (2024)
di: Leidinger, Alina, et al.
Pubblicazione: (2024)
Rethinking Associative Memory Mechanism in Induction Head
di: Wang, Shuo, et al.
Pubblicazione: (2024)
di: Wang, Shuo, et al.
Pubblicazione: (2024)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
di: Choenni, Rochelle, et al.
Pubblicazione: (2024)
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
di: Rajaee, Sara, et al.
Pubblicazione: (2025)
di: Rajaee, Sara, et al.
Pubblicazione: (2025)
On the Emergence of Induction Heads for In-Context Learning
di: Musat, Tiberiu, et al.
Pubblicazione: (2025)
di: Musat, Tiberiu, et al.
Pubblicazione: (2025)
The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
di: Shani, Chen, et al.
Pubblicazione: (2026)
di: Shani, Chen, et al.
Pubblicazione: (2026)
Cross-modal Information Flow in Multimodal Large Language Models
di: Zhang, Zhi, et al.
Pubblicazione: (2024)
di: Zhang, Zhi, et al.
Pubblicazione: (2024)
NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
di: Zhang, Zhi, et al.
Pubblicazione: (2025)
di: Zhang, Zhi, et al.
Pubblicazione: (2025)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
di: Zhang, Zhi, et al.
Pubblicazione: (2023)
di: Zhang, Zhi, et al.
Pubblicazione: (2023)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
di: Doan, Nhi Hoai, et al.
Pubblicazione: (2025)
di: Doan, Nhi Hoai, et al.
Pubblicazione: (2025)
Hummus: A Dataset of Humorous Multimodal Metaphor Use
di: Tong, Xiaoyu, et al.
Pubblicazione: (2025)
di: Tong, Xiaoyu, et al.
Pubblicazione: (2025)
Identifying Semantic Induction Heads to Understand In-Context Learning
di: Ren, Jie, et al.
Pubblicazione: (2024)
di: Ren, Jie, et al.
Pubblicazione: (2024)
Temporal Dependencies in In-Context Learning: The Role of Induction Heads
di: Bajaj, Anooshka, et al.
Pubblicazione: (2026)
di: Bajaj, Anooshka, et al.
Pubblicazione: (2026)
Predicting the Emergence of Induction Heads in Language Model Pretraining
di: Aoyama, Tatsuya, et al.
Pubblicazione: (2025)
di: Aoyama, Tatsuya, et al.
Pubblicazione: (2025)
A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
di: Verhoeven, Ivo, et al.
Pubblicazione: (2024)
di: Verhoeven, Ivo, et al.
Pubblicazione: (2024)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
di: Chen, Siyu, et al.
Pubblicazione: (2024)
di: Chen, Siyu, et al.
Pubblicazione: (2024)
In-Context Learning in Speech Language Models: Analyzing the Role of Acoustic Features, Linguistic Structure, and Induction Heads
di: Pouw, Charlotte, et al.
Pubblicazione: (2026)
di: Pouw, Charlotte, et al.
Pubblicazione: (2026)
Interpretable Next-token Prediction via the Generalized Induction Head
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
di: Wang, Shuxun, et al.
Pubblicazione: (2025)
di: Wang, Shuxun, et al.
Pubblicazione: (2025)
Evaluating Memory Structure in LLM Agents
di: Shutova, Alina, et al.
Pubblicazione: (2026)
di: Shutova, Alina, et al.
Pubblicazione: (2026)
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
di: Sabry, Mohammed, et al.
Pubblicazione: (2025)
di: Sabry, Mohammed, et al.
Pubblicazione: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
di: Jian, Yichang, et al.
Pubblicazione: (2026)
di: Jian, Yichang, et al.
Pubblicazione: (2026)
Prompt-MII: Meta-Learning Instruction Induction for LLMs
di: Xiao, Emily, et al.
Pubblicazione: (2025)
di: Xiao, Emily, et al.
Pubblicazione: (2025)
Learning Evidence of Depression Symptoms via Prompt Induction
di: Bao, Eliseo, et al.
Pubblicazione: (2026)
di: Bao, Eliseo, et al.
Pubblicazione: (2026)
Mechanism of Task-oriented Information Removal in In-context Learning
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Going Beyond Word Matching: Syntax Improves In-context Example Selection for Machine Translation
di: Tang, Chenming, et al.
Pubblicazione: (2024)
di: Tang, Chenming, et al.
Pubblicazione: (2024)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
di: Shelmanov, Artem, et al.
Pubblicazione: (2025)
di: Shelmanov, Artem, et al.
Pubblicazione: (2025)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
Semi-Supervised Learning for Bilingual Lexicon Induction
di: Garnier, Paul, et al.
Pubblicazione: (2024)
di: Garnier, Paul, et al.
Pubblicazione: (2024)
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
di: Vazhentsev, Artem, et al.
Pubblicazione: (2025)
di: Vazhentsev, Artem, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
di: Choenni, Rochelle, et al.
Pubblicazione: (2024) -
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
di: Choenni, Rochelle, et al.
Pubblicazione: (2024) -
A framework for annotating and modelling intentions behind metaphor use
di: Michelli, Gianluca, et al.
Pubblicazione: (2024) -
Density Matrices for Metaphor Understanding
di: Owers, Jay, et al.
Pubblicazione: (2024) -
How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning
di: Choenni, Rochelle, et al.
Pubblicazione: (2023)