What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Aaditya K., Moskovitz, Ted, Hill, Felix, Chan, Stephanie C. Y., Saxe, Andrew M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
di: Singh, Aaditya K., et al.
Pubblicazione: (2025)
di: Singh, Aaditya K., et al.
Pubblicazione: (2025)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
di: Dragutinović, Sara, et al.
Pubblicazione: (2025)
di: Dragutinović, Sara, et al.
Pubblicazione: (2025)
The broader spectrum of in-context learning
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
HARP: A challenging human-annotated math reasoning benchmark
di: Yue, Albert S., et al.
Pubblicazione: (2024)
di: Yue, Albert S., et al.
Pubblicazione: (2024)
Training Dynamics of In-Context Learning in Linear Attention
di: Zhang, Yedi, et al.
Pubblicazione: (2025)
di: Zhang, Yedi, et al.
Pubblicazione: (2025)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
di: Lee, Jin Hwa, et al.
Pubblicazione: (2025)
di: Lee, Jin Hwa, et al.
Pubblicazione: (2025)
On the generalization of language models from in-context learning and finetuning: a controlled study
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025)
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025)
Biological Engineering: What does it mean? Where does it (need to) go?
di: Nuber, Ulrike A., et al.
Pubblicazione: (2025)
di: Nuber, Ulrike A., et al.
Pubblicazione: (2025)
AutoLTS: Automating Cycling Stress Assessment via Contrastive Learning and Spatial Post-processing
di: Lin, Bo, et al.
Pubblicazione: (2023)
di: Lin, Bo, et al.
Pubblicazione: (2023)
Machine Learning-Augmented Optimization of Large Bilevel and Two-stage Stochastic Programs: Application to Cycling Network Design
di: Chan, Timothy C. Y., et al.
Pubblicazione: (2022)
di: Chan, Timothy C. Y., et al.
Pubblicazione: (2022)
OPEC and IEA go head‐to‐head over demand and investment
Pubblicazione: (2024)
Pubblicazione: (2024)
Learned feature representations are biased by complexity, learning order, position, and more
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024)
Relational reasoning and inductive bias in transformers and large language models
di: Geerts, Jesse, et al.
Pubblicazione: (2025)
di: Geerts, Jesse, et al.
Pubblicazione: (2025)
Locating acts of mechanistic reasoning in student team conversations with mechanistic machine learning
di: Gili, Kaitlin, et al.
Pubblicazione: (2026)
di: Gili, Kaitlin, et al.
Pubblicazione: (2026)
The in-context inductive biases of vision-language models differ across modalities
di: Allen, Kelsey, et al.
Pubblicazione: (2025)
di: Allen, Kelsey, et al.
Pubblicazione: (2025)
Scaling sparse feature circuit finding for in-context learning
di: Kharlapenko, Dmitrii, et al.
Pubblicazione: (2025)
di: Kharlapenko, Dmitrii, et al.
Pubblicazione: (2025)
Calibrated Peer Review Assignments in Science Courses: Are They Designed to Promote Critical Thinking and Writing Skills?
di: Reynolds, Julie, et al.
Pubblicazione: (2008)
di: Reynolds, Julie, et al.
Pubblicazione: (2008)
Does learning the right latent variables necessarily improve in-context learning?
di: Mittal, Sarthak, et al.
Pubblicazione: (2024)
di: Mittal, Sarthak, et al.
Pubblicazione: (2024)
An evolutionary perspective on modes of learning in Transformers
di: Ku, Alexander Y., et al.
Pubblicazione: (2025)
di: Ku, Alexander Y., et al.
Pubblicazione: (2025)
Turning mechanistic models into forecasters by using machine learning
di: Chakraborty, Amit K., et al.
Pubblicazione: (2026)
di: Chakraborty, Amit K., et al.
Pubblicazione: (2026)
Early learning of the optimal constant solution in neural networks and humans
di: Rubruck, Jirko, et al.
Pubblicazione: (2024)
di: Rubruck, Jirko, et al.
Pubblicazione: (2024)
Kinetic inductance coupling for circuit QED with spins
di: Günzler, S., et al.
Pubblicazione: (2025)
di: Günzler, S., et al.
Pubblicazione: (2025)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
Meta predictive learning model of languages in neural circuits
di: Li, Chan, et al.
Pubblicazione: (2023)
di: Li, Chan, et al.
Pubblicazione: (2023)
What is an inductive mean?
di: Nielsen, Frank
Pubblicazione: (2024)
di: Nielsen, Frank
Pubblicazione: (2024)
DNA analyst's refusal to answer an activity level question did not violate the defendant's right to confrontation
di: Ted R. Hunt
Pubblicazione: (2026)
di: Ted R. Hunt
Pubblicazione: (2026)
Sexualidades disidentes en la narrativa española reciente: políticas sexuales, cine y subversión en Mae West y yo (2011) de Eduardo Mendicutti
di: Facundo Saxe
Pubblicazione: (2016)
di: Facundo Saxe
Pubblicazione: (2016)
Isabel Sarli, Alba Mujica y Sabaleros (1959). Del archivo de sentimientos al multiverso de disidencias sexuales
di: Facundo Saxe
Pubblicazione: (2025)
di: Facundo Saxe
Pubblicazione: (2025)
Shapiro effect in inductive quantum circuits with charge discreteness
di: Kristopher Chandía Valenzuela
Pubblicazione: (2006)
di: Kristopher Chandía Valenzuela
Pubblicazione: (2006)
One-layer transformers fail to solve the induction heads task
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
What makes child workers go to school?: a case study from West Bengal
di: Manoranjan Pal, et al.
Pubblicazione: (2011)
di: Manoranjan Pal, et al.
Pubblicazione: (2011)
Targeting VMPFC‐amygdala circuit with TMS in substance use disorder: A mechanistic framework
di: Ghazaleh Soleimani, et al.
Pubblicazione: (2025)
di: Ghazaleh Soleimani, et al.
Pubblicazione: (2025)
What for…A CAE Past‐Presidential Address
di: Edmund ‘Ted’ Hamann
Pubblicazione: (2025)
di: Edmund ‘Ted’ Hamann
Pubblicazione: (2025)
Language models show human-like content effects on reasoning tasks
di: Dasgupta, Ishita, et al.
Pubblicazione: (2022)
di: Dasgupta, Ishita, et al.
Pubblicazione: (2022)
When Representations Align: Universality in Representation Learning Dynamics
di: van Rossem, Loek, et al.
Pubblicazione: (2024)
di: van Rossem, Loek, et al.
Pubblicazione: (2024)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
di: van Rossem, Loek, et al.
Pubblicazione: (2025)
di: van Rossem, Loek, et al.
Pubblicazione: (2025)
Human rights in diverse education contexts
di: Beckmann, Johan, et al.
Pubblicazione: (2020)
di: Beckmann, Johan, et al.
Pubblicazione: (2020)
A Machine Learning Approach to Meteor Classification
di: Hemmelgarn, Samantha, et al.
Pubblicazione: (2026)
di: Hemmelgarn, Samantha, et al.
Pubblicazione: (2026)
Respiratory system compliance during anesthesia induction and postoperative mechanical ventilation needs: An observational study
di: Yukiko Yamazaki, et al.
Pubblicazione: (2024)
di: Yukiko Yamazaki, et al.
Pubblicazione: (2024)
Melatonin regulates endoplasmic reticulum stress in diverse pathophysiological contexts: A comprehensive mechanistic review
di: Luiz Gustavo de Almeida Chuffa, et al.
Pubblicazione: (2024)
di: Luiz Gustavo de Almeida Chuffa, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
di: Singh, Aaditya K., et al.
Pubblicazione: (2025) -
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
di: Dragutinović, Sara, et al.
Pubblicazione: (2025) -
The broader spectrum of in-context learning
di: Lampinen, Andrew Kyle, et al.
Pubblicazione: (2024) -
HARP: A challenging human-annotated math reasoning benchmark
di: Yue, Albert S., et al.
Pubblicazione: (2024) -
Training Dynamics of In-Context Learning in Linear Attention
di: Zhang, Yedi, et al.
Pubblicazione: (2025)