Contextualize-then-Aggregate: Circuits for In-Context Learning in Gemma-2 2B
Fuente:
arXiv
Saved in:
| Main Authors: | Bakalova, Aleksandra, Veitsman, Yana, Huang, Xinting, Hahn, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
The Expressive Capacity of State Space Models: A Formal Language Perspective
by: Sarrof, Yash, et al.
Published: (2024)
by: Sarrof, Yash, et al.
Published: (2024)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026)
by: Huang, Xinting, et al.
Published: (2026)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
by: Huang, Xinting, et al.
Published: (2025)
by: Huang, Xinting, et al.
Published: (2025)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
Recent Advancements and Challenges of Turkic Central Asian Language Processing
by: Veitsman, Yana, et al.
Published: (2024)
by: Veitsman, Yana, et al.
Published: (2024)
How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
by: Wang, Entang, et al.
Published: (2026)
by: Wang, Entang, et al.
Published: (2026)
ShieldGemma: Generative AI Content Moderation Based on Gemma
by: Zeng, Wenjun, et al.
Published: (2024)
by: Zeng, Wenjun, et al.
Published: (2024)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
by: Veitsman, Yana, et al.
Published: (2026)
by: Veitsman, Yana, et al.
Published: (2026)
Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
by: Huang, Yiran, et al.
Published: (2026)
by: Huang, Yiran, et al.
Published: (2026)
Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge
by: Šliogeris, Vytenis, et al.
Published: (2025)
by: Šliogeris, Vytenis, et al.
Published: (2025)
Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B
by: Bektursun, Abay
Published: (2026)
by: Bektursun, Abay
Published: (2026)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
by: Hsu, Aliyah R., et al.
Published: (2024)
by: Hsu, Aliyah R., et al.
Published: (2024)
TxGemma: Efficient and Agentic LLMs for Therapeutics
by: Wang, Eric, et al.
Published: (2025)
by: Wang, Eric, et al.
Published: (2025)
Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation
by: Zhang, Biao, et al.
Published: (2025)
by: Zhang, Biao, et al.
Published: (2025)
PaliGemma: A versatile 3B VLM for transfer
by: Beyer, Lucas, et al.
Published: (2024)
by: Beyer, Lucas, et al.
Published: (2024)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
by: Botev, Aleksandar, et al.
Published: (2024)
by: Botev, Aleksandar, et al.
Published: (2024)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025)
by: Africa, David Demitri
Published: (2025)
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
by: Kiulian, Artur, et al.
Published: (2024)
by: Kiulian, Artur, et al.
Published: (2024)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
by: Cheng, Yun, et al.
Published: (2026)
by: Cheng, Yun, et al.
Published: (2026)
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
by: Tiwari, Utkarsh, et al.
Published: (2025)
by: Tiwari, Utkarsh, et al.
Published: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
by: Anand, Nikhil, et al.
Published: (2026)
by: Anand, Nikhil, et al.
Published: (2026)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
by: Morlat, Geoffroy, et al.
Published: (2025)
by: Morlat, Geoffroy, et al.
Published: (2025)
Many-Shot In-Context Learning
by: Agarwal, Rishabh, et al.
Published: (2024)
by: Agarwal, Rishabh, et al.
Published: (2024)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
by: Li, Chengye, et al.
Published: (2025)
by: Li, Chengye, et al.
Published: (2025)
Prompt Optimization via Adversarial In-Context Learning
by: Do, Xuan Long, et al.
Published: (2023)
by: Do, Xuan Long, et al.
Published: (2023)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
by: Gur, Izzeddin, et al.
Published: (2023)
by: Gur, Izzeddin, et al.
Published: (2023)
SkillAggregation: Reference-free LLM-Dependent Aggregation
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
LLM Circuit Analyses Are Consistent Across Training and Scale
by: Tigges, Curt, et al.
Published: (2024)
by: Tigges, Curt, et al.
Published: (2024)
Semantic Captioning: Benchmark Dataset and Graph-Aware Few-Shot In-Context Learning for SQL2Text
by: Al-Lawati, Ali, et al.
Published: (2025)
by: Al-Lawati, Ali, et al.
Published: (2025)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
by: Zeng, Zhiyuan, et al.
Published: (2026)
by: Zeng, Zhiyuan, et al.
Published: (2026)
Judge Circuits
by: Feldhus, Nils, et al.
Published: (2026)
by: Feldhus, Nils, et al.
Published: (2026)
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Similar Items
-
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025) -
The Expressive Capacity of State Space Models: A Formal Language Perspective
by: Sarrof, Yash, et al.
Published: (2024) -
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026) -
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
by: Huang, Xinting, et al.
Published: (2025) -
InversionView: A General-Purpose Method for Reading Information from Neural Activations
by: Huang, Xinting, et al.
Published: (2024)