Clustering Head: A Visual Case Study of the Training Dynamics in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Odonnat, Ambroise, Bouaziz, Wassim, Cabannes, Vivien |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Easing Optimization Paths: a Circuit Perspective
by: Odonnat, Ambroise, et al.
Published: (2025)
by: Odonnat, Ambroise, et al.
Published: (2025)
Provable Benefits of In-Tool Learning for Large Language Models
by: Houliston, Sam, et al.
Published: (2025)
by: Houliston, Sam, et al.
Published: (2025)
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Leveraging Ensemble Diversity for Robust Self-Training in the Presence of Sample Selection Bias
by: Odonnat, Ambroise, et al.
Published: (2023)
by: Odonnat, Ambroise, et al.
Published: (2023)
Optimal Self-Consistency for Efficient Reasoning with Large Language Models
by: Feng, Austin, et al.
Published: (2025)
by: Feng, Austin, et al.
Published: (2025)
Vision Transformer Finetuning Benefits from Non-Smooth Components
by: Odonnat, Ambroise, et al.
Published: (2026)
by: Odonnat, Ambroise, et al.
Published: (2026)
Touring sampling with pushforward maps
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
The Galerkin method beats Graph-Based Approaches for Spectral Algorithms
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026)
by: Arnal, Charles, et al.
Published: (2026)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention
by: Ilbert, Romain, et al.
Published: (2024)
by: Ilbert, Romain, et al.
Published: (2024)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
MANO: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution Shifts
by: Xie, Renchunzi, et al.
Published: (2024)
by: Xie, Renchunzi, et al.
Published: (2024)
Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift
by: Xie, Renchunzi, et al.
Published: (2024)
by: Xie, Renchunzi, et al.
Published: (2024)
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2025)
by: Bouaziz, Wassim, et al.
Published: (2025)
Layer by layer, module by module: Choose both for optimal OOD probing of ViT
by: Odonnat, Ambroise, et al.
Published: (2026)
by: Odonnat, Ambroise, et al.
Published: (2026)
Learning with Hidden Factorial Structure
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Data Taggants: Dataset Ownership Verification via Harmless Targeted Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2024)
by: Bouaziz, Wassim, et al.
Published: (2024)
Inverting Gradient Attacks Makes Powerful Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2024)
by: Bouaziz, Wassim, et al.
Published: (2024)
Targeted Data Poisoning for Black-Box Audio Datasets Ownership Verification
by: Bouaziz, Wassim, et al.
Published: (2025)
by: Bouaziz, Wassim, et al.
Published: (2025)
Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series Forecasting
by: Ilbert, Romain, et al.
Published: (2024)
by: Ilbert, Romain, et al.
Published: (2024)
Scaling Laws for Associative Memories
by: Cabannes, Vivien, et al.
Published: (2023)
by: Cabannes, Vivien, et al.
Published: (2023)
Large Language Models as Markov Chains
by: Zekri, Oussama, et al.
Published: (2024)
by: Zekri, Oussama, et al.
Published: (2024)
Mode Estimation with Partial Feedback
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
Zero-shot Model-based Reinforcement Learning using Large Language Models
by: Benechehab, Abdelhakim, et al.
Published: (2024)
by: Benechehab, Abdelhakim, et al.
Published: (2024)
SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities
by: Lalou, Yanis, et al.
Published: (2024)
by: Lalou, Yanis, et al.
Published: (2024)
CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data
by: Xie, Shifeng, et al.
Published: (2025)
by: Xie, Shifeng, et al.
Published: (2025)
A General Framework for Joint Multi-State Models
by: Laplante, Félix, et al.
Published: (2025)
by: Laplante, Félix, et al.
Published: (2025)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Weather-Aware Transformer for Real-Time Route Optimization in Drone-as-a-Service Operations
by: Mohamed, Kamal, et al.
Published: (2026)
by: Mohamed, Kamal, et al.
Published: (2026)
On Monotonicity in AI Alignment
by: Bareilles, Gilles, et al.
Published: (2025)
by: Bareilles, Gilles, et al.
Published: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
$\mathbb{X}$-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs
by: Sobal, Vlad, et al.
Published: (2024)
by: Sobal, Vlad, et al.
Published: (2024)
Clustering and Alignment: Understanding the Training Dynamics in Modular Addition
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Enhancing Transformer Training Efficiency with Dynamic Dropout
by: Yan, Hanrui, et al.
Published: (2024)
by: Yan, Hanrui, et al.
Published: (2024)
NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
by: Grooten, Bram, et al.
Published: (2025)
by: Grooten, Bram, et al.
Published: (2025)
Inference of Multiscale Gaussian Graphical Model
by: Sanou, Do Edmond, et al.
Published: (2022)
by: Sanou, Do Edmond, et al.
Published: (2022)
Similar Items
-
Easing Optimization Paths: a Circuit Perspective
by: Odonnat, Ambroise, et al.
Published: (2025) -
Provable Benefits of In-Tool Learning for Large Language Models
by: Houliston, Sam, et al.
Published: (2025) -
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024) -
Leveraging Ensemble Diversity for Robust Self-Training in the Presence of Sample Selection Bias
by: Odonnat, Ambroise, et al.
Published: (2023) -
Optimal Self-Consistency for Efficient Reasoning with Large Language Models
by: Feng, Austin, et al.
Published: (2025)