Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
Fuente:
arXiv
Guardado en:
| Autores principales: | Sakamoto, Keitaro, Sato, Issei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
Benign Overfitting in Token Selection of Attention Mechanism
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
por: Sakamoto, Keitaro, et al.
Publicado: (2024)
Multiplicative Logit Adjustment Approximates Neural-Collapse-Aware Decision Boundary Adjustment
por: Hasegawa, Naoya, et al.
Publicado: (2024)
por: Hasegawa, Naoya, et al.
Publicado: (2024)
Explaining Grokking in Transformers through the Lens of Inductive Bias
por: Singh, Jaisidh, et al.
Publicado: (2026)
por: Singh, Jaisidh, et al.
Publicado: (2026)
Can Test-time Computation Mitigate Reproduction Bias in Neural Symbolic Regression?
por: Sato, Shun, et al.
Publicado: (2025)
por: Sato, Shun, et al.
Publicado: (2025)
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
por: Han, Ting, et al.
Publicado: (2025)
por: Han, Ting, et al.
Publicado: (2025)
Grokking Explained: A Statistical Phenomenon
por: Carvalho, Breno W., et al.
Publicado: (2025)
por: Carvalho, Breno W., et al.
Publicado: (2025)
Understanding Generalization in Physics Informed Models through Affine Variety Dimensions
por: Koshizuka, Takeshi, et al.
Publicado: (2025)
por: Koshizuka, Takeshi, et al.
Publicado: (2025)
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
por: Nam, Yoonsoo, et al.
Publicado: (2025)
por: Nam, Yoonsoo, et al.
Publicado: (2025)
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
por: Prakash, Hari K, et al.
Publicado: (2026)
por: Prakash, Hari K, et al.
Publicado: (2026)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
Exploring Weight Balancing on Long-Tailed Recognition Problem
por: Hasegawa, Naoya, et al.
Publicado: (2023)
por: Hasegawa, Naoya, et al.
Publicado: (2023)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
por: Xu, Kevin, et al.
Publicado: (2024)
por: Xu, Kevin, et al.
Publicado: (2024)
Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
por: Tomihari, Akiyoshi, et al.
Publicado: (2024)
por: Tomihari, Akiyoshi, et al.
Publicado: (2024)
Top-Down Bayesian Posterior Sampling for Sum-Product Networks
por: Yokoi, Soma, et al.
Publicado: (2024)
por: Yokoi, Soma, et al.
Publicado: (2024)
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
por: Prakash, Hari K., et al.
Publicado: (2025)
por: Prakash, Hari K., et al.
Publicado: (2025)
Understanding the Expressivity and Trainability of Fourier Neural Operator: A Mean-Field Perspective
por: Koshizuka, Takeshi, et al.
Publicado: (2023)
por: Koshizuka, Takeshi, et al.
Publicado: (2023)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
por: Xu, Kevin, et al.
Publicado: (2025)
por: Xu, Kevin, et al.
Publicado: (2025)
Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
por: Fujikawa, Shota, et al.
Publicado: (2026)
por: Fujikawa, Shota, et al.
Publicado: (2026)
Rethinking Associative Memory Mechanism in Induction Head
por: Wang, Shuo, et al.
Publicado: (2024)
por: Wang, Shuo, et al.
Publicado: (2024)
On the Optimal Memorization Capacity of Transformers
por: Kajitsuka, Tokio, et al.
Publicado: (2024)
por: Kajitsuka, Tokio, et al.
Publicado: (2024)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
por: Tanaka, Yuto, et al.
Publicado: (2026)
por: Tanaka, Yuto, et al.
Publicado: (2026)
To Grok Grokking: Provable Grokking in Ridge Regression
por: Xu, Mingyue, et al.
Publicado: (2026)
por: Xu, Mingyue, et al.
Publicado: (2026)
Mutual Information Collapse Explains Disentanglement Failure in $β$-VAEs
por: Vu, Minh, et al.
Publicado: (2026)
por: Vu, Minh, et al.
Publicado: (2026)
A Formal Comparison Between Chain of Thought and Latent Thought
por: Xu, Kevin, et al.
Publicado: (2025)
por: Xu, Kevin, et al.
Publicado: (2025)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
por: Tian, Yuandong
Publicado: (2025)
por: Tian, Yuandong
Publicado: (2025)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
por: Zhou, Xinyu, et al.
Publicado: (2025)
por: Zhou, Xinyu, et al.
Publicado: (2025)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
por: Sakamoto, Keitaro, et al.
Publicado: (2026)
por: Sakamoto, Keitaro, et al.
Publicado: (2026)
Understanding Transformer Optimization via Gradient Heterogeneity
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
por: Tomihari, Akiyoshi, et al.
Publicado: (2025)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
por: Kajitsuka, Tokio, et al.
Publicado: (2023)
por: Kajitsuka, Tokio, et al.
Publicado: (2023)
Deep Grokking: Would Deep Neural Networks Generalize Better?
por: Fan, Simin, et al.
Publicado: (2024)
por: Fan, Simin, et al.
Publicado: (2024)
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
por: Miller, Jack, et al.
Publicado: (2023)
por: Miller, Jack, et al.
Publicado: (2023)
Explaining and Preventing Alignment Collapse in Iterative RLHF
por: Gauthier, Etienne, et al.
Publicado: (2026)
por: Gauthier, Etienne, et al.
Publicado: (2026)
The Complexity Dynamics of Grokking
por: DeMoss, Branton, et al.
Publicado: (2024)
por: DeMoss, Branton, et al.
Publicado: (2024)
Measuring Sharpness in Grokking
por: Miller, Jack, et al.
Publicado: (2024)
por: Miller, Jack, et al.
Publicado: (2024)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
por: Minegishi, Gouki, et al.
Publicado: (2023)
por: Minegishi, Gouki, et al.
Publicado: (2023)
Aligning Multimodal Representations through an Information Bottleneck
por: Almudévar, Antonio, et al.
Publicado: (2025)
por: Almudévar, Antonio, et al.
Publicado: (2025)
Directional Neural Collapse Explains Few-Shot Transfer in Self-Supervised Learning
por: Luthra, Achleshwar, et al.
Publicado: (2026)
por: Luthra, Achleshwar, et al.
Publicado: (2026)
Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds
por: Song, Yiding, et al.
Publicado: (2026)
por: Song, Yiding, et al.
Publicado: (2026)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
por: Kothapalli, Vignesh, et al.
Publicado: (2024)
por: Kothapalli, Vignesh, et al.
Publicado: (2024)
Ejemplares similares
-
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
por: Sakamoto, Keitaro, et al.
Publicado: (2024) -
Benign Overfitting in Token Selection of Attention Mechanism
por: Sakamoto, Keitaro, et al.
Publicado: (2024) -
Multiplicative Logit Adjustment Approximates Neural-Collapse-Aware Decision Boundary Adjustment
por: Hasegawa, Naoya, et al.
Publicado: (2024) -
Explaining Grokking in Transformers through the Lens of Inductive Bias
por: Singh, Jaisidh, et al.
Publicado: (2026) -
Can Test-time Computation Mitigate Reproduction Bias in Neural Symbolic Regression?
por: Sato, Shun, et al.
Publicado: (2025)