Generalization vs. Memorization in the Presence of Statistical Biases in Transformers
Fuente:
arXiv
Saved in:
| Main Author: | Mitros, John |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Space Factorization in LoRA
by: Kumar, Shashi, et al.
Published: (2025)
by: Kumar, Shashi, et al.
Published: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024)
by: Bayat, Reza, et al.
Published: (2024)
On the Optimal Memorization Capacity of Transformers
by: Kajitsuka, Tokio, et al.
Published: (2024)
by: Kajitsuka, Tokio, et al.
Published: (2024)
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023)
by: Mahdavi, Sadegh, et al.
Published: (2023)
Generalization vs. Memorization in Autoregressive Deep Learning: Or, Examining Temporal Decay of Gradient Coherence
by: Amarel, James, et al.
Published: (2025)
by: Amarel, James, et al.
Published: (2025)
Uncovering Memorization Effect in the Presence of Spurious Correlations
by: You, Chenyu, et al.
Published: (2025)
by: You, Chenyu, et al.
Published: (2025)
When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural Nets
by: Zeno, Chen, et al.
Published: (2025)
by: Zeno, Chen, et al.
Published: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Centralized Selection with Preferences in the Presence of Biases
by: Celis, L. Elisa, et al.
Published: (2024)
by: Celis, L. Elisa, et al.
Published: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
by: Li, Aochong Oliver, et al.
Published: (2025)
by: Li, Aochong Oliver, et al.
Published: (2025)
Prior Aware Memorization: An Efficient Metric for Distinguishing Memorization from Generalization in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2026)
by: Tiwari, Trishita, et al.
Published: (2026)
Generalization-Memorization Machines
by: Wang, Zhen, et al.
Published: (2022)
by: Wang, Zhen, et al.
Published: (2022)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences
by: Li, Siquan, et al.
Published: (2026)
by: Li, Siquan, et al.
Published: (2026)
On the Dynamics & Transferability of Latent Generalization during Memorization
by: Ketha, Simran, et al.
Published: (2026)
by: Ketha, Simran, et al.
Published: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
A Geometric Framework for Understanding Memorization in Generative Models
by: Ross, Brendan Leigh, et al.
Published: (2024)
by: Ross, Brendan Leigh, et al.
Published: (2024)
Memorization in Self-Supervised Learning Improves Downstream Generalization
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
SolidMark: Evaluating Image Memorization in Generative Models
by: Kriplani, Nicky, et al.
Published: (2025)
by: Kriplani, Nicky, et al.
Published: (2025)
Provable Separations between Memorization and Generalization in Diffusion Models
by: Ye, Zeqi, et al.
Published: (2025)
by: Ye, Zeqi, et al.
Published: (2025)
Manifold Generalization Provably Proceeds Memorization in Diffusion Models
by: Shen, Zebang, et al.
Published: (2026)
by: Shen, Zebang, et al.
Published: (2026)
Memorize Early, Then Query: Inlier-Memorization-Guided Active Outlier Detection
by: Kang, Minseo, et al.
Published: (2026)
by: Kang, Minseo, et al.
Published: (2026)
Generative Modeling of Weights: Generalization or Memorization?
by: Zeng, Boya, et al.
Published: (2025)
by: Zeng, Boya, et al.
Published: (2025)
Generalization and Memorization in Rectified Flow
by: Rao, Mingxing, et al.
Published: (2026)
by: Rao, Mingxing, et al.
Published: (2026)
Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
by: Attias, Idan, et al.
Published: (2024)
by: Attias, Idan, et al.
Published: (2024)
Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
by: Li, Ziyue, et al.
Published: (2025)
by: Li, Ziyue, et al.
Published: (2025)
Memorization Sinks: Isolating Memorization during LLM Training
by: Ghosal, Gaurav R., et al.
Published: (2025)
by: Ghosal, Gaurav R., et al.
Published: (2025)
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
by: Yu, Haonan, et al.
Published: (2025)
by: Yu, Haonan, et al.
Published: (2025)
Critical Windows of Complexity Control: When Transformers Decide to Reason or Memorize
by: Ali, Sarwan
Published: (2026)
by: Ali, Sarwan
Published: (2026)
Memorizing Long-tail Data Can Help Generalization Through Composition
by: Zhou, Mo, et al.
Published: (2025)
by: Zhou, Mo, et al.
Published: (2025)
Slower Generalization, Faster Memorization: A Sweet Spot in Algorithmic Learning
by: So, Shin, et al.
Published: (2026)
by: So, Shin, et al.
Published: (2026)
Memorization and Regularization in Generative Diffusion Models
by: Baptista, Ricardo, et al.
Published: (2025)
by: Baptista, Ricardo, et al.
Published: (2025)
Efficiently Attacking Memorization Scores
by: Do, Tue, et al.
Published: (2025)
by: Do, Tue, et al.
Published: (2025)
On the Edge of Memorization in Diffusion Models
by: Buchanan, Sam, et al.
Published: (2025)
by: Buchanan, Sam, et al.
Published: (2025)
Memorization in Graph Neural Networks
by: Jamadandi, Adarsh, et al.
Published: (2025)
by: Jamadandi, Adarsh, et al.
Published: (2025)
How Much Training Data is Memorized in Overparameterized Autoencoders? An Inverse Problem Perspective on Memorization Evaluation
by: Abitbul, Koren, et al.
Published: (2023)
by: Abitbul, Koren, et al.
Published: (2023)
Decoding Generalization from Memorization in Deep Neural Networks
by: Ketha, Simran, et al.
Published: (2025)
by: Ketha, Simran, et al.
Published: (2025)
A Statistical Framework for Alignment with Biased AI Feedback
by: Xia, Xintao, et al.
Published: (2026)
by: Xia, Xintao, et al.
Published: (2026)
A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
by: Shi, Lianghe, et al.
Published: (2025)
by: Shi, Lianghe, et al.
Published: (2025)
Similar Items
-
Latent Space Factorization in LoRA
by: Kumar, Shashi, et al.
Published: (2025) -
The Pitfalls of Memorization: When Memorization Hurts Generalization
by: Bayat, Reza, et al.
Published: (2024) -
On the Optimal Memorization Capacity of Transformers
by: Kajitsuka, Tokio, et al.
Published: (2024) -
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
by: Zhang, Zhongwang, et al.
Published: (2024) -
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023)