Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Favero, Alessandro, Sclocchi, Antonio, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
by: Favero, Alessandro, et al.
Published: (2025)
by: Favero, Alessandro, et al.
Published: (2025)
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)
by: Sclocchi, Antonio, et al.
Published: (2023)
Learn from your own latents and not from tokens: A sample-complexity theory
by: Korchinski, Daniel J., et al.
Published: (2026)
by: Korchinski, Daniel J., et al.
Published: (2026)
Bigger is not Always Better: Scaling Properties of Latent Diffusion Models
by: Mei, Kangfu, et al.
Published: (2024)
by: Mei, Kangfu, et al.
Published: (2024)
Bigger Isn't Always Better: Towards a General Prior for Medical Image Reconstruction
by: Glaszner, Lukas, et al.
Published: (2025)
by: Glaszner, Lukas, et al.
Published: (2025)
ML Interpretability: Simple Isn't Easy
by: Räz, Tim
Published: (2022)
by: Räz, Tim
Published: (2022)
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Model
by: Cagnetta, Francesco, et al.
Published: (2023)
by: Cagnetta, Francesco, et al.
Published: (2023)
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026)
by: Freiesleben, Timo, et al.
Published: (2026)
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
by: Chen, Yiwei, et al.
Published: (2025)
by: Chen, Yiwei, et al.
Published: (2025)
Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
by: Nava, Andres, et al.
Published: (2026)
by: Nava, Andres, et al.
Published: (2026)
Why Isn't Relational Learning Taking Over the World?
by: Poole, David
Published: (2025)
by: Poole, David
Published: (2025)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024)
by: Tomasini, Umberto, et al.
Published: (2024)
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
by: Mustaqim, S. M., et al.
Published: (2025)
by: Mustaqim, S. M., et al.
Published: (2025)
Causal Inference Isn't Special: Why It's Just Another Prediction Problem
by: Fernández-Loría, Carlos
Published: (2025)
by: Fernández-Loría, Carlos
Published: (2025)
How Much Training Data is Memorized in Overparameterized Autoencoders? An Inverse Problem Perspective on Memorization Evaluation
by: Abitbul, Koren, et al.
Published: (2023)
by: Abitbul, Koren, et al.
Published: (2023)
Safety Isn't Always First: A Disturbing Look at Chemistry Books.
by: Manning, Pat, et al.
Published: (1986)
by: Manning, Pat, et al.
Published: (1986)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
by: Barale, Claire, et al.
Published: (2025)
by: Barale, Claire, et al.
Published: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)
by: Galeone, Cosimo, et al.
Published: (2026)
Towards a theory of how the structure of language is acquired by deep neural networks
by: Cagnetta, Francesco, et al.
Published: (2024)
by: Cagnetta, Francesco, et al.
Published: (2024)
AI Isn't Creating Anything
by: Lee Skallerup Bessette
Published: (2025)
by: Lee Skallerup Bessette
Published: (2025)
Is Bigger Always Better? Efficiency Analysis in Resource-Constrained Small Object Detection
by: Mbobda-Kuate, Kwame, et al.
Published: (2026)
by: Mbobda-Kuate, Kwame, et al.
Published: (2026)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024)
by: Yoon, Junsang, et al.
Published: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
On the Edge of Memorization in Diffusion Models
by: Buchanan, Sam, et al.
Published: (2025)
by: Buchanan, Sam, et al.
Published: (2025)
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026)
by: Parley, Jack T., et al.
Published: (2026)
Deriving Neural Scaling Laws from the statistics of natural language
by: Cagnetta, Francesco, et al.
Published: (2026)
by: Cagnetta, Francesco, et al.
Published: (2026)
The Interpolating Information Criterion for Overparameterized Models
by: Hodgkinson, Liam, et al.
Published: (2023)
by: Hodgkinson, Liam, et al.
Published: (2023)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
by: Nguyen, Binh, et al.
Published: (2025)
by: Nguyen, Binh, et al.
Published: (2025)
Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR
by: Vempati, Shashank, et al.
Published: (2025)
by: Vempati, Shashank, et al.
Published: (2025)
Just on Time: Token-Level Early Stopping for Diffusion Language Models
by: Kohut, Zahar, et al.
Published: (2026)
by: Kohut, Zahar, et al.
Published: (2026)
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking
by: LeVine, Will, et al.
Published: (2025)
by: LeVine, Will, et al.
Published: (2025)
Memorize Early, Then Query: Inlier-Memorization-Guided Active Outlier Detection
by: Kang, Minseo, et al.
Published: (2026)
by: Kang, Minseo, et al.
Published: (2026)
Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
by: Guragain, Anmol
Published: (2026)
by: Guragain, Anmol
Published: (2026)
Similar Items
-
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024) -
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024) -
How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
by: Favero, Alessandro, et al.
Published: (2025) -
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025) -
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)