When Models Don't Collapse: On the Consistency of Iterative MLE
Fuente:
arXiv
Guardado en:
| Autores principales: | Barzilai, Daniel, Shamir, Ohad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Limitations of SGD for Multi-Index Models Beyond Statistical Queries
por: Barzilai, Daniel, et al.
Publicado: (2026)
por: Barzilai, Daniel, et al.
Publicado: (2026)
Generalization in Kernel Regression Under Realistic Assumptions
por: Barzilai, Daniel, et al.
Publicado: (2023)
por: Barzilai, Daniel, et al.
Publicado: (2023)
Simple Relative Deviation Bounds for Covariance and Gram Matrices
por: Barzilai, Daniel, et al.
Publicado: (2024)
por: Barzilai, Daniel, et al.
Publicado: (2024)
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
por: Barzilai, Daniel, et al.
Publicado: (2025)
por: Barzilai, Daniel, et al.
Publicado: (2025)
Are Convex Optimization Curves Convex?
por: Barzilai, Guy, et al.
Publicado: (2025)
por: Barzilai, Guy, et al.
Publicado: (2025)
Gradient Descent's Last Iterate is Often (slightly) Suboptimal
por: Kornowski, Guy, et al.
Publicado: (2026)
por: Kornowski, Guy, et al.
Publicado: (2026)
REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
por: Almog, Gal, et al.
Publicado: (2025)
por: Almog, Gal, et al.
Publicado: (2025)
Hardness of Learning Fixed Parities with Neural Networks
por: Shoshani, Itamar, et al.
Publicado: (2025)
por: Shoshani, Itamar, et al.
Publicado: (2025)
An Algorithm with Optimal Dimension-Dependence for Zero-Order Nonsmooth Nonconvex Stochastic Optimization
por: Kornowski, Guy, et al.
Publicado: (2023)
por: Kornowski, Guy, et al.
Publicado: (2023)
On the Complexity of Finding Small Subgradients in Nonsmooth Optimization
por: Kornowski, Guy, et al.
Publicado: (2022)
por: Kornowski, Guy, et al.
Publicado: (2022)
Open Problem: Anytime Convergence Rate of Gradient Descent
por: Kornowski, Guy, et al.
Publicado: (2024)
por: Kornowski, Guy, et al.
Publicado: (2024)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
por: Johnson, Daniel D., et al.
Publicado: (2024)
por: Johnson, Daniel D., et al.
Publicado: (2024)
LoRA and Privacy: When Random Projections Help (and When They Don't)
por: Hu, Yaxi, et al.
Publicado: (2026)
por: Hu, Yaxi, et al.
Publicado: (2026)
When Is Compositional Reasoning Learnable from Verifiable Rewards?
por: Barzilai, Daniel, et al.
Publicado: (2026)
por: Barzilai, Daniel, et al.
Publicado: (2026)
Logarithmic Width Suffices for Robust Memorization
por: Egosi, Amitsour, et al.
Publicado: (2025)
por: Egosi, Amitsour, et al.
Publicado: (2025)
Implicit Regularization Towards Rank Minimization in ReLU Networks
por: Timor, Nadav, et al.
Publicado: (2022)
por: Timor, Nadav, et al.
Publicado: (2022)
The Oracle Complexity of Simplex-based Matrix Games
por: Kornowski, Guy, et al.
Publicado: (2024)
por: Kornowski, Guy, et al.
Publicado: (2024)
On the Hardness of Meaningful Local Guarantees in Nonsmooth Nonconvex Optimization
por: Kornowski, Guy, et al.
Publicado: (2024)
por: Kornowski, Guy, et al.
Publicado: (2024)
From Tempered to Benign Overfitting in ReLU Neural Networks
por: Kornowski, Guy, et al.
Publicado: (2023)
por: Kornowski, Guy, et al.
Publicado: (2023)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
por: Yang, Yongyi, et al.
Publicado: (2026)
por: Yang, Yongyi, et al.
Publicado: (2026)
When Respondents Don't Care Anymore: Identifying the Onset of Careless Responding
por: Welz, Max, et al.
Publicado: (2023)
por: Welz, Max, et al.
Publicado: (2023)
What LLMs Think When You Don't Tell Them What to Think About?
por: Kwon, Yongchan, et al.
Publicado: (2026)
por: Kwon, Yongchan, et al.
Publicado: (2026)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
por: Hernandez, Adriano
Publicado: (2024)
por: Hernandez, Adriano
Publicado: (2024)
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
por: Qiang, Rushi, et al.
Publicado: (2025)
por: Qiang, Rushi, et al.
Publicado: (2025)
Depth Separation in Norm-Bounded Infinite-Width Neural Networks
por: Parkinson, Suzanna, et al.
Publicado: (2024)
por: Parkinson, Suzanna, et al.
Publicado: (2024)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
por: Zhang, Jiefu, et al.
Publicado: (2026)
por: Zhang, Jiefu, et al.
Publicado: (2026)
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
por: Fadeeva, Ekaterina, et al.
Publicado: (2025)
Position: Don't be Afraid of Over-Smoothing And Over-Squashing
por: Kormann, Niklas, et al.
Publicado: (2026)
por: Kormann, Niklas, et al.
Publicado: (2026)
Grow, Don't Overwrite: Fine-tuning Without Forgetting
por: Adila, Dyah, et al.
Publicado: (2026)
por: Adila, Dyah, et al.
Publicado: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
por: Gjølbye, Anders, et al.
Publicado: (2026)
por: Gjølbye, Anders, et al.
Publicado: (2026)
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
por: Cohen, Roi, et al.
Publicado: (2024)
por: Cohen, Roi, et al.
Publicado: (2024)
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
por: Jabbour, Jason, et al.
Publicado: (2025)
por: Jabbour, Jason, et al.
Publicado: (2025)
Why Do Some Language Models Fake Alignment While Others Don't?
por: Sheshadri, Abhay, et al.
Publicado: (2025)
por: Sheshadri, Abhay, et al.
Publicado: (2025)
Don't be so negative! Score-based Generative Modeling with Oracle-assisted Guidance
por: Naderiparizi, Saeid, et al.
Publicado: (2023)
por: Naderiparizi, Saeid, et al.
Publicado: (2023)
Reasoning Models Don't Always Say What They Think
por: Chen, Yanda, et al.
Publicado: (2025)
por: Chen, Yanda, et al.
Publicado: (2025)
You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
por: Moutakanni, Théo, et al.
Publicado: (2024)
por: Moutakanni, Théo, et al.
Publicado: (2024)
We Still Don't Understand High-Dimensional Bayesian Optimization
por: Doumont, Colin, et al.
Publicado: (2025)
por: Doumont, Colin, et al.
Publicado: (2025)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
por: Shahout, Rana, et al.
Publicado: (2024)
por: Shahout, Rana, et al.
Publicado: (2024)
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
por: Faria, Gonçalo, et al.
Publicado: (2025)
por: Faria, Gonçalo, et al.
Publicado: (2025)
Don't Forget Imagination!
por: Vityaev, Evgenii E., et al.
Publicado: (2025)
por: Vityaev, Evgenii E., et al.
Publicado: (2025)
Ejemplares similares
-
Limitations of SGD for Multi-Index Models Beyond Statistical Queries
por: Barzilai, Daniel, et al.
Publicado: (2026) -
Generalization in Kernel Regression Under Realistic Assumptions
por: Barzilai, Daniel, et al.
Publicado: (2023) -
Simple Relative Deviation Bounds for Covariance and Gram Matrices
por: Barzilai, Daniel, et al.
Publicado: (2024) -
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
por: Barzilai, Daniel, et al.
Publicado: (2025) -
Are Convex Optimization Curves Convex?
por: Barzilai, Guy, et al.
Publicado: (2025)