Gespeichert in:
| Hauptverfasser: | Duzgun, Ahmet Cagri, Jelassi, Samy, Li, Yuanzhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.00968 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Does Overparameterization Affect Machine Unlearning of Deep Neural Networks?
von: Alon, Gal, et al.
Veröffentlicht: (2025)
von: Alon, Gal, et al.
Veröffentlicht: (2025)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
von: Jelassi, Samy, et al.
Veröffentlicht: (2026)
von: Jelassi, Samy, et al.
Veröffentlicht: (2026)
Collective Model Intelligence Requires Compatible Specialization
von: Pari, Jyothish, et al.
Veröffentlicht: (2024)
von: Pari, Jyothish, et al.
Veröffentlicht: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
How Does Quantization Affect Multilingual LLMs?
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
von: Li, Kenneth, et al.
Veröffentlicht: (2024)
Mixture of Parrots: Experts improve memorization more than reasoning
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
von: Li, Binghui, et al.
Veröffentlicht: (2024)
von: Li, Binghui, et al.
Veröffentlicht: (2024)
Theoretical Limitations of Ensembles in the Age of Overparameterization
von: Dern, Niclas, et al.
Veröffentlicht: (2024)
von: Dern, Niclas, et al.
Veröffentlicht: (2024)
The Interpolating Information Criterion for Overparameterized Models
von: Hodgkinson, Liam, et al.
Veröffentlicht: (2023)
von: Hodgkinson, Liam, et al.
Veröffentlicht: (2023)
Privacy for Free in the Overparameterized Regime
von: Bombari, Simone, et al.
Veröffentlicht: (2024)
von: Bombari, Simone, et al.
Veröffentlicht: (2024)
Machine Unlearning under Overparameterization
von: Block, Jacob L., et al.
Veröffentlicht: (2025)
von: Block, Jacob L., et al.
Veröffentlicht: (2025)
How Does Code Pretraining Affect Language Model Task Performance?
von: Petty, Jackson, et al.
Veröffentlicht: (2024)
von: Petty, Jackson, et al.
Veröffentlicht: (2024)
How Much Training Data is Memorized in Overparameterized Autoencoders? An Inverse Problem Perspective on Memorization Evaluation
von: Abitbul, Koren, et al.
Veröffentlicht: (2023)
von: Abitbul, Koren, et al.
Veröffentlicht: (2023)
How Does Response Length Affect Long-Form Factuality
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
The Role of Symmetry in Optimizing Overparameterized Networks
von: Sareen, Kusha, et al.
Veröffentlicht: (2026)
von: Sareen, Kusha, et al.
Veröffentlicht: (2026)
Provable Generalization in Overparameterized Neural Nets
von: Dhingra, Aviral
Veröffentlicht: (2025)
von: Dhingra, Aviral
Veröffentlicht: (2025)
A Note on Generalization in Variational Autoencoders: How Effective Is Synthetic Data & Overparameterization?
von: Xiao, Tim Z., et al.
Veröffentlicht: (2023)
von: Xiao, Tim Z., et al.
Veröffentlicht: (2023)
On the Clean Generalization and Robust Overfitting in Adversarial Training from Two Theoretical Views: Representation Complexity and Training Dynamics
von: Li, Binghui, et al.
Veröffentlicht: (2023)
von: Li, Binghui, et al.
Veröffentlicht: (2023)
On the Benefits of Weight Normalization for Overparameterized Matrix Sensing
von: Wei, Yudong, et al.
Veröffentlicht: (2025)
von: Wei, Yudong, et al.
Veröffentlicht: (2025)
Overparameterized Multiple Linear Regression as Hyper-Curve Fitting
von: Atza, E., et al.
Veröffentlicht: (2024)
von: Atza, E., et al.
Veröffentlicht: (2024)
Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime
von: Ghane, Reza, et al.
Veröffentlicht: (2026)
von: Ghane, Reza, et al.
Veröffentlicht: (2026)
An Analytical Model for Overparameterized Learning Under Class Imbalance
von: Mor, Eliav, et al.
Veröffentlicht: (2025)
von: Mor, Eliav, et al.
Veröffentlicht: (2025)
Feature Impact Analysis on Top Long-Jump Performances with Quantile Random Forest and Explainable AI Techniques
von: Gan, Qi, et al.
Veröffentlicht: (2025)
von: Gan, Qi, et al.
Veröffentlicht: (2025)
Critical Influence of Overparameterization on Sharpness-aware Minimization
von: Shin, Sungbin, et al.
Veröffentlicht: (2023)
von: Shin, Sungbin, et al.
Veröffentlicht: (2023)
Bayesian Inference for Consistent Predictions in Overparameterized Nonlinear Regression
von: Wakayama, Tomoya
Veröffentlicht: (2024)
von: Wakayama, Tomoya
Veröffentlicht: (2024)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
von: Zhang, Yaoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Yaoyu, et al.
Veröffentlicht: (2024)
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Task Shift: From Classification to Regression in Overparameterized Linear Models
von: LaBonte, Tyler, et al.
Veröffentlicht: (2025)
von: LaBonte, Tyler, et al.
Veröffentlicht: (2025)
Estimation of Toeplitz Covariance Matrices using Overparameterized Gradient Descent
von: Busbib, Daniel, et al.
Veröffentlicht: (2025)
von: Busbib, Daniel, et al.
Veröffentlicht: (2025)
Precise Asymptotic Generalization for Multiclass Classification with Overparameterized Linear Models
von: Wu, David X., et al.
Veröffentlicht: (2023)
von: Wu, David X., et al.
Veröffentlicht: (2023)
Implicit Regularization and Generalization in Overparameterized Neural Networks
von: Johannsen, Zeran
Veröffentlicht: (2026)
von: Johannsen, Zeran
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Does Overparameterization Affect Machine Unlearning of Deep Neural Networks?
von: Alon, Gal, et al.
Veröffentlicht: (2025) -
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024) -
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
von: Jelassi, Samy, et al.
Veröffentlicht: (2026) -
Collective Model Intelligence Requires Compatible Specialization
von: Pari, Jyothish, et al.
Veröffentlicht: (2024) -
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
von: Qin, Tian, et al.
Veröffentlicht: (2025)