Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yize, Thrampoulidis, Christos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Supervised Contrastive Representation Learning: Landscape Analysis with Unconstrained Features
von: Behnia, Tina, et al.
Veröffentlicht: (2024)
von: Behnia, Tina, et al.
Veröffentlicht: (2024)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
von: Zhao, Yize, et al.
Veröffentlicht: (2024)
von: Zhao, Yize, et al.
Veröffentlicht: (2024)
How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
von: Stromberg, Nathan, et al.
Veröffentlicht: (2025)
von: Stromberg, Nathan, et al.
Veröffentlicht: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2024)
von: Deng, Wenlong, et al.
Veröffentlicht: (2024)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
von: Zhao, Yize, et al.
Veröffentlicht: (2025)
von: Zhao, Yize, et al.
Veröffentlicht: (2025)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
von: Garrod, Connall, et al.
Veröffentlicht: (2025)
von: Garrod, Connall, et al.
Veröffentlicht: (2025)
Memorization Capacity of Multi-Head Attention in Transformers
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2023)
Memory capacity of two layer neural networks with smooth activations
von: Madden, Liam, et al.
Veröffentlicht: (2023)
von: Madden, Liam, et al.
Veröffentlicht: (2023)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
von: Thrampoulidis, Christos, et al.
Veröffentlicht: (2025)
von: Thrampoulidis, Christos, et al.
Veröffentlicht: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
von: Behnia, Tina, et al.
Veröffentlicht: (2025)
von: Behnia, Tina, et al.
Veröffentlicht: (2025)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
von: Taheri, Hossein, et al.
Veröffentlicht: (2024)
von: Taheri, Hossein, et al.
Veröffentlicht: (2024)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
von: Fan, Chen, et al.
Veröffentlicht: (2025)
von: Fan, Chen, et al.
Veröffentlicht: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
von: Deng, Wenlong, et al.
Veröffentlicht: (2023)
von: Deng, Wenlong, et al.
Veröffentlicht: (2023)
You Only Train Once
von: Sakaridis, Christos
Veröffentlicht: (2025)
von: Sakaridis, Christos
Veröffentlicht: (2025)
Neural Collapse for Cross-entropy Class-Imbalanced Learning with Unconstrained ReLU Feature Model
von: Dang, Hien, et al.
Veröffentlicht: (2024)
von: Dang, Hien, et al.
Veröffentlicht: (2024)
On the Optimization and Generalization of Multi-head Attention
von: Deora, Puneesh, et al.
Veröffentlicht: (2023)
von: Deora, Puneesh, et al.
Veröffentlicht: (2023)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
von: Sernau, Luke
Veröffentlicht: (2024)
von: Sernau, Luke
Veröffentlicht: (2024)
Next-token prediction capacity: general upper bounds and a lower bound for transformers
von: Madden, Liam, et al.
Veröffentlicht: (2024)
von: Madden, Liam, et al.
Veröffentlicht: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2025)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2025)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
von: Deora, Puneesh, et al.
Veröffentlicht: (2025)
von: Deora, Puneesh, et al.
Veröffentlicht: (2025)
Geometric Analysis of Unconstrained Feature Models with $d=K$
von: Shen, Yi, et al.
Veröffentlicht: (2024)
von: Shen, Yi, et al.
Veröffentlicht: (2024)
Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
Instance-dependent Early Stopping
von: Yuan, Suqin, et al.
Veröffentlicht: (2025)
von: Yuan, Suqin, et al.
Veröffentlicht: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
Class-attribute Priors: Adapting Optimization to Heterogeneity and Fairness Objective
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
Early Stopping Tabular In-Context Learning
von: Küken, Jaris, et al.
Veröffentlicht: (2025)
von: Küken, Jaris, et al.
Veröffentlicht: (2025)
Noisy Early Stopping for Noisy Labels
von: Toner, William, et al.
Veröffentlicht: (2024)
von: Toner, William, et al.
Veröffentlicht: (2024)
Neural Multivariate Regression: Qualitative Insights from the Unconstrained Feature Model
von: Andriopoulos, George, et al.
Veröffentlicht: (2025)
von: Andriopoulos, George, et al.
Veröffentlicht: (2025)
FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness
von: Amer, Hossam, et al.
Veröffentlicht: (2026)
von: Amer, Hossam, et al.
Veröffentlicht: (2026)
Parameter-Free Dynamic Regret for Unconstrained Linear Bandits
von: Rumi, Alberto, et al.
Veröffentlicht: (2026)
von: Rumi, Alberto, et al.
Veröffentlicht: (2026)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
Early Stopping Based on Repeated Significance
von: Bax, Eric, et al.
Veröffentlicht: (2024)
von: Bax, Eric, et al.
Veröffentlicht: (2024)
GradStop: Exploring Training Dynamics in Unsupervised Outlier Detection through Gradient
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuang, et al.
Veröffentlicht: (2024)
Transformers as Support Vector Machines
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
Gradient-Variation Regret Bounds for Unconstrained Online Learning
von: Zhao, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhao, Yuheng, et al.
Veröffentlicht: (2026)
Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data
von: Hong, Wanli, et al.
Veröffentlicht: (2024)
von: Hong, Wanli, et al.
Veröffentlicht: (2024)
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026)
von: Doron-Arad, Ilan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Supervised Contrastive Representation Learning: Landscape Analysis with Unconstrained Features
von: Behnia, Tina, et al.
Veröffentlicht: (2024) -
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
von: Zhao, Yize, et al.
Veröffentlicht: (2024) -
How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025) -
Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
von: Stromberg, Nathan, et al.
Veröffentlicht: (2025) -
Implicit Optimization Bias of Next-Token Prediction in Linear Models
von: Thrampoulidis, Christos
Veröffentlicht: (2024)