Thumb on the Scale: Optimal Loss Weighting in Last Layer Retraining
Fuente:
arXiv
Saved in:
| Main Authors: | Stromberg, Nathan, Thrampoulidis, Christos, Sankar, Lalitha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Theoretical Guarantees of Data Augmented Last Layer Retraining Methods
by: Welfert, Monica, et al.
Published: (2024)
by: Welfert, Monica, et al.
Published: (2024)
Correcting Class Imbalance in Prior-Data Fitted Networks for Tabular Classification
by: McDowell, Samuel, et al.
Published: (2026)
by: McDowell, Samuel, et al.
Published: (2026)
Lower Bounds on the MMSE of Adversarially Inferring Sensitive Features
by: Welfert, Monica, et al.
Published: (2025)
by: Welfert, Monica, et al.
Published: (2025)
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
by: Stromberg, Nathan, et al.
Published: (2024)
by: Stromberg, Nathan, et al.
Published: (2024)
Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
by: Zhao, Yize, et al.
Published: (2026)
by: Zhao, Yize, et al.
Published: (2026)
Robustness to Subpopulation Shift with Domain Label Noise via Regularized Annotation of Domains
by: Stromberg, Nathan, et al.
Published: (2024)
by: Stromberg, Nathan, et al.
Published: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
CORAL: Disentangling Latent Representations in Long-Tailed Diffusion
by: Rodriguez, Esther, et al.
Published: (2025)
by: Rodriguez, Esther, et al.
Published: (2025)
Supervised Contrastive Representation Learning: Landscape Analysis with Unconstrained Features
by: Behnia, Tina, et al.
Published: (2024)
by: Behnia, Tina, et al.
Published: (2024)
On the Unreasonable Effectiveness of Last-layer Retraining
by: Hill, John C., et al.
Published: (2025)
by: Hill, John C., et al.
Published: (2025)
Memorization Capacity of Multi-Head Attention in Transformers
by: Mahdavi, Sadegh, et al.
Published: (2023)
by: Mahdavi, Sadegh, et al.
Published: (2023)
AugLoss: A Robust Augmentation-based Fine Tuning Methodology
by: Otstot, Kyle, et al.
Published: (2022)
by: Otstot, Kyle, et al.
Published: (2022)
Memory capacity of two layer neural networks with smooth activations
by: Madden, Liam, et al.
Published: (2023)
by: Madden, Liam, et al.
Published: (2023)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
by: Behnia, Tina, et al.
Published: (2025)
by: Behnia, Tina, et al.
Published: (2025)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
by: Taheri, Hossein, et al.
Published: (2024)
by: Taheri, Hossein, et al.
Published: (2024)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
by: Deng, Wenlong, et al.
Published: (2023)
by: Deng, Wenlong, et al.
Published: (2023)
A Semi-Supervised Approach for Power System Event Identification
by: Taghipourbazargani, Nima, et al.
Published: (2023)
by: Taghipourbazargani, Nima, et al.
Published: (2023)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)
by: Garrod, Connall, et al.
Published: (2025)
Parameter Optimization with Conscious Allocation (POCA)
by: Inman, Joshua, et al.
Published: (2023)
by: Inman, Joshua, et al.
Published: (2023)
An Adversarial Approach to Evaluating the Robustness of Event Identification Models
by: Bahwal, Obai, et al.
Published: (2024)
by: Bahwal, Obai, et al.
Published: (2024)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
by: Zhao, Yize, et al.
Published: (2024)
by: Zhao, Yize, et al.
Published: (2024)
On the Optimization and Generalization of Multi-head Attention
by: Deora, Puneesh, et al.
Published: (2023)
by: Deora, Puneesh, et al.
Published: (2023)
Next-token prediction capacity: general upper bounds and a lower bound for transformers
by: Madden, Liam, et al.
Published: (2024)
by: Madden, Liam, et al.
Published: (2024)
How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
POCAII: Parameter Optimization with Conscious Allocation using Iterative Intelligence
by: Inman, Joshua, et al.
Published: (2025)
by: Inman, Joshua, et al.
Published: (2025)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Class-attribute Priors: Adapting Optimization to Heterogeneity and Fairness Objective
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
GeoClip: Geometry-Aware Clipping for Differentially Private SGD
by: Gilani, Atefeh, et al.
Published: (2025)
by: Gilani, Atefeh, et al.
Published: (2025)
CLAPS: Aleatoric-Epistemic Scaling via Last-Layer Laplace for Conformal Regression
by: Kim, Dongseok, et al.
Published: (2025)
by: Kim, Dongseok, et al.
Published: (2025)
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
by: Gilani, Atefeh, et al.
Published: (2026)
by: Gilani, Atefeh, et al.
Published: (2026)
Robust Model Selection of Gaussian Graphical Models
by: Zahin, Abrar, et al.
Published: (2022)
by: Zahin, Abrar, et al.
Published: (2022)
Closed-Form Last Layer Optimization
by: Galashov, Alexandre, et al.
Published: (2025)
by: Galashov, Alexandre, et al.
Published: (2025)
Is the Last Layer Sufficient for Uncertainty Quantification?
by: Wilson, Joseph, et al.
Published: (2026)
by: Wilson, Joseph, et al.
Published: (2026)
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
Reveal-or-Obscure: A Differentially Private Sampling Algorithm for Discrete Distributions
by: Tasnim, Naima, et al.
Published: (2025)
by: Tasnim, Naima, et al.
Published: (2025)
Last Layer Empirical Bayes
by: Villecroze, Valentin, et al.
Published: (2025)
by: Villecroze, Valentin, et al.
Published: (2025)
Similar Items
-
Theoretical Guarantees of Data Augmented Last Layer Retraining Methods
by: Welfert, Monica, et al.
Published: (2024) -
Correcting Class Imbalance in Prior-Data Fitted Networks for Tabular Classification
by: McDowell, Samuel, et al.
Published: (2026) -
Lower Bounds on the MMSE of Adversarially Inferring Sensitive Features
by: Welfert, Monica, et al.
Published: (2025) -
Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading
by: Stromberg, Nathan, et al.
Published: (2024) -
Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
by: Zhao, Yize, et al.
Published: (2026)