Avoiding spurious sharpness minimization broadens applicability of SAM
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Sidak Pal, Mobahi, Hossein, Agarwala, Atish, Dauphin, Yann |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Neglected Hessian component explains mysteries in Sharpness regularization
di: Dauphin, Yann N., et al.
Pubblicazione: (2024)
di: Dauphin, Yann N., et al.
Pubblicazione: (2024)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
di: Singh, Sidak Pal, et al.
Pubblicazione: (2024)
di: Singh, Sidak Pal, et al.
Pubblicazione: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
High dimensional theory of two-phase optimizers
di: Agarwala, Atish
Pubblicazione: (2026)
di: Agarwala, Atish
Pubblicazione: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
di: Roulet, Vincent, et al.
Pubblicazione: (2025)
di: Roulet, Vincent, et al.
Pubblicazione: (2025)
Introduction to speech recognition
di: Dauphin, Gabriel
Pubblicazione: (2024)
di: Dauphin, Gabriel
Pubblicazione: (2024)
Accelerating Neural Network Training Along Sharp and Flat Directions
di: Zakarin, Daniyar, et al.
Pubblicazione: (2025)
di: Zakarin, Daniyar, et al.
Pubblicazione: (2025)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
di: Khromov, Grigory, et al.
Pubblicazione: (2023)
di: Khromov, Grigory, et al.
Pubblicazione: (2023)
A density estimation perspective on learning from pairwise human preferences
di: Dumoulin, Vincent, et al.
Pubblicazione: (2023)
di: Dumoulin, Vincent, et al.
Pubblicazione: (2023)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
di: Agarwala, Atish, et al.
Pubblicazione: (2024)
di: Agarwala, Atish, et al.
Pubblicazione: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
di: Ormaniec, Weronika, et al.
Pubblicazione: (2024)
di: Ormaniec, Weronika, et al.
Pubblicazione: (2024)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
di: Zhao, Jim, et al.
Pubblicazione: (2024)
di: Zhao, Jim, et al.
Pubblicazione: (2024)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
di: Roulet, Vincent, et al.
Pubblicazione: (2023)
di: Roulet, Vincent, et al.
Pubblicazione: (2023)
Reasoning Boosts Opinion Alignment in LLMs
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
On the Foundations of Shortcut Learning
di: Hermann, Katherine L., et al.
Pubblicazione: (2023)
di: Hermann, Katherine L., et al.
Pubblicazione: (2023)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
di: Arefin, Md Rifat, et al.
Pubblicazione: (2024)
di: Arefin, Md Rifat, et al.
Pubblicazione: (2024)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
di: Marshall, Noah, et al.
Pubblicazione: (2024)
di: Marshall, Noah, et al.
Pubblicazione: (2024)
What do near-optimal learning rate schedules look like?
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
di: Xiao, Ke Liang, et al.
Pubblicazione: (2024)
di: Xiao, Ke Liang, et al.
Pubblicazione: (2024)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
Mining Mental Health Signals: A Comparative Study of Four Machine Learning Methods for Depression Detection from Social Media Posts in Sorani Kurdish
di: Mohammed, Idrees, et al.
Pubblicazione: (2025)
di: Mohammed, Idrees, et al.
Pubblicazione: (2025)
Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
di: Reddy, Karan, et al.
Pubblicazione: (2025)
di: Reddy, Karan, et al.
Pubblicazione: (2025)
Towards Meta-Pruning via Optimal Transport
di: Theus, Alexander, et al.
Pubblicazione: (2024)
di: Theus, Alexander, et al.
Pubblicazione: (2024)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
FEval-TTC: Fair Evaluation Protocol for Test-Time Compute
di: Rumiantsev, Pavel, et al.
Pubblicazione: (2025)
di: Rumiantsev, Pavel, et al.
Pubblicazione: (2025)
LLMs can learn self-restraint through iterative self-reflection
di: Piché, Alexandre, et al.
Pubblicazione: (2024)
di: Piché, Alexandre, et al.
Pubblicazione: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
di: Skean, Oscar, et al.
Pubblicazione: (2024)
di: Skean, Oscar, et al.
Pubblicazione: (2024)
Exploring Precision and Recall to assess the quality and diversity of LLMs
di: Bronnec, Florian Le, et al.
Pubblicazione: (2024)
di: Bronnec, Florian Le, et al.
Pubblicazione: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
di: Zarch, Hossein Entezari, et al.
Pubblicazione: (2025)
Data-Aware Random Feature Kernel for Transformers
di: Farzam, Amirhossein, et al.
Pubblicazione: (2026)
di: Farzam, Amirhossein, et al.
Pubblicazione: (2026)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
di: Sundaram, Shobhita, et al.
Pubblicazione: (2026)
di: Sundaram, Shobhita, et al.
Pubblicazione: (2026)
Clinical Context-aware Radiology Report Generation from Medical Images using Transformers
di: Singh, Sonit
Pubblicazione: (2024)
di: Singh, Sonit
Pubblicazione: (2024)
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
di: Verine, Alexandre, et al.
Pubblicazione: (2025)
di: Verine, Alexandre, et al.
Pubblicazione: (2025)
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
di: Fan, Chongyu, et al.
Pubblicazione: (2025)
di: Fan, Chongyu, et al.
Pubblicazione: (2025)
Local vs Global continual learning
di: Lanzillotta, Giulia, et al.
Pubblicazione: (2024)
di: Lanzillotta, Giulia, et al.
Pubblicazione: (2024)
Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
di: Lin, Huawei, et al.
Pubblicazione: (2025)
di: Lin, Huawei, et al.
Pubblicazione: (2025)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
di: Singh, Vaibhav, et al.
Pubblicazione: (2025)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
di: Rashidi, Sina, et al.
Pubblicazione: (2025)
di: Rashidi, Sina, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Neglected Hessian component explains mysteries in Sharpness regularization
di: Dauphin, Yann N., et al.
Pubblicazione: (2024) -
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
di: Singh, Sidak Pal, et al.
Pubblicazione: (2024) -
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023) -
High dimensional theory of two-phase optimizers
di: Agarwala, Atish
Pubblicazione: (2026) -
Per-example gradients: a new frontier for understanding and improving optimizers
di: Roulet, Vincent, et al.
Pubblicazione: (2025)