Enregistré dans:
| Auteurs principaux: | Singh, Sidak Pal, Mobahi, Hossein, Agarwala, Atish, Dauphin, Yann |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2502.02407 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Neglected Hessian component explains mysteries in Sharpness regularization
par: Dauphin, Yann N., et autres
Publié: (2024)
par: Dauphin, Yann N., et autres
Publié: (2024)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
par: Singh, Sidak Pal, et autres
Publié: (2024)
par: Singh, Sidak Pal, et autres
Publié: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
par: Bozic, Vukasin, et autres
Publié: (2023)
par: Bozic, Vukasin, et autres
Publié: (2023)
High dimensional theory of two-phase optimizers
par: Agarwala, Atish
Publié: (2026)
par: Agarwala, Atish
Publié: (2026)
Per-example gradients: a new frontier for understanding and improving optimizers
par: Roulet, Vincent, et autres
Publié: (2025)
par: Roulet, Vincent, et autres
Publié: (2025)
Introduction to speech recognition
par: Dauphin, Gabriel
Publié: (2024)
par: Dauphin, Gabriel
Publié: (2024)
Accelerating Neural Network Training Along Sharp and Flat Directions
par: Zakarin, Daniyar, et autres
Publié: (2025)
par: Zakarin, Daniyar, et autres
Publié: (2025)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
par: Khromov, Grigory, et autres
Publié: (2023)
par: Khromov, Grigory, et autres
Publié: (2023)
A density estimation perspective on learning from pairwise human preferences
par: Dumoulin, Vincent, et autres
Publié: (2023)
par: Dumoulin, Vincent, et autres
Publié: (2023)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
par: Agarwala, Atish, et autres
Publié: (2024)
par: Agarwala, Atish, et autres
Publié: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
par: Beaglehole, Daniel, et autres
Publié: (2024)
par: Beaglehole, Daniel, et autres
Publié: (2024)
On the Interplay Between Stepsize Tuning and Progressive Sharpening
par: Roulet, Vincent, et autres
Publié: (2023)
par: Roulet, Vincent, et autres
Publié: (2023)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
par: Ormaniec, Weronika, et autres
Publié: (2024)
par: Ormaniec, Weronika, et autres
Publié: (2024)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
par: Zhao, Jim, et autres
Publié: (2024)
par: Zhao, Jim, et autres
Publié: (2024)
On the Foundations of Shortcut Learning
par: Hermann, Katherine L., et autres
Publié: (2023)
par: Hermann, Katherine L., et autres
Publié: (2023)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
par: Marshall, Noah, et autres
Publié: (2024)
par: Marshall, Noah, et autres
Publié: (2024)
What do near-optimal learning rate schedules look like?
par: Naganuma, Hiroki, et autres
Publié: (2026)
par: Naganuma, Hiroki, et autres
Publié: (2026)
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
par: Xiao, Ke Liang, et autres
Publié: (2024)
par: Xiao, Ke Liang, et autres
Publié: (2024)
Reasoning Boosts Opinion Alignment in LLMs
par: Berdoz, Frédéric, et autres
Publié: (2026)
par: Berdoz, Frédéric, et autres
Publié: (2026)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
par: Arefin, Md Rifat, et autres
Publié: (2024)
par: Arefin, Md Rifat, et autres
Publié: (2024)
Towards Meta-Pruning via Optimal Transport
par: Theus, Alexander, et autres
Publié: (2024)
par: Theus, Alexander, et autres
Publié: (2024)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
par: Zhou, Jin Peng, et autres
Publié: (2025)
par: Zhou, Jin Peng, et autres
Publié: (2025)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
par: Qiu, Shikai, et autres
Publié: (2025)
par: Qiu, Shikai, et autres
Publié: (2025)
Mining Mental Health Signals: A Comparative Study of Four Machine Learning Methods for Depression Detection from Social Media Posts in Sorani Kurdish
par: Mohammed, Idrees, et autres
Publié: (2025)
par: Mohammed, Idrees, et autres
Publié: (2025)
Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
par: Reddy, Karan, et autres
Publié: (2025)
par: Reddy, Karan, et autres
Publié: (2025)
Data-Aware Random Feature Kernel for Transformers
par: Farzam, Amirhossein, et autres
Publié: (2026)
par: Farzam, Amirhossein, et autres
Publié: (2026)
Local vs Global continual learning
par: Lanzillotta, Giulia, et autres
Publié: (2024)
par: Lanzillotta, Giulia, et autres
Publié: (2024)
Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
par: Lin, Huawei, et autres
Publié: (2025)
par: Lin, Huawei, et autres
Publié: (2025)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
par: Singh, Vaibhav, et autres
Publié: (2025)
par: Singh, Vaibhav, et autres
Publié: (2025)
FEval-TTC: Fair Evaluation Protocol for Test-Time Compute
par: Rumiantsev, Pavel, et autres
Publié: (2025)
par: Rumiantsev, Pavel, et autres
Publié: (2025)
LLMs can learn self-restraint through iterative self-reflection
par: Piché, Alexandre, et autres
Publié: (2024)
par: Piché, Alexandre, et autres
Publié: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
par: Skean, Oscar, et autres
Publié: (2024)
par: Skean, Oscar, et autres
Publié: (2024)
Exploring Precision and Recall to assess the quality and diversity of LLMs
par: Bronnec, Florian Le, et autres
Publié: (2024)
par: Bronnec, Florian Le, et autres
Publié: (2024)
Robustmix: Improving Robustness by Regularizing the Frequency Bias of Deep Nets
par: Ngnawe, Jonas, et autres
Publié: (2023)
par: Ngnawe, Jonas, et autres
Publié: (2023)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
par: Roulet, Vincent, et autres
Publié: (2024)
par: Roulet, Vincent, et autres
Publié: (2024)
A Gauge Theory of Superposition: Toward a Sheaf-Theoretic Atlas of Neural Representations
par: Javidnia, Hossein
Publié: (2026)
par: Javidnia, Hossein
Publié: (2026)
Semantic Sections: An Atlas-Native Feature Ontology for Obstructed Representation Spaces
par: Javidnia, Hossein
Publié: (2026)
par: Javidnia, Hossein
Publié: (2026)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
par: Rashidi, Sina, et autres
Publié: (2025)
par: Rashidi, Sina, et autres
Publié: (2025)
A Comprehensive Approach to Misspelling Correction with BERT and Levenshtein Distance
par: Naziri, Amirreza, et autres
Publié: (2024)
par: Naziri, Amirreza, et autres
Publié: (2024)
Phases of Muon: When Muon Eclipses SignSGD
par: Paquette, Elliot, et autres
Publié: (2026)
par: Paquette, Elliot, et autres
Publié: (2026)
Documents similaires
-
Neglected Hessian component explains mysteries in Sharpness regularization
par: Dauphin, Yann N., et autres
Publié: (2024) -
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
par: Singh, Sidak Pal, et autres
Publié: (2024) -
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
par: Bozic, Vukasin, et autres
Publié: (2023) -
High dimensional theory of two-phase optimizers
par: Agarwala, Atish
Publié: (2026) -
Per-example gradients: a new frontier for understanding and improving optimizers
par: Roulet, Vincent, et autres
Publié: (2025)