Salvato in:
| Autore principale: | Harvey, Thomas R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2509.03594 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
di: Schotthöfer, Steffen, et al.
Pubblicazione: (2024)
di: Schotthöfer, Steffen, et al.
Pubblicazione: (2024)
AdamZ: An Enhanced Optimisation Method for Neural Network Training
di: Zaznov, Ilia, et al.
Pubblicazione: (2024)
di: Zaznov, Ilia, et al.
Pubblicazione: (2024)
Hidden Convexity of Fair PCA and Fast Solver via Eigenvalue Optimization
di: Shen, Junhui, et al.
Pubblicazione: (2025)
di: Shen, Junhui, et al.
Pubblicazione: (2025)
TaskMet: Task-Driven Metric Learning for Model Learning
di: Bansal, Dishank, et al.
Pubblicazione: (2023)
di: Bansal, Dishank, et al.
Pubblicazione: (2023)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
di: Cattaneo, Matias D., et al.
Pubblicazione: (2025)
di: Cattaneo, Matias D., et al.
Pubblicazione: (2025)
Unveiling Hidden Pivotal Players with GoalNet: A GNN-Based Soccer Player Evaluation System
di: Jiang, Jacky Hao, et al.
Pubblicazione: (2025)
di: Jiang, Jacky Hao, et al.
Pubblicazione: (2025)
Training Infinitely Deep and Wide Transformers
di: Barboni, Raphaël, et al.
Pubblicazione: (2026)
di: Barboni, Raphaël, et al.
Pubblicazione: (2026)
Unsupervised Machine Learning Hybrid Approach Integrating Linear Programming in Loss Function: A Robust Optimization Technique
di: Kiruluta, Andrew, et al.
Pubblicazione: (2024)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2024)
Accelerating RLHF Training with Reward Variance Increase
di: Yang, Zonglin, et al.
Pubblicazione: (2025)
di: Yang, Zonglin, et al.
Pubblicazione: (2025)
Anytime Training with Schedule-Free Spectral Optimization
di: Apte, Anuj, et al.
Pubblicazione: (2026)
di: Apte, Anuj, et al.
Pubblicazione: (2026)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
di: Han, X. Y., et al.
Pubblicazione: (2025)
di: Han, X. Y., et al.
Pubblicazione: (2025)
Challenges in Training PINNs: A Loss Landscape Perspective
di: Rathore, Pratik, et al.
Pubblicazione: (2024)
di: Rathore, Pratik, et al.
Pubblicazione: (2024)
Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
di: Sheng, Jiayuan, et al.
Pubblicazione: (2025)
di: Sheng, Jiayuan, et al.
Pubblicazione: (2025)
Training Safe Neural Networks with Global SDP Bounds
di: Soletskyi, Roman, et al.
Pubblicazione: (2024)
di: Soletskyi, Roman, et al.
Pubblicazione: (2024)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
di: Zhang, Fangzhao, et al.
Pubblicazione: (2026)
di: Zhang, Fangzhao, et al.
Pubblicazione: (2026)
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization
di: Zhang, Shihao, et al.
Pubblicazione: (2025)
di: Zhang, Shihao, et al.
Pubblicazione: (2025)
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack
di: Gai, Kuo, et al.
Pubblicazione: (2024)
di: Gai, Kuo, et al.
Pubblicazione: (2024)
Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method
di: Bright-Thonney, Samuel, et al.
Pubblicazione: (2026)
di: Bright-Thonney, Samuel, et al.
Pubblicazione: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
di: Meterez, Alexandru, et al.
Pubblicazione: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
di: Wang, Jinbo, et al.
Pubblicazione: (2025)
di: Wang, Jinbo, et al.
Pubblicazione: (2025)
Decision-Focused Forecasting: A Differentiable Multistage Optimisation Architecture
di: Peršak, Egon, et al.
Pubblicazione: (2024)
di: Peršak, Egon, et al.
Pubblicazione: (2024)
Min-Max Optimisation for Nonconvex-Nonconcave Functions Using a Random Zeroth-Order Extragradient Algorithm
di: Farzin, Amir Ali, et al.
Pubblicazione: (2025)
di: Farzin, Amir Ali, et al.
Pubblicazione: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
di: Song, Minhak, et al.
Pubblicazione: (2025)
di: Song, Minhak, et al.
Pubblicazione: (2025)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
di: Hrycej, Tomas, et al.
Pubblicazione: (2025)
di: Hrycej, Tomas, et al.
Pubblicazione: (2025)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
di: Hong, Seong-Hyun, et al.
Pubblicazione: (2024)
di: Hong, Seong-Hyun, et al.
Pubblicazione: (2024)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
di: Ghosh, Ipsita, et al.
Pubblicazione: (2025)
di: Ghosh, Ipsita, et al.
Pubblicazione: (2025)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
di: Farhat, Yehya, et al.
Pubblicazione: (2023)
di: Farhat, Yehya, et al.
Pubblicazione: (2023)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
di: Yu, Dingzhi, et al.
Pubblicazione: (2026)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
di: Liu, Xin, et al.
Pubblicazione: (2022)
di: Liu, Xin, et al.
Pubblicazione: (2022)
Stochastic Security as a Performance Metric for Quantum-enhanced Generative AI
di: Crum, Noah A., et al.
Pubblicazione: (2023)
di: Crum, Noah A., et al.
Pubblicazione: (2023)
Primitive Agentic First-Order Optimization
di: Sala, R.
Pubblicazione: (2024)
di: Sala, R.
Pubblicazione: (2024)
Optimal Power Grid Operations with Foundation Models
di: Puech, Alban, et al.
Pubblicazione: (2024)
di: Puech, Alban, et al.
Pubblicazione: (2024)
Combining Reinforcement Learning and Optimal Transport for the Traveling Salesman Problem
di: Goh, Yong Liang, et al.
Pubblicazione: (2022)
di: Goh, Yong Liang, et al.
Pubblicazione: (2022)
DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization
di: Leenders, Nick, et al.
Pubblicazione: (2025)
di: Leenders, Nick, et al.
Pubblicazione: (2025)
Faster Reinforcement Learning by Freezing Slow States
di: Wang, Yijia, et al.
Pubblicazione: (2023)
di: Wang, Yijia, et al.
Pubblicazione: (2023)
gridfm-datakit-v1: A Python Library for Scalable and Realistic Power Flow and Optimal Power Flow Data Generation
di: Puech, Alban, et al.
Pubblicazione: (2025)
di: Puech, Alban, et al.
Pubblicazione: (2025)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
di: Ozaslan, Ibrahim K., et al.
Pubblicazione: (2024)
di: Ozaslan, Ibrahim K., et al.
Pubblicazione: (2024)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
di: Zhang, Yihao, et al.
Pubblicazione: (2024)
di: Zhang, Yihao, et al.
Pubblicazione: (2024)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
di: Ding, Dongsheng, et al.
Pubblicazione: (2022)
di: Ding, Dongsheng, et al.
Pubblicazione: (2022)
Documenti analoghi
-
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
di: Schotthöfer, Steffen, et al.
Pubblicazione: (2024) -
AdamZ: An Enhanced Optimisation Method for Neural Network Training
di: Zaznov, Ilia, et al.
Pubblicazione: (2024) -
Hidden Convexity of Fair PCA and Fast Solver via Eigenvalue Optimization
di: Shen, Junhui, et al.
Pubblicazione: (2025) -
TaskMet: Task-Driven Metric Learning for Model Learning
di: Bansal, Dishank, et al.
Pubblicazione: (2023) -
How Memory in Optimization Algorithms Implicitly Modifies the Loss
di: Cattaneo, Matias D., et al.
Pubblicazione: (2025)