Dimer-Enhanced Optimization: A First-Order Approach to Escaping Saddle Points in Neural Network Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Yue, Cao, Zanxia, Liu, Yingchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
von: Bantzis, Ioannis, et al.
Veröffentlicht: (2025)
von: Bantzis, Ioannis, et al.
Veröffentlicht: (2025)
Hierarchical Zero-Order Optimization for Deep Neural Networks
von: Cao, Sansheng, et al.
Veröffentlicht: (2026)
von: Cao, Sansheng, et al.
Veröffentlicht: (2026)
DouRN: Improving DouZero by Residual Neural Networks
von: Chen, Yiquan, et al.
Veröffentlicht: (2024)
von: Chen, Yiquan, et al.
Veröffentlicht: (2024)
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Simultaneous Training of First- and Second-Order Optimizers in Population-Based Reinforcement Learning
von: Pfeiffer, Felix, et al.
Veröffentlicht: (2024)
von: Pfeiffer, Felix, et al.
Veröffentlicht: (2024)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
Cross-Entropy Optimization for Hyperparameter Optimization in Stochastic Gradient-based Approaches to Train Deep Neural Networks
von: Li, Kevin, et al.
Veröffentlicht: (2024)
von: Li, Kevin, et al.
Veröffentlicht: (2024)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
von: Ma, Yuhan, et al.
Veröffentlicht: (2024)
von: Ma, Yuhan, et al.
Veröffentlicht: (2024)
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
von: Flügel, Katharina, et al.
Veröffentlicht: (2023)
von: Flügel, Katharina, et al.
Veröffentlicht: (2023)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
von: Chaubard, Francois, et al.
Veröffentlicht: (2025)
von: Chaubard, Francois, et al.
Veröffentlicht: (2025)
Efficiently Escaping Saddle Points for Policy Optimization
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2023)
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2023)
Spectral Higher-Order Neural Networks
von: Peri, Gianluca, et al.
Veröffentlicht: (2026)
von: Peri, Gianluca, et al.
Veröffentlicht: (2026)
NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
von: Hao, Yongchang, et al.
Veröffentlicht: (2024)
von: Hao, Yongchang, et al.
Veröffentlicht: (2024)
Enhancing Trustworthiness of Graph Neural Networks with Rank-Based Conformal Training
von: Wang, Ting, et al.
Veröffentlicht: (2025)
von: Wang, Ting, et al.
Veröffentlicht: (2025)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
von: Long, Jikai, et al.
Veröffentlicht: (2025)
von: Long, Jikai, et al.
Veröffentlicht: (2025)
QuadraNet V2: Efficient and Sustainable Training of High-Order Neural Networks with Quadratic Adaptation
von: Xu, Chenhui, et al.
Veröffentlicht: (2024)
von: Xu, Chenhui, et al.
Veröffentlicht: (2024)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
von: Zhang, Di, et al.
Veröffentlicht: (2025)
von: Zhang, Di, et al.
Veröffentlicht: (2025)
Quantum Optimization for Training Quantum Neural Networks
von: Liao, Yidong, et al.
Veröffentlicht: (2021)
von: Liao, Yidong, et al.
Veröffentlicht: (2021)
Dimension-Free Saddle-Point Escape in Muon
von: Long, Yanlin, et al.
Veröffentlicht: (2026)
von: Long, Yanlin, et al.
Veröffentlicht: (2026)
Visual Perceptual to Conceptual First-Order Rule Learning Networks
von: Gao, Kun, et al.
Veröffentlicht: (2026)
von: Gao, Kun, et al.
Veröffentlicht: (2026)
Super Level Sets and Exponential Decay: A Synergistic Approach to Stable Neural Network Training
von: Chaudhary, Jatin, et al.
Veröffentlicht: (2024)
von: Chaudhary, Jatin, et al.
Veröffentlicht: (2024)
OptEx: Expediting First-Order Optimization with Approximately Parallelized Iterations
von: Shu, Yao, et al.
Veröffentlicht: (2024)
von: Shu, Yao, et al.
Veröffentlicht: (2024)
Interactive Training: Feedback-Driven Neural Network Optimization
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
Online Neural Networks for Change-Point Detection
von: Hushchyn, Mikhail, et al.
Veröffentlicht: (2020)
von: Hushchyn, Mikhail, et al.
Veröffentlicht: (2020)
SCPL: Enhancing Neural Network Training Throughput with Decoupled Local Losses and Model Parallelism
von: Ho, Ming-Yao, et al.
Veröffentlicht: (2026)
von: Ho, Ming-Yao, et al.
Veröffentlicht: (2026)
Topological Neural Networks: Mitigating the Bottlenecks of Graph Neural Networks via Higher-Order Interactions
von: Giusti, Lorenzo
Veröffentlicht: (2024)
von: Giusti, Lorenzo
Veröffentlicht: (2024)
A Post-Training Enhanced Optimization Approach for Small Language Models
von: Zhai, Keke
Veröffentlicht: (2024)
von: Zhai, Keke
Veröffentlicht: (2024)
Dispelling the Curse of Singularities in Neural Network Optimizations
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
Enhancing Adversarial Training via Reweighting Optimization Trajectory
von: Huang, Tianjin, et al.
Veröffentlicht: (2023)
von: Huang, Tianjin, et al.
Veröffentlicht: (2023)
Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural Networks
von: Xiao, Mingqing, et al.
Veröffentlicht: (2024)
von: Xiao, Mingqing, et al.
Veröffentlicht: (2024)
Identifying Backdoored Graphs in Graph Neural Network Training: An Explanation-Based Approach with Novel Metrics
von: Downer, Jane, et al.
Veröffentlicht: (2024)
von: Downer, Jane, et al.
Veröffentlicht: (2024)
Gradient Alignment in Physics-informed Neural Networks: A Second-Order Optimization Perspective
von: Wang, Sifan, et al.
Veröffentlicht: (2025)
von: Wang, Sifan, et al.
Veröffentlicht: (2025)
Combinatorial Optimization with Automated Graph Neural Networks
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Full Bayesian Significance Testing for Neural Networks
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
von: Wu, Frank Zhengqing, et al.
Veröffentlicht: (2024)
von: Wu, Frank Zhengqing, et al.
Veröffentlicht: (2024)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
von: Hu, Rui, et al.
Veröffentlicht: (2024)
von: Hu, Rui, et al.
Veröffentlicht: (2024)
Adversarial Instance Generation and Robust Training for Neural Combinatorial Optimization with Multiple Objectives
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
Z-Error Loss for Training Neural Networks
von: Godin, Guillaume
Veröffentlicht: (2025)
von: Godin, Guillaume
Veröffentlicht: (2025)
Energy Consumption in Parallel Neural Network Training
von: Huber, Philipp, et al.
Veröffentlicht: (2025)
von: Huber, Philipp, et al.
Veröffentlicht: (2025)
Training Neural Networks for Modularity aids Interpretability
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
von: Bantzis, Ioannis, et al.
Veröffentlicht: (2025) -
Hierarchical Zero-Order Optimization for Deep Neural Networks
von: Cao, Sansheng, et al.
Veröffentlicht: (2026) -
DouRN: Improving DouZero by Residual Neural Networks
von: Chen, Yiquan, et al.
Veröffentlicht: (2024) -
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
von: Zheng, Zihao, et al.
Veröffentlicht: (2026) -
Simultaneous Training of First- and Second-Order Optimizers in Population-Based Reinforcement Learning
von: Pfeiffer, Felix, et al.
Veröffentlicht: (2024)