Saved in:
| Main Author: | Akiyama, Shunta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.22667 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models
by: Sonoda, Sho, et al.
Published: (2026)
by: Sonoda, Sho, et al.
Published: (2026)
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023)
by: Phunyaphibarn, Prin, et al.
Published: (2023)
Exponential Sample Complexity Separation between Flat and Hierarchical Agentic Theorem Provers
by: Sonoda, Sho, et al.
Published: (2026)
by: Sonoda, Sho, et al.
Published: (2026)
Geometry and Local Recovery of Global Minima of Two-layer Neural Networks at Overparameterization
by: Zhang, Leyang, et al.
Published: (2023)
by: Zhang, Leyang, et al.
Published: (2023)
Differentially Private Random Block Coordinate Descent
by: Maranjyan, Artavazd, et al.
Published: (2024)
by: Maranjyan, Artavazd, et al.
Published: (2024)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
by: Hsiao, Yen-Che, et al.
Published: (2024)
by: Hsiao, Yen-Che, et al.
Published: (2024)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
by: Lee, Dongyeop, et al.
Published: (2025)
by: Lee, Dongyeop, et al.
Published: (2025)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Neural Networks with Complex-Valued Weights Have No Spurious Local Minima
by: Liu, Xingtu
Published: (2021)
by: Liu, Xingtu
Published: (2021)
Coordinate Descent for Network Linearization
by: Rakhlin, Vlad, et al.
Published: (2025)
by: Rakhlin, Vlad, et al.
Published: (2025)
Hidden State Differential Private Mini-Batch Block Coordinate Descent for Multi-convexity Optimization
by: Chen, Ding, et al.
Published: (2024)
by: Chen, Ding, et al.
Published: (2024)
Gradient Descent with Provably Tuned Learning-rate Schedules
by: Sharma, Dravyansh
Published: (2025)
by: Sharma, Dravyansh
Published: (2025)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Gradient Descent Finds Over-Parameterized Neural Networks with Sharp Generalization for Nonparametric Regression
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Asynchronous Decentralized SGD under Non-Convexity: A Block-Coordinate Descent Framework
by: Zhou, Yijie, et al.
Published: (2025)
by: Zhou, Yijie, et al.
Published: (2025)
FedBCD:Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
by: Liu, Junkang, et al.
Published: (2026)
by: Liu, Junkang, et al.
Published: (2026)
DP-FedPGN: Finding Global Flat Minima for Differentially Private Federated Learning via Penalizing Gradient Norm
by: Liu, Junkang, et al.
Published: (2025)
by: Liu, Junkang, et al.
Published: (2025)
Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering Phenomenon
by: Liang, Tongtong, et al.
Published: (2025)
by: Liang, Tongtong, et al.
Published: (2025)
Randomized Block-Coordinate Optimistic Gradient Algorithms for Root-Finding Problems
by: Tran-Dinh, Quoc, et al.
Published: (2023)
by: Tran-Dinh, Quoc, et al.
Published: (2023)
SMiLE: Provably Enforcing Global Relational Properties in Neural Networks
by: Francobaldi, Matteo, et al.
Published: (2025)
by: Francobaldi, Matteo, et al.
Published: (2025)
Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
by: Gan, Eric
Published: (2026)
by: Gan, Eric
Published: (2026)
A Block Coordinate Descent Method for Nonsmooth Composite Optimization under Orthogonality Constraints
by: Yuan, Ganzhao
Published: (2023)
by: Yuan, Ganzhao
Published: (2023)
Generative Autoencoding of Dropout Patterns
by: Maeda, Shunta
Published: (2023)
by: Maeda, Shunta
Published: (2023)
Global Minima by Penalized Full-dimensional Scaling
by: de Leeuw, Jan
Published: (2024)
by: de Leeuw, Jan
Published: (2024)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
by: Chu, Ya-Chi, et al.
Published: (2025)
by: Chu, Ya-Chi, et al.
Published: (2025)
Provably Faster Gradient Descent via Long Steps
by: Grimmer, Benjamin
Published: (2023)
by: Grimmer, Benjamin
Published: (2023)
Provably Bounding Neural Network Preimages
by: Kotha, Suhas, et al.
Published: (2023)
by: Kotha, Suhas, et al.
Published: (2023)
Hidden Minima in Two-Layer ReLU Networks
by: Arjevani, Yossi
Published: (2023)
by: Arjevani, Yossi
Published: (2023)
Are Flat Minima an Illusion?
by: Bennett, Michael Timothy
Published: (2026)
by: Bennett, Michael Timothy
Published: (2026)
Asynchronous Distributed Reinforcement Learning for LQR Control via Zeroth-Order Block Coordinate Descent
by: Jing, Gangshan, et al.
Published: (2021)
by: Jing, Gangshan, et al.
Published: (2021)
Architecture-Aware Minimization (A$^2$M): How to Find Flat Minima in Neural Architecture Search
by: Gambella, Matteo, et al.
Published: (2025)
by: Gambella, Matteo, et al.
Published: (2025)
Stochastic Gradient Descent for Two-layer Neural Networks
by: Cao, Dinghao, et al.
Published: (2024)
by: Cao, Dinghao, et al.
Published: (2024)
Variational Stochastic Gradient Descent for Deep Neural Networks
by: Chen, Haotian, et al.
Published: (2024)
by: Chen, Haotian, et al.
Published: (2024)
Noise Stability Optimization for Finding Flat Minima: A Hessian-based Regularization Approach
by: Zhang, Hongyang R., et al.
Published: (2023)
by: Zhang, Hongyang R., et al.
Published: (2023)
Provable Privacy Attacks on Trained Shallow Neural Networks
by: Smorodinsky, Guy, et al.
Published: (2024)
by: Smorodinsky, Guy, et al.
Published: (2024)
Provably Powerful Graph Neural Networks for Directed Multigraphs
by: Egressy, Béni, et al.
Published: (2023)
by: Egressy, Béni, et al.
Published: (2023)
Designing a Linearized Potential Function in Neural Network Optimization Using Csiszár Type of Tsallis Entropy
by: Akiyama, Keito
Published: (2024)
by: Akiyama, Keito
Published: (2024)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
by: Vaswani, Sharan, et al.
Published: (2025)
by: Vaswani, Sharan, et al.
Published: (2025)
Similar Items
-
Why Agentic Theorem Prover Works: A Statistical Provability Theory of Mathematical Reasoning Models
by: Sonoda, Sho, et al.
Published: (2026) -
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
by: Phunyaphibarn, Prin, et al.
Published: (2023) -
Exponential Sample Complexity Separation between Flat and Hierarchical Agentic Theorem Provers
by: Sonoda, Sho, et al.
Published: (2026) -
Geometry and Local Recovery of Global Minima of Two-layer Neural Networks at Overparameterization
by: Zhang, Leyang, et al.
Published: (2023) -
Differentially Private Random Block Coordinate Descent
by: Maranjyan, Artavazd, et al.
Published: (2024)