Simmering: Sufficient is better than optimal for training neural networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Babayan, Irina, Aliahmadi, Hazhir, van Anders, Greg |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Entropic Auto-Encoding via Implicit Free-Energy Minimization
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2026)
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2026)
Flashpoints Signal Hidden Inherent Instabilities in Land-Use Planning
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2023)
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2023)
Adaptive multiple optimal learning factors for neural network training
di: Challagundla, Jeshwanth
Pubblicazione: (2024)
di: Challagundla, Jeshwanth
Pubblicazione: (2024)
Transforming Design Spaces Using Pareto-Laplace Filters
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2024)
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2024)
Filters reveal emergent structure in computational morphogenesis
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2025)
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2025)
When narrower is better: the narrow width limit of Bayesian parallel branching neural networks
di: Zhang, Zechen, et al.
Pubblicazione: (2024)
di: Zhang, Zechen, et al.
Pubblicazione: (2024)
Confidence-gated training for efficient early-exit neural networks
di: Mokssit, Saad, et al.
Pubblicazione: (2025)
di: Mokssit, Saad, et al.
Pubblicazione: (2025)
A comparative analysis of a neural network with calculated weights and a neural network with random generation of weights based on the training dataset size
di: Geidarov, Polad
Pubblicazione: (2025)
di: Geidarov, Polad
Pubblicazione: (2025)
Language models are better than humans at next-token prediction
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
Do graph neural network states contain graph properties?
di: Pelletreau-Duris, Tom, et al.
Pubblicazione: (2024)
di: Pelletreau-Duris, Tom, et al.
Pubblicazione: (2024)
Target noise: A pre-training based neural network initialization for efficient high resolution learning
di: Wang, Shaowen, et al.
Pubblicazione: (2026)
di: Wang, Shaowen, et al.
Pubblicazione: (2026)
On permutation-invariant neural networks
di: Kimura, Masanari, et al.
Pubblicazione: (2024)
di: Kimura, Masanari, et al.
Pubblicazione: (2024)
Sobolev acceleration for neural networks
di: Oh, Jong Kwon, et al.
Pubblicazione: (2025)
di: Oh, Jong Kwon, et al.
Pubblicazione: (2025)
Attention mechanisms in neural networks
di: Hays, Hasi
Pubblicazione: (2026)
di: Hays, Hasi
Pubblicazione: (2026)
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
di: Keisler, Julie, et al.
Pubblicazione: (2023)
di: Keisler, Julie, et al.
Pubblicazione: (2023)
Two are better than one: Context window extension with multi-grained self-injection
di: Han, Wei, et al.
Pubblicazione: (2024)
di: Han, Wei, et al.
Pubblicazione: (2024)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
di: Yang, Zonglin, et al.
Pubblicazione: (2026)
di: Yang, Zonglin, et al.
Pubblicazione: (2026)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
di: Nunez, Jeanmely Rojas, et al.
Pubblicazione: (2026)
Action-Sufficient Goal Representations
di: Hyeon, Jinu, et al.
Pubblicazione: (2026)
di: Hyeon, Jinu, et al.
Pubblicazione: (2026)
Principles of Lipschitz continuity in neural networks
di: Luo, Róisín
Pubblicazione: (2026)
di: Luo, Róisín
Pubblicazione: (2026)
Linearity-based neural network compression
di: Dobler, Silas, et al.
Pubblicazione: (2025)
di: Dobler, Silas, et al.
Pubblicazione: (2025)
Astral: training physics-informed neural networks with error majorants
di: Fanaskov, Vladimir, et al.
Pubblicazione: (2024)
di: Fanaskov, Vladimir, et al.
Pubblicazione: (2024)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
di: Joshi, Ameya
Pubblicazione: (2025)
di: Joshi, Ameya
Pubblicazione: (2025)
A framework for measuring the training efficiency of a neural architecture
di: Cueto-Mendoza, Eduardo, et al.
Pubblicazione: (2024)
di: Cueto-Mendoza, Eduardo, et al.
Pubblicazione: (2024)
Towards graph neural networks for provably solving convex optimization problems
di: Qian, Chendi, et al.
Pubblicazione: (2025)
di: Qian, Chendi, et al.
Pubblicazione: (2025)
Applying graph neural network to SupplyGraph for supply chain network
di: Han, Kihwan
Pubblicazione: (2024)
di: Han, Kihwan
Pubblicazione: (2024)
Sufficient Invariant Learning for Distribution Shift
di: Kim, Taero, et al.
Pubblicazione: (2022)
di: Kim, Taero, et al.
Pubblicazione: (2022)
Understanding the dynamics of the frequency bias in neural networks
di: Molina, Juan, et al.
Pubblicazione: (2024)
di: Molina, Juan, et al.
Pubblicazione: (2024)
Graph neural networks informed locally by thermodynamics
di: Tierz, Alicia, et al.
Pubblicazione: (2024)
di: Tierz, Alicia, et al.
Pubblicazione: (2024)
Graph neural networks and non-commuting operators
di: Velasco, Mauricio, et al.
Pubblicazione: (2024)
di: Velasco, Mauricio, et al.
Pubblicazione: (2024)
LayerCollapse: Adaptive compression of neural networks
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
Self-Improving Pretraining: using post-trained models to pretrain better models
di: Tan, Ellen Xiaoqing, et al.
Pubblicazione: (2026)
di: Tan, Ellen Xiaoqing, et al.
Pubblicazione: (2026)
Outlier-robust neural network training: variation regularization meets trimmed loss to prevent functional breakdown
di: Okuno, Akifumi, et al.
Pubblicazione: (2023)
di: Okuno, Akifumi, et al.
Pubblicazione: (2023)
Crystal structure prediction using graph neural combinatorial optimization
di: Gerolymatos, Stavros, et al.
Pubblicazione: (2026)
di: Gerolymatos, Stavros, et al.
Pubblicazione: (2026)
Sufficient and Necessary Explanations (and What Lies in Between)
di: Bharti, Beepul, et al.
Pubblicazione: (2024)
di: Bharti, Beepul, et al.
Pubblicazione: (2024)
Randomness and signal propagation in physics-informed neural networks (PINNs): A neural PDE perspective
di: Tucny, Jean-Michel, et al.
Pubblicazione: (2025)
di: Tucny, Jean-Michel, et al.
Pubblicazione: (2025)
Planning in a recurrent neural network that plays Sokoban
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2024)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2024)
Understanding polysemanticity in neural networks through coding theory
di: Marshall, Simon C., et al.
Pubblicazione: (2024)
di: Marshall, Simon C., et al.
Pubblicazione: (2024)
Variational autoencoder-based neural network model compression
di: Cheng, Liang, et al.
Pubblicazione: (2024)
di: Cheng, Liang, et al.
Pubblicazione: (2024)
Conditional computation in neural networks: principles and research trends
di: Scardapane, Simone, et al.
Pubblicazione: (2024)
di: Scardapane, Simone, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Entropic Auto-Encoding via Implicit Free-Energy Minimization
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2026) -
Flashpoints Signal Hidden Inherent Instabilities in Land-Use Planning
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2023) -
Adaptive multiple optimal learning factors for neural network training
di: Challagundla, Jeshwanth
Pubblicazione: (2024) -
Transforming Design Spaces Using Pareto-Laplace Filters
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2024) -
Filters reveal emergent structure in computational morphogenesis
di: Aliahmadi, Hazhir, et al.
Pubblicazione: (2025)