Improving Line Search Methods for Large Scale Neural Network Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kenneweg, Philip, Kenneweg, Tristan, Hammer, Barbara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faster Convergence for Transformer Fine-tuning with Line Search Methods
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
Neural Architecture Search for Sentence Classification with BERT
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2024)
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2024)
JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2025)
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Remaining Lifespan Prediction from Images
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2025)
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2025)
Who Has The Final Say? Conformity Dynamics in ChatGPT's Selections
von: Arlinghaus, Clarissa Sabrina, et al.
Veröffentlicht: (2025)
von: Arlinghaus, Clarissa Sabrina, et al.
Veröffentlicht: (2025)
Towards Understanding the Influence of Training Samples on Explanations
von: Artelt, André, et al.
Veröffentlicht: (2024)
von: Artelt, André, et al.
Veröffentlicht: (2024)
Physics-Informed Graph Neural Networks for Water Distribution Systems
von: Ashraf, Inaam, et al.
Veröffentlicht: (2024)
von: Ashraf, Inaam, et al.
Veröffentlicht: (2024)
Fairness-Enhancing Ensemble Classification in Water Distribution Networks
von: Strotherm, Janine, et al.
Veröffentlicht: (2024)
von: Strotherm, Janine, et al.
Veröffentlicht: (2024)
Extending Fair Null-Space Projections for Continuous Attributes to Kernel Methods
von: Störck, Felix, et al.
Veröffentlicht: (2025)
von: Störck, Felix, et al.
Veröffentlicht: (2025)
Federated Loss Exploration for Improved Convergence on Non-IID Data
von: Internò, Christian, et al.
Veröffentlicht: (2025)
von: Internò, Christian, et al.
Veröffentlicht: (2025)
RapidGNN: Energy and Communication-Efficient Distributed Training on Large-Scale Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
Monte Carlo Permutation Search
von: Cazenave, Tristan
Veröffentlicht: (2025)
von: Cazenave, Tristan
Veröffentlicht: (2025)
Exact Computation of Any-Order Shapley Interactions for Graph Neural Networks
von: Muschalik, Maximilian, et al.
Veröffentlicht: (2025)
von: Muschalik, Maximilian, et al.
Veröffentlicht: (2025)
Improved Training of Physics-Informed Neural Networks with Model Ensembles
von: Haitsiukevich, Katsiaryna, et al.
Veröffentlicht: (2022)
von: Haitsiukevich, Katsiaryna, et al.
Veröffentlicht: (2022)
DGPO: RL-Steered Graph Diffusion for Neural Architecture Generation
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2026)
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2026)
Neural Network Verification with PyRAT
von: Lemesle, Augustin, et al.
Veröffentlicht: (2024)
von: Lemesle, Augustin, et al.
Veröffentlicht: (2024)
Continuous Fair SMOTE -- Fairness-Aware Stream Learning from Imbalanced Data
von: Lammers, Kathrin, et al.
Veröffentlicht: (2025)
von: Lammers, Kathrin, et al.
Veröffentlicht: (2025)
IDInit: A Universal and Stable Initialization Method for Neural Network Training
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
von: Do, Khoi, et al.
Veröffentlicht: (2023)
von: Do, Khoi, et al.
Veröffentlicht: (2023)
Decoupling Search and Learning in Neural Net Training
von: Vegesna, Akshay, et al.
Veröffentlicht: (2025)
von: Vegesna, Akshay, et al.
Veröffentlicht: (2025)
Debiasing Sentence Embedders through Contrastive Word Pairs
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
Training Artificial Neural Networks by Coordinate Search Algorithm
von: Rokhsatyazdi, Ehsan, et al.
Veröffentlicht: (2024)
von: Rokhsatyazdi, Ehsan, et al.
Veröffentlicht: (2024)
The Effect of Data Poisoning on Counterfactual Explanations
von: Artelt, André, et al.
Veröffentlicht: (2024)
von: Artelt, André, et al.
Veröffentlicht: (2024)
One-Class Intrusion Detection with Dynamic Graphs
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2025)
von: Liuliakov, Aleksei, et al.
Veröffentlicht: (2025)
Position: Rethinking Post-Hoc Search-Based Neural Approaches for Solving Large-Scale Traveling Salesman Problems
von: Xia, Yifan, et al.
Veröffentlicht: (2024)
von: Xia, Yifan, et al.
Veröffentlicht: (2024)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
von: Ma, Yuhan, et al.
Veröffentlicht: (2024)
von: Ma, Yuhan, et al.
Veröffentlicht: (2024)
LOGIN: A Large Language Model Consulted Graph Neural Network Training Framework
von: Qiao, Yiran, et al.
Veröffentlicht: (2024)
von: Qiao, Yiran, et al.
Veröffentlicht: (2024)
Graph Neural Architecture Search with GPT-4
von: Wang, Haishuai, et al.
Veröffentlicht: (2023)
von: Wang, Haishuai, et al.
Veröffentlicht: (2023)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025)
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025)
Scaling Combinatorial Optimization Neural Improvement Heuristics with Online Search and Adaptation
von: Verdù, Federico Julian Camerota, et al.
Veröffentlicht: (2024)
von: Verdù, Federico Julian Camerota, et al.
Veröffentlicht: (2024)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
von: Cui, Brandon, et al.
Veröffentlicht: (2026)
von: Cui, Brandon, et al.
Veröffentlicht: (2026)
Fast Training of Sinusoidal Neural Fields via Scaling Initialization
von: Yeom, Taesun, et al.
Veröffentlicht: (2024)
von: Yeom, Taesun, et al.
Veröffentlicht: (2024)
On Newton's Method to Unlearn Neural Networks
von: Bui, Nhung, et al.
Veröffentlicht: (2024)
von: Bui, Nhung, et al.
Veröffentlicht: (2024)
Training Neural Networks for Modularity aids Interpretability
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Gradient-Free Training of Quantized Neural Networks
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
von: Cohen, Noa, et al.
Veröffentlicht: (2024)
Z-Error Loss for Training Neural Networks
von: Godin, Guillaume
Veröffentlicht: (2025)
von: Godin, Guillaume
Veröffentlicht: (2025)
Automatic Stability and Recovery for Neural Network Training
von: Or, Barak
Veröffentlicht: (2026)
von: Or, Barak
Veröffentlicht: (2026)
Ähnliche Einträge
-
Faster Convergence for Transformer Fine-tuning with Line Search Methods
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024) -
No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024) -
Neural Architecture Search for Sentence Classification with BERT
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024) -
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024) -
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup
von: Kenneweg, Tristan, et al.
Veröffentlicht: (2024)