Dynamic Learning Rate Scheduling based on Loss Changes Leads to Faster Convergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Subramanian, Shreyas, Krishnamoorthy, Bala, Murthy, Pranav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VERAFI: Verified Agentic Financial Intelligence through Neurosymbolic Policy Generation
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
On the Convergence of Loss and Uncertainty-based Active Learning Algorithms
von: Haimovich, Daniel, et al.
Veröffentlicht: (2023)
von: Haimovich, Daniel, et al.
Veröffentlicht: (2023)
Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
von: Jhandi, Polaris, et al.
Veröffentlicht: (2025)
von: Jhandi, Polaris, et al.
Veröffentlicht: (2025)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
Dynamic Masking Rate Schedules for MLM Pretraining
von: Ankner, Zachary, et al.
Veröffentlicht: (2023)
von: Ankner, Zachary, et al.
Veröffentlicht: (2023)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025)
SAIL: Faster-than-Demonstration Execution of Imitation Learning Policies
von: Arachchige, Nadun Ranawaka, et al.
Veröffentlicht: (2025)
von: Arachchige, Nadun Ranawaka, et al.
Veröffentlicht: (2025)
On Tuning Neural ODE for Stability, Consistency and Faster Convergence
von: Akhtar, Sheikh Waqas
Veröffentlicht: (2023)
von: Akhtar, Sheikh Waqas
Veröffentlicht: (2023)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
von: Defazio, Aaron
Veröffentlicht: (2026)
von: Defazio, Aaron
Veröffentlicht: (2026)
Double Successive Over-Relaxation Q-Learning with an Extension to Deep Reinforcement Learning
von: R, Shreyas S
Veröffentlicht: (2024)
von: R, Shreyas S
Veröffentlicht: (2024)
Cyclical Log Annealing as a Learning Rate Scheduler
von: Naveen, Philip
Veröffentlicht: (2024)
von: Naveen, Philip
Veröffentlicht: (2024)
Faster Convergence for Transformer Fine-tuning with Line Search Methods
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
von: Kenneweg, Philip, et al.
Veröffentlicht: (2024)
If You Want Coherence, Orchestrate a Team of Rivals: Multi-Agent Models of Organizational Intelligence
von: Vijayaraghavan, Gopal, et al.
Veröffentlicht: (2026)
von: Vijayaraghavan, Gopal, et al.
Veröffentlicht: (2026)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
von: Defazio, Aaron, et al.
Veröffentlicht: (2023)
von: Defazio, Aaron, et al.
Veröffentlicht: (2023)
Convergence Rate Maximization for Split Learning-based Control of EMG Prosthetic Devices
von: Marinova, Matea, et al.
Veröffentlicht: (2024)
von: Marinova, Matea, et al.
Veröffentlicht: (2024)
On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2023)
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2023)
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
von: Shen, Yikang, et al.
Veröffentlicht: (2024)
Meta-Sealing: A Revolutionizing Integrity Assurance Protocol for Transparent, Tamper-Proof, and Trustworthy AI System
von: Krishnamoorthy, Mahesh Vaijainthymala
Veröffentlicht: (2024)
von: Krishnamoorthy, Mahesh Vaijainthymala
Veröffentlicht: (2024)
Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation
von: Ye, Zilyu, et al.
Veröffentlicht: (2024)
von: Ye, Zilyu, et al.
Veröffentlicht: (2024)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
von: Schotthöfer, Steffen, et al.
Veröffentlicht: (2024)
von: Schotthöfer, Steffen, et al.
Veröffentlicht: (2024)
Representation Stability in a Minimal Continual Learning Agent
von: Subramanian, Vishnu
Veröffentlicht: (2026)
von: Subramanian, Vishnu
Veröffentlicht: (2026)
Heterogeneous Learning Rate Scheduling for Neural Architecture Search on Long-Tailed Datasets
von: Tang, Chenxia
Veröffentlicht: (2024)
von: Tang, Chenxia
Veröffentlicht: (2024)
The Practimum-Optimum Algorithm for Manufacturing Scheduling: A Paradigm Shift Leading to Breakthroughs in Scale and Performance
von: BenBassat, Moshe
Veröffentlicht: (2024)
von: BenBassat, Moshe
Veröffentlicht: (2024)
Learning NEAT Emergent Behaviors in Robot Swarms
von: Rajbhandari, Pranav, et al.
Veröffentlicht: (2023)
von: Rajbhandari, Pranav, et al.
Veröffentlicht: (2023)
Neural Optimizer Equation, Decay Function, and Learning Rate Schedule Joint Evolution
von: Morgan, Brandon, et al.
Veröffentlicht: (2024)
von: Morgan, Brandon, et al.
Veröffentlicht: (2024)
Improving Diffusion Models's Data-Corruption Resistance using Scheduled Pseudo-Huber Loss
von: Khrapov, Artem, et al.
Veröffentlicht: (2024)
von: Khrapov, Artem, et al.
Veröffentlicht: (2024)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Neuroscience-Inspired Memory Replay for Continual Learning: A Comparative Study of Predictive Coding and Backpropagation-Based Strategies
von: Nalagatla, Goutham, et al.
Veröffentlicht: (2025)
von: Nalagatla, Goutham, et al.
Veröffentlicht: (2025)
Anytime Training with Schedule-Free Spectral Optimization
von: Apte, Anuj, et al.
Veröffentlicht: (2026)
von: Apte, Anuj, et al.
Veröffentlicht: (2026)
STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Phat, et al.
Veröffentlicht: (2026)
Predictive Batch Scheduling: Accelerating Language Model Training Through Loss-Aware Sample Prioritization
von: Rasal, Sumedh
Veröffentlicht: (2026)
von: Rasal, Sumedh
Veröffentlicht: (2026)
Genetic Programming with Reinforcement Learning Trained Transformer for Real-World Dynamic Scheduling Problems
von: Chen, Xinan, et al.
Veröffentlicht: (2025)
von: Chen, Xinan, et al.
Veröffentlicht: (2025)
Certificate-Guided Evaluation of Reinforcement Learning Generalization
von: Subramanian, Vignesh, et al.
Veröffentlicht: (2026)
von: Subramanian, Vignesh, et al.
Veröffentlicht: (2026)
Adversarial Machine Learning: Attacks, Defenses, and Open Challenges
von: Jha, Pranav K
Veröffentlicht: (2025)
von: Jha, Pranav K
Veröffentlicht: (2025)
Federated Loss Exploration for Improved Convergence on Non-IID Data
von: Internò, Christian, et al.
Veröffentlicht: (2025)
von: Internò, Christian, et al.
Veröffentlicht: (2025)
On the Rate of Convergence of Kolmogorov-Arnold Network Regression Estimators
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VERAFI: Verified Agentic Financial Intelligence through Neurosymbolic Policy Generation
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026) -
Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025) -
On the Convergence of Loss and Uncertainty-based Active Learning Algorithms
von: Haimovich, Daniel, et al.
Veröffentlicht: (2023) -
Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
von: Jhandi, Polaris, et al.
Veröffentlicht: (2025)