Step-size Optimization for Continual Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Degris, Thomas, Javed, Khurram, Sharifnassab, Arsalan, Liu, Yuxin, Sutton, Richard |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
por: Sharifnassab, Arsalan, et al.
Publicado: (2024)
por: Sharifnassab, Arsalan, et al.
Publicado: (2024)
Swift-Sarsa: Fast and Robust Linear Control
por: Javed, Khurram, et al.
Publicado: (2025)
por: Javed, Khurram, et al.
Publicado: (2025)
Intentional Updates for Streaming Reinforcement Learning
por: Sharifnassab, Arsalan, et al.
Publicado: (2026)
por: Sharifnassab, Arsalan, et al.
Publicado: (2026)
Soft Preference Optimization: Aligning Language Models to Expert Distributions
por: Sharifnassab, Arsalan, et al.
Publicado: (2024)
por: Sharifnassab, Arsalan, et al.
Publicado: (2024)
Order Optimal Bounds for One-Shot Federated Learning over non-Convex Loss Functions
por: Sharifnassab, Arsalan, et al.
Publicado: (2021)
por: Sharifnassab, Arsalan, et al.
Publicado: (2021)
RIFT: A Scalable Methodology for LLM Accelerator Fault Assessment using Reinforcement Learning
por: Khalil, Khurram, et al.
Publicado: (2025)
por: Khalil, Khurram, et al.
Publicado: (2025)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
por: Yamin, Khurram, et al.
Publicado: (2025)
por: Yamin, Khurram, et al.
Publicado: (2025)
Machine Learning-based Android Intrusion Detection System
por: Tahreem, Madiha, et al.
Publicado: (2024)
por: Tahreem, Madiha, et al.
Publicado: (2024)
SRTFD: Scalable Real-Time Fault Diagnosis through Online Continual Learning
por: Zhao, Dandan, et al.
Publicado: (2024)
por: Zhao, Dandan, et al.
Publicado: (2024)
SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation
por: Vetcha, Nitin, et al.
Publicado: (2026)
por: Vetcha, Nitin, et al.
Publicado: (2026)
Continual Optimization with Symmetry Teleportation for Multi-Task Learning
por: Zhou, Zhipeng, et al.
Publicado: (2025)
por: Zhou, Zhipeng, et al.
Publicado: (2025)
Structure-based RNA Design by Step-wise Optimization of Latent Diffusion Model
por: Si, Qi, et al.
Publicado: (2026)
por: Si, Qi, et al.
Publicado: (2026)
Explainability and Continual Learning meet Federated Learning at the Network Edge
por: Tsouparopoulos, Thomas, et al.
Publicado: (2025)
por: Tsouparopoulos, Thomas, et al.
Publicado: (2025)
Reward Centering
por: Naik, Abhishek, et al.
Publicado: (2024)
por: Naik, Abhishek, et al.
Publicado: (2024)
Ferret: An Efficient Online Continual Learning Framework under Varying Memory Constraints
por: Zhou, Yuhao, et al.
Publicado: (2025)
por: Zhou, Yuhao, et al.
Publicado: (2025)
Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
por: Liu, Zequn, et al.
Publicado: (2026)
por: Liu, Zequn, et al.
Publicado: (2026)
Optimizing Mastery Learning by Fast-Forwarding Over-Practice Steps
por: Xia, Meng, et al.
Publicado: (2025)
por: Xia, Meng, et al.
Publicado: (2025)
Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
por: Fan, Jiajun, et al.
Publicado: (2025)
por: Fan, Jiajun, et al.
Publicado: (2025)
MANGO: Meta-Adaptive Network Gradient Optimization for Online Continual Learning
por: Awasthi, Ankita, et al.
Publicado: (2026)
por: Awasthi, Ankita, et al.
Publicado: (2026)
Continuous Invariance Learning
por: Lin, Yong, et al.
Publicado: (2023)
por: Lin, Yong, et al.
Publicado: (2023)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
por: Ding, Zihan, et al.
Publicado: (2024)
por: Ding, Zihan, et al.
Publicado: (2024)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
por: Chen, Luoxin, et al.
Publicado: (2026)
por: Chen, Luoxin, et al.
Publicado: (2026)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
por: Lin, Yen-Ting, et al.
Publicado: (2025)
por: Lin, Yen-Ting, et al.
Publicado: (2025)
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
por: Lei, Shiye, et al.
Publicado: (2026)
por: Lei, Shiye, et al.
Publicado: (2026)
Enhancing Autonomous Online Intrusion Detection for IoT with Balanced Learning, Reliable Pseudo-Labels, and Lightweight Architectures
por: Afzaal, Hanzala, et al.
Publicado: (2026)
por: Afzaal, Hanzala, et al.
Publicado: (2026)
Uncertainty-aware Human Mobility Modeling and Anomaly Detection
por: Wen, Haomin, et al.
Publicado: (2024)
por: Wen, Haomin, et al.
Publicado: (2024)
Low-redundancy Distillation for Continual Learning
por: Liu, RuiQi, et al.
Publicado: (2023)
por: Liu, RuiQi, et al.
Publicado: (2023)
A Survey on Deep Tabular Learning
por: Somvanshi, Shriyank, et al.
Publicado: (2024)
por: Somvanshi, Shriyank, et al.
Publicado: (2024)
Explainable AI-Guided Efficient Approximate DNN Generation for Multi-Pod Systolic Arrays
por: Siddique, Ayesha, et al.
Publicado: (2025)
por: Siddique, Ayesha, et al.
Publicado: (2025)
Learning Dynamic Representations via An Optimally-Weighted Maximum Mean Discrepancy Optimization Framework for Continual Learning
por: Huang, KaiHui, et al.
Publicado: (2025)
por: Huang, KaiHui, et al.
Publicado: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024)
por: Lai, Xin, et al.
Publicado: (2024)
The Importance of Being Lazy: Scaling Limits of Continual Learning
por: Graldi, Jacopo, et al.
Publicado: (2025)
por: Graldi, Jacopo, et al.
Publicado: (2025)
Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
por: Chen, Jack, et al.
Publicado: (2025)
por: Chen, Jack, et al.
Publicado: (2025)
Continual Learning as Computationally Constrained Reinforcement Learning
por: Kumar, Saurabh, et al.
Publicado: (2023)
por: Kumar, Saurabh, et al.
Publicado: (2023)
Continuous-Utility Direct Preference Optimization
por: Mohsin, Muhammad Ahmed, et al.
Publicado: (2026)
por: Mohsin, Muhammad Ahmed, et al.
Publicado: (2026)
Take a Step and Reconsider: Sequence Decoding for Self-Improved Neural Combinatorial Optimization
por: Pirnay, Jonathan, et al.
Publicado: (2024)
por: Pirnay, Jonathan, et al.
Publicado: (2024)
SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
por: Xu, Yuyang, et al.
Publicado: (2025)
por: Xu, Yuyang, et al.
Publicado: (2025)
Escaping Optimization Stagnation: Taking Steps Beyond Task Arithmetic via Difference Vectors
por: Wang, Jinping, et al.
Publicado: (2025)
por: Wang, Jinping, et al.
Publicado: (2025)
Dual-CBA: Improving Online Continual Learning via Dual Continual Bias Adaptors from a Bi-level Optimization Perspective
por: Wang, Quanziang, et al.
Publicado: (2024)
por: Wang, Quanziang, et al.
Publicado: (2024)
Ejemplares similares
-
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
por: Sharifnassab, Arsalan, et al.
Publicado: (2024) -
Swift-Sarsa: Fast and Robust Linear Control
por: Javed, Khurram, et al.
Publicado: (2025) -
Intentional Updates for Streaming Reinforcement Learning
por: Sharifnassab, Arsalan, et al.
Publicado: (2026) -
Soft Preference Optimization: Aligning Language Models to Expert Distributions
por: Sharifnassab, Arsalan, et al.
Publicado: (2024) -
Order Optimal Bounds for One-Shot Federated Learning over non-Convex Loss Functions
por: Sharifnassab, Arsalan, et al.
Publicado: (2021)