On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Joo, Taejong, Xia, Wenhan, Kim, Cheolmin, Zhang, Ming, Ie, Eugene |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context
von: Joo, Taejong, et al.
Veröffentlicht: (2025)
von: Joo, Taejong, et al.
Veröffentlicht: (2025)
How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum
von: Lin, Chu-Cheng, et al.
Veröffentlicht: (2026)
von: Lin, Chu-Cheng, et al.
Veröffentlicht: (2026)
Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
von: Lin, Chu-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Chu-Cheng, et al.
Veröffentlicht: (2025)
The Surprising Effectiveness of Skip-Tuning in Diffusion Sampling
von: Ma, Jiajun, et al.
Veröffentlicht: (2024)
von: Ma, Jiajun, et al.
Veröffentlicht: (2024)
Graph of Agents: Principled Long Context Modeling by Emergent Multi-Agent Collaboration
von: Joo, Taejong, et al.
Veröffentlicht: (2025)
von: Joo, Taejong, et al.
Veröffentlicht: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
von: Zhang, Shenao, et al.
Veröffentlicht: (2025)
SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
von: Kallusky, Dominik, et al.
Veröffentlicht: (2025)
von: Kallusky, Dominik, et al.
Veröffentlicht: (2025)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
von: Hugessen, Adriana, et al.
Veröffentlicht: (2024)
von: Hugessen, Adriana, et al.
Veröffentlicht: (2024)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
von: Akyürek, Ekin, et al.
Veröffentlicht: (2024)
von: Akyürek, Ekin, et al.
Veröffentlicht: (2024)
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling
von: Haas, Moritz, et al.
Veröffentlicht: (2025)
von: Haas, Moritz, et al.
Veröffentlicht: (2025)
IW-GAE: Importance Weighted Group Accuracy Estimation for Improved Calibration and Model Selection in Unsupervised Domain Adaptation
von: Joo, Taejong, et al.
Veröffentlicht: (2023)
von: Joo, Taejong, et al.
Veröffentlicht: (2023)
Improving self-training under distribution shifts via anchored confidence with theoretical guarantees
von: Joo, Taejong, et al.
Veröffentlicht: (2024)
von: Joo, Taejong, et al.
Veröffentlicht: (2024)
MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
von: Iacob, Alex, et al.
Veröffentlicht: (2025)
von: Iacob, Alex, et al.
Veröffentlicht: (2025)
Surprisal Driven $k$-NN for Robust and Interpretable Nonparametric Learning
von: Banerjee, Amartya, et al.
Veröffentlicht: (2023)
von: Banerjee, Amartya, et al.
Veröffentlicht: (2023)
Anomaly Detection with Adaptive and Aggressive Rejection for Contaminated Training Data
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
World Model Robustness via Surprise Recognition
von: Zollicoffer, Geigh, et al.
Veröffentlicht: (2025)
von: Zollicoffer, Geigh, et al.
Veröffentlicht: (2025)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
von: Liu, Huihan, et al.
Veröffentlicht: (2026)
von: Liu, Huihan, et al.
Veröffentlicht: (2026)
LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models
von: Wei, Chenxing, et al.
Veröffentlicht: (2026)
von: Wei, Chenxing, et al.
Veröffentlicht: (2026)
On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models
von: Farhat, Sean, et al.
Veröffentlicht: (2024)
von: Farhat, Sean, et al.
Veröffentlicht: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
Guiding Masked Representation Learning to Capture Spatio-Temporal Relationship of Electrocardiogram
von: Na, Yeongyeon, et al.
Veröffentlicht: (2024)
von: Na, Yeongyeon, et al.
Veröffentlicht: (2024)
Batch Acquisition Function Evaluations and Decouple Optimizer Updates for Faster Bayesian Optimization
von: Irie, Kaichi, et al.
Veröffentlicht: (2025)
von: Irie, Kaichi, et al.
Veröffentlicht: (2025)
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
von: Wu, Wenhan, et al.
Veröffentlicht: (2025)
von: Wu, Wenhan, et al.
Veröffentlicht: (2025)
DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
von: Zhu, WenZhuo, et al.
Veröffentlicht: (2025)
von: Zhu, WenZhuo, et al.
Veröffentlicht: (2025)
Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection
von: Özer, Kadir-Kaan, et al.
Veröffentlicht: (2026)
von: Özer, Kadir-Kaan, et al.
Veröffentlicht: (2026)
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
von: Cheng, Peng, et al.
Veröffentlicht: (2026)
von: Cheng, Peng, et al.
Veröffentlicht: (2026)
Automated Random Embedding for Practical Bayesian Optimization with Unknown Effective Dimension
von: Qian, Hong, et al.
Veröffentlicht: (2026)
von: Qian, Hong, et al.
Veröffentlicht: (2026)
Simple and Effective Masked Diffusion Language Models
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2024)
von: Sahoo, Subham Sekhar, et al.
Veröffentlicht: (2024)
On the Surprising Effectiveness of Attention Transfer for Vision Transformers
von: Li, Alexander C., et al.
Veröffentlicht: (2024)
von: Li, Alexander C., et al.
Veröffentlicht: (2024)
Celo2: Towards Learned Optimization Free Lunch
von: Moudgil, Abhinav, et al.
Veröffentlicht: (2026)
von: Moudgil, Abhinav, et al.
Veröffentlicht: (2026)
Iterative Mask Filling: An Effective Text Augmentation Method Using Masked Language Modeling
von: Kesgin, Himmet Toprak, et al.
Veröffentlicht: (2024)
von: Kesgin, Himmet Toprak, et al.
Veröffentlicht: (2024)
CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning
von: Wang, Shuo, et al.
Veröffentlicht: (2026)
von: Wang, Shuo, et al.
Veröffentlicht: (2026)
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2025)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
von: Kim, Hoyong, et al.
Veröffentlicht: (2023)
von: Kim, Hoyong, et al.
Veröffentlicht: (2023)
MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Models
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
Mask-PINNs: Mitigating Internal Covariate Shift in Physics-Informed Neural Networks
von: Jiang, Feilong, et al.
Veröffentlicht: (2025)
von: Jiang, Feilong, et al.
Veröffentlicht: (2025)
To Predict or Not To Predict? Proportionally Masked Autoencoders for Tabular Data Imputation
von: Kim, Jungkyu, et al.
Veröffentlicht: (2024)
von: Kim, Jungkyu, et al.
Veröffentlicht: (2024)
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2026)
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2026)
Refining Adaptive Zeroth-Order Optimization at Ease
von: Shu, Yao, et al.
Veröffentlicht: (2025)
von: Shu, Yao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context
von: Joo, Taejong, et al.
Veröffentlicht: (2025) -
How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum
von: Lin, Chu-Cheng, et al.
Veröffentlicht: (2026) -
Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
von: Lin, Chu-Cheng, et al.
Veröffentlicht: (2025) -
The Surprising Effectiveness of Skip-Tuning in Diffusion Sampling
von: Ma, Jiajun, et al.
Veröffentlicht: (2024) -
Graph of Agents: Principled Long Context Modeling by Emergent Multi-Agent Collaboration
von: Joo, Taejong, et al.
Veröffentlicht: (2025)