AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Fu-Ming, Fan, Yingfang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Robustness of Decision-Focused Learning
by: Farhat, Yehya
Published: (2023)
by: Farhat, Yehya
Published: (2023)
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
by: Heitzig, Jobst, et al.
Published: (2025)
by: Heitzig, Jobst, et al.
Published: (2025)
Self-Composing Neural Operators with Depth and Accuracy Scaling via Adaptive Train-and-Unroll Approach
by: He, Juncai, et al.
Published: (2025)
by: He, Juncai, et al.
Published: (2025)
GraphCNNpred: A stock market indices prediction using a Graph based deep learning system
by: Jin, Yuhui
Published: (2024)
by: Jin, Yuhui
Published: (2024)
Detection of Chagas Disease from the ECG: The George B. Moody PhysioNet Challenge 2025
by: Reyna, Matthew A., et al.
Published: (2025)
by: Reyna, Matthew A., et al.
Published: (2025)
From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production
by: Shlomov, Segev, et al.
Published: (2025)
by: Shlomov, Segev, et al.
Published: (2025)
Any four real numbers are on all fours with analogy
by: Lepage, Yves, et al.
Published: (2024)
by: Lepage, Yves, et al.
Published: (2024)
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Distributional Drift Adaptation with Temporal Conditional Variational Autoencoder for Multivariate Time Series Forecasting
by: He, Hui, et al.
Published: (2022)
by: He, Hui, et al.
Published: (2022)
Toward Storage-Aware Learning with Compressed Data An Empirical Exploratory Study on JPEG
by: Lee, Kichang, et al.
Published: (2025)
by: Lee, Kichang, et al.
Published: (2025)
Emergence of heavy tails in homogenized stochastic gradient descent
by: Jiao, Zhe, et al.
Published: (2024)
by: Jiao, Zhe, et al.
Published: (2024)
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
by: Xie, Zeyu, et al.
Published: (2026)
by: Xie, Zeyu, et al.
Published: (2026)
Moving Object Proposals with Deep Learned Optical Flow for Video Object Segmentation
by: Shi, Ge, et al.
Published: (2024)
by: Shi, Ge, et al.
Published: (2024)
RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation
by: Huang, Yuhao, et al.
Published: (2026)
by: Huang, Yuhao, et al.
Published: (2026)
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks
by: Ahmadian, Rouhollah, et al.
Published: (2024)
by: Ahmadian, Rouhollah, et al.
Published: (2024)
The Maximum Number of Sets for 12 Cards is 14
by: Stevens, Justin, et al.
Published: (2025)
by: Stevens, Justin, et al.
Published: (2025)
Big Data Intelligence Using Distributed Deep Neural Networks
by: Ongati, Felix, et al.
Published: (2019)
by: Ongati, Felix, et al.
Published: (2019)
Verifiable Dropout: Turning Randomness into a Verifiable Claim
by: Lee, Kichang, et al.
Published: (2025)
by: Lee, Kichang, et al.
Published: (2025)
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
by: Alizadeh, Meysam, et al.
Published: (2025)
by: Alizadeh, Meysam, et al.
Published: (2025)
Robust Multivariate Time Series Forecasting against Intra- and Inter-Series Transitional Shift
by: He, Hui, et al.
Published: (2024)
by: He, Hui, et al.
Published: (2024)
Enabling Trustworthy Federated Learning in Industrial IoT: Bridging the Gap Between Interpretability and Robustness
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
by: Jagatheesaperumal, Senthil Kumar, et al.
Published: (2024)
ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal Processing Analysis for Sounds, & Strains Emerging Technology
by: Ali, Ratul, et al.
Published: (2023)
by: Ali, Ratul, et al.
Published: (2023)
Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
ViMo: Generating Motions from Casual Videos
by: Qiu, Liangdong, et al.
Published: (2024)
by: Qiu, Liangdong, et al.
Published: (2024)
Temperature Scaling Attack Disrupting Model Confidence in Federated Learning
by: Lee, Kichang, et al.
Published: (2026)
by: Lee, Kichang, et al.
Published: (2026)
PairHuman: A High-Fidelity Photographic Dataset for Customized Dual-Person Generation
by: Pan, Ting, et al.
Published: (2025)
by: Pan, Ting, et al.
Published: (2025)
VCRScore: Image captioning metric based on V\&L Transformers, CLIP, and precision-recall
by: Ruiz, Guillermo, et al.
Published: (2025)
by: Ruiz, Guillermo, et al.
Published: (2025)
Automated Model Selection for Generalized Linear Models
by: Schwendinger, Benjamin, et al.
Published: (2024)
by: Schwendinger, Benjamin, et al.
Published: (2024)
Efficient Sliced Wasserstein Distance Computation via Adaptive Bayesian Optimization
by: Acharya, Manish, et al.
Published: (2025)
by: Acharya, Manish, et al.
Published: (2025)
Reinforced Inverse Scattering
by: Jiang, Hanyang, et al.
Published: (2022)
by: Jiang, Hanyang, et al.
Published: (2022)
Conversion rate prediction in online advertising: modeling techniques, performance evaluation and future directions
by: Xue, Tao, et al.
Published: (2025)
by: Xue, Tao, et al.
Published: (2025)
Toward a benchmark for CTR prediction in online advertising: datasets, evaluation protocols and perspectives
by: Gao, Shan, et al.
Published: (2025)
by: Gao, Shan, et al.
Published: (2025)
STAR: Speech-to-Audio Generation via Representation Learning
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Exploring ChatGPT and its Impact on Society
by: Haque, Md. Asraful, et al.
Published: (2024)
by: Haque, Md. Asraful, et al.
Published: (2024)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
Contemporary Agent Technology: LLM-Driven Advancements vs Classic Multi-Agent Systems
by: Bădică, Costin, et al.
Published: (2025)
by: Bădică, Costin, et al.
Published: (2025)
Adaptive Dataset Quantization: A New Direction for Dataset Pruning
by: Yu, Chenyue, et al.
Published: (2025)
by: Yu, Chenyue, et al.
Published: (2025)
Similar Items
-
On the Robustness of Decision-Focused Learning
by: Farhat, Yehya
Published: (2023) -
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
by: Heitzig, Jobst, et al.
Published: (2025) -
Self-Composing Neural Operators with Depth and Accuracy Scaling via Adaptive Train-and-Unroll Approach
by: He, Juncai, et al.
Published: (2025) -
GraphCNNpred: A stock market indices prediction using a Graph based deep learning system
by: Jin, Yuhui
Published: (2024) -
Detection of Chagas Disease from the ECG: The George B. Moody PhysioNet Challenge 2025
by: Reyna, Matthew A., et al.
Published: (2025)