Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Sow, Daouda, Woisetschläger, Herbert, Bulusu, Saikiran, Wang, Shiqiang, Jacobsen, Hans-Arno, Liang, Yingbin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
by: Wang, Shiqiang, et al.
Published: (2026)
by: Wang, Shiqiang, et al.
Published: (2026)
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
by: Woisetschläger, Herbert, et al.
Published: (2025)
by: Woisetschläger, Herbert, et al.
Published: (2025)
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
by: Mulitze, Felix, et al.
Published: (2025)
by: Mulitze, Felix, et al.
Published: (2025)
A Survey on Efficient Federated Learning Methods for Foundation Model Training
by: Woisetschläger, Herbert, et al.
Published: (2024)
by: Woisetschläger, Herbert, et al.
Published: (2024)
MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees
by: Zhang, Ryan, et al.
Published: (2024)
by: Zhang, Ryan, et al.
Published: (2024)
Agentic Performance at the Edge: Insights from Benchmarking
by: Wang, Shiqiang, et al.
Published: (2026)
by: Wang, Shiqiang, et al.
Published: (2026)
Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization
by: Chen, Xuxi, et al.
Published: (2024)
by: Chen, Xuxi, et al.
Published: (2024)
Federated Learning Priorities Under the European Union Artificial Intelligence Act
by: Woisetschläger, Herbert, et al.
Published: (2024)
by: Woisetschläger, Herbert, et al.
Published: (2024)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Algorithm Design for Online Meta-Learning with Task Boundary Detection
by: Sow, Daouda, et al.
Published: (2023)
by: Sow, Daouda, et al.
Published: (2023)
Federated Learning and AI Regulation in the European Union: Who is Responsible? -- An Interdisciplinary Analysis
by: Woisetschläger, Herbert, et al.
Published: (2024)
by: Woisetschläger, Herbert, et al.
Published: (2024)
SIGN: Schema-Induced Games for Naming
by: Zhang, Ryan, et al.
Published: (2025)
by: Zhang, Ryan, et al.
Published: (2025)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
by: Yang, Puning, et al.
Published: (2025)
by: Yang, Puning, et al.
Published: (2025)
Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games
by: Feng, Songtao, et al.
Published: (2023)
by: Feng, Songtao, et al.
Published: (2023)
Adversarial Robustness of Partitioned Quantum Classifiers
by: Kananian, Pouya, et al.
Published: (2025)
by: Kananian, Pouya, et al.
Published: (2025)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
by: Galinkin, Erick, et al.
Published: (2024)
by: Galinkin, Erick, et al.
Published: (2024)
When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining
by: Karpukhin, Ivan, et al.
Published: (2026)
by: Karpukhin, Ivan, et al.
Published: (2026)
IFFair: Influence Function-driven Sample Reweighting for Fair Classification
by: Yang, Jingran, et al.
Published: (2025)
by: Yang, Jingran, et al.
Published: (2025)
Beyond Losses Reweighting: Empowering Multi-Task Learning via the Generalization Perspective
by: Phan, Hoang, et al.
Published: (2022)
by: Phan, Hoang, et al.
Published: (2022)
FaLW: A Forgetting-aware Loss Reweighting for Long-tailed Unlearning
by: Yu, Liheng, et al.
Published: (2026)
by: Yu, Liheng, et al.
Published: (2026)
Choosing a Classical Planner with Graph Neural Networks
by: Vatter, Jana, et al.
Published: (2024)
by: Vatter, Jana, et al.
Published: (2024)
GALA: Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis?
by: Tian, Yifang, et al.
Published: (2025)
by: Tian, Yifang, et al.
Published: (2025)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Improved Constrained Generation by Bridging Pretrained Generative Models
by: Liang, Xiaoxuan, et al.
Published: (2026)
by: Liang, Xiaoxuan, et al.
Published: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View
by: Wang, Jinping, et al.
Published: (2026)
by: Wang, Jinping, et al.
Published: (2026)
LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo
by: Sawkar, Mandira, et al.
Published: (2025)
by: Sawkar, Mandira, et al.
Published: (2025)
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
by: Geng, Jiahui, et al.
Published: (2025)
by: Geng, Jiahui, et al.
Published: (2025)
A Step Toward Federated Pretraining of Multimodal Large Language Models
by: Xiong, Baochen, et al.
Published: (2026)
by: Xiong, Baochen, et al.
Published: (2026)
Multimodal Physiological Signals Representation Learning via Multiscale Contrasting for Depression Recognition
by: Shao, Kai, et al.
Published: (2024)
by: Shao, Kai, et al.
Published: (2024)
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
by: Shi, Ming, et al.
Published: (2023)
by: Shi, Ming, et al.
Published: (2023)
Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
Joint Continual Learning of Local Language Models and Cloud Offloading Decisions with Budget Constraints
by: Chen, Evan, et al.
Published: (2026)
by: Chen, Evan, et al.
Published: (2026)
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
by: Portes, Jacob, et al.
Published: (2025)
by: Portes, Jacob, et al.
Published: (2025)
Geometry Preserving Loss Functions Promote Improved Adaptation of Blackbox Generative Model
by: Mitra, Sinjini, et al.
Published: (2026)
by: Mitra, Sinjini, et al.
Published: (2026)
Sample Transform Cost-Based Training-Free Hallucination Detector for Large Language Models
by: Ding, Zeyang, et al.
Published: (2026)
by: Ding, Zeyang, et al.
Published: (2026)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
Robust Uncertainty Quantification for Self-Evolving Large Language Models via Continual Domain Pretraining
by: Zhou, Xiaofan, et al.
Published: (2025)
by: Zhou, Xiaofan, et al.
Published: (2025)
Similar Items
-
Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
by: Wang, Shiqiang, et al.
Published: (2026) -
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
by: Woisetschläger, Herbert, et al.
Published: (2025) -
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
by: Mulitze, Felix, et al.
Published: (2025) -
A Survey on Efficient Federated Learning Methods for Foundation Model Training
by: Woisetschläger, Herbert, et al.
Published: (2024) -
MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees
by: Zhang, Ryan, et al.
Published: (2024)