An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Qizhen, Garg, Ankush, Foerster, Jakob, Chatterji, Niladri, Malik, Kshitiz, Lewis, Mike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
von: Frei, Spencer, et al.
Veröffentlicht: (2022)
von: Frei, Spencer, et al.
Veröffentlicht: (2022)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
von: Roberts, Nicholas, et al.
Veröffentlicht: (2025)
von: Roberts, Nicholas, et al.
Veröffentlicht: (2025)
BTS: Harmonizing Specialized Experts into a Generalist LLM
von: Zhang, Qizhen, et al.
Veröffentlicht: (2025)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2025)
Noisy Zero-Shot Coordination: Breaking The Common Knowledge Assumption In Zero-Shot Coordination Games
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
von: Anwar, Usman, et al.
Veröffentlicht: (2024)
Analysing the Sample Complexity of Opponent Shaping
von: Fung, Kitty, et al.
Veröffentlicht: (2024)
von: Fung, Kitty, et al.
Veröffentlicht: (2024)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
JaxUED: A simple and useable UED library in Jax
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
von: Coward, Samuel, et al.
Veröffentlicht: (2024)
Learning to Forget with Information Divergence Reweighted Objectives for Noisy Labels
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2025)
von: Birrell, Jeremiah, et al.
Veröffentlicht: (2025)
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning
von: Chatterji, Satchit, et al.
Veröffentlicht: (2024)
von: Chatterji, Satchit, et al.
Veröffentlicht: (2024)
Performance of Small Language Model Pretraining on FABRIC: An Empirical Study
von: Rao, Praveen
Veröffentlicht: (2026)
von: Rao, Praveen
Veröffentlicht: (2026)
Improving Regret Approximation for Unsupervised Dynamic Environment Generation
von: Mead, Harry, et al.
Veröffentlicht: (2026)
von: Mead, Harry, et al.
Veröffentlicht: (2026)
Empirical Risk Minimization with $f$-Divergence Regularization
von: Daunas, Francisco, et al.
Veröffentlicht: (2026)
von: Daunas, Francisco, et al.
Veröffentlicht: (2026)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
Loss Functions and Operators Generated by f-Divergences
von: Roulet, Vincent, et al.
Veröffentlicht: (2025)
von: Roulet, Vincent, et al.
Veröffentlicht: (2025)
Revisiting Auxiliary Losses for Conditional Depth Routing: An Empirical Study
von: Lin, Qingwei
Veröffentlicht: (2026)
von: Lin, Qingwei
Veröffentlicht: (2026)
Accelerated Smoothing: A Scalable Approach to Randomized Smoothing
von: Bhardwaj, Devansh, et al.
Veröffentlicht: (2024)
von: Bhardwaj, Devansh, et al.
Veröffentlicht: (2024)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
von: Nishimori, Soichiro, et al.
Veröffentlicht: (2025)
PARDEN, Can You Repeat That? Defending against Jailbreaks via Repetition
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
HelloFresh: LLM Evaluations on Streams of Real-World Human Editorial Actions across X Community Notes and Wikipedia edits
von: Franzmeyer, Tim, et al.
Veröffentlicht: (2024)
von: Franzmeyer, Tim, et al.
Veröffentlicht: (2024)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
General Formulation and PCL-Analysis for Restless Bandits with Limited Observability
von: Liu, Keqin, et al.
Veröffentlicht: (2023)
von: Liu, Keqin, et al.
Veröffentlicht: (2023)
Decoupled Kullback-Leibler Divergence Loss
von: Cui, Jiequan, et al.
Veröffentlicht: (2023)
von: Cui, Jiequan, et al.
Veröffentlicht: (2023)
A Model-Based Solution to the Offline Multi-Agent Reinforcement Learning Coordination Problem
von: Barde, Paul, et al.
Veröffentlicht: (2023)
von: Barde, Paul, et al.
Veröffentlicht: (2023)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
von: Zhang, Zishi, et al.
Veröffentlicht: (2026)
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
von: Matthews, Michael, et al.
Veröffentlicht: (2024)
von: Matthews, Michael, et al.
Veröffentlicht: (2024)
Introducing Fractional Classification Loss for Robust Learning with Noisy Labels
von: Kurucu, Mert Can, et al.
Veröffentlicht: (2025)
von: Kurucu, Mert Can, et al.
Veröffentlicht: (2025)
Robust Loss Functions for Training Decision Trees with Noisy Labels
von: Wilton, Jonathan, et al.
Veröffentlicht: (2023)
von: Wilton, Jonathan, et al.
Veröffentlicht: (2023)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy Labels
von: Hu, Ruofan, et al.
Veröffentlicht: (2025)
von: Hu, Ruofan, et al.
Veröffentlicht: (2025)
Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima
von: Chen, Huanran, et al.
Veröffentlicht: (2026)
von: Chen, Huanran, et al.
Veröffentlicht: (2026)
Distribution-Free Pretraining of Classification Losses via Evolutionary Dynamics
von: Xiang, Meng, et al.
Veröffentlicht: (2026)
von: Xiang, Meng, et al.
Veröffentlicht: (2026)
On the connection between Noise-Contrastive Estimation and Contrastive Divergence
von: Olmin, Amanda, et al.
Veröffentlicht: (2024)
von: Olmin, Amanda, et al.
Veröffentlicht: (2024)
Predicting Large Model Test Losses with a Noisy Quadratic System
von: Li, Chuning, et al.
Veröffentlicht: (2026)
von: Li, Chuning, et al.
Veröffentlicht: (2026)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
Explainability-Guided Adversarial Attacks on Transformer-Based Malware Detectors Using Control Flow Graphs
von: Wheeler, Andrew, et al.
Veröffentlicht: (2026)
von: Wheeler, Andrew, et al.
Veröffentlicht: (2026)
Towards Knowledge Guided Pretraining Approaches for Multimodal Foundation Models: Applications in Remote Sensing
von: Ravirathinam, Praveen, et al.
Veröffentlicht: (2024)
von: Ravirathinam, Praveen, et al.
Veröffentlicht: (2024)
Divergence of Empirical Neural Tangent Kernel in Classification Problems
von: Yu, Zixiong, et al.
Veröffentlicht: (2025)
von: Yu, Zixiong, et al.
Veröffentlicht: (2025)
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
von: Yang, Haotian, et al.
Veröffentlicht: (2025)
von: Yang, Haotian, et al.
Veröffentlicht: (2025)
Missing Data Imputation by Reducing Mutual Information with Rectified Flows
von: Yu, Jiahao, et al.
Veröffentlicht: (2025)
von: Yu, Jiahao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
von: Frei, Spencer, et al.
Veröffentlicht: (2022) -
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
von: Roberts, Nicholas, et al.
Veröffentlicht: (2025) -
BTS: Harmonizing Specialized Experts into a Generalist LLM
von: Zhang, Qizhen, et al.
Veröffentlicht: (2025) -
Noisy Zero-Shot Coordination: Breaking The Common Knowledge Assumption In Zero-Shot Coordination Games
von: Anwar, Usman, et al.
Veröffentlicht: (2024) -
Analysing the Sample Complexity of Opponent Shaping
von: Fung, Kitty, et al.
Veröffentlicht: (2024)