Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Feng, Yunzhen, Jain, Parag, Hartshorn, Anthony, Duan, Yaqi, Kempe, Julia |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
par: Feng, Yunzhen, et autres
Publié: (2025)
par: Feng, Yunzhen, et autres
Publié: (2025)
PILAF: Optimal Human Preference Sampling for Reward Modeling
par: Feng, Yunzhen, et autres
Publié: (2025)
par: Feng, Yunzhen, et autres
Publié: (2025)
Model Collapse Demystified: The Case of Regression
par: Dohmatob, Elvis, et autres
Publié: (2024)
par: Dohmatob, Elvis, et autres
Publié: (2024)
Strong Model Collapse
par: Dohmatob, Elvis, et autres
Publié: (2024)
par: Dohmatob, Elvis, et autres
Publié: (2024)
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
par: Feng, Yunzhen, et autres
Publié: (2024)
par: Feng, Yunzhen, et autres
Publié: (2024)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
par: Dohmatob, Elvis, et autres
Publié: (2024)
par: Dohmatob, Elvis, et autres
Publié: (2024)
Attacking Bayes: On the Adversarial Robustness of Bayesian Neural Networks
par: Feng, Yunzhen, et autres
Publié: (2024)
par: Feng, Yunzhen, et autres
Publié: (2024)
Don't Waste It: Guiding Generative Recommenders with Structured Human Priors via Multi-Head Decoding
par: Zhang, Yunkai, et autres
Publié: (2025)
par: Zhang, Yunkai, et autres
Publié: (2025)
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
par: Guo, Junyu, et autres
Publié: (2025)
par: Guo, Junyu, et autres
Publié: (2025)
Don't Waste Your Time: Early Stopping Cross-Validation
par: Bergman, Edward, et autres
Publié: (2024)
par: Bergman, Edward, et autres
Publié: (2024)
Efficient RL Training for LLMs with Experience Replay
par: Arnal, Charles, et autres
Publié: (2026)
par: Arnal, Charles, et autres
Publié: (2026)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
par: Shahnazari, Kourosh, et autres
Publié: (2025)
par: Shahnazari, Kourosh, et autres
Publié: (2025)
Don't Be So Positive: Negative Step Sizes in Second-Order Methods
par: Shea, Betty, et autres
Publié: (2024)
par: Shea, Betty, et autres
Publié: (2024)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
par: Pal, Arka, et autres
Publié: (2025)
par: Pal, Arka, et autres
Publié: (2025)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
par: Johnson, Daniel D., et autres
Publié: (2024)
par: Johnson, Daniel D., et autres
Publié: (2024)
Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective
par: Yao, Yunzhen, et autres
Publié: (2025)
par: Yao, Yunzhen, et autres
Publié: (2025)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
par: Wang, Qinsi, et autres
Publié: (2025)
par: Wang, Qinsi, et autres
Publié: (2025)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
par: Sokar, Ghada, et autres
Publié: (2024)
par: Sokar, Ghada, et autres
Publié: (2024)
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
par: Han, Boyu, et autres
Publié: (2025)
par: Han, Boyu, et autres
Publié: (2025)
Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
par: Ravie, Navin Sriram, et autres
Publié: (2026)
par: Ravie, Navin Sriram, et autres
Publié: (2026)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
par: Hernandez, Adriano
Publié: (2024)
par: Hernandez, Adriano
Publié: (2024)
CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning
par: Sun, Ke, et autres
Publié: (2026)
par: Sun, Ke, et autres
Publié: (2026)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
par: Zhang, Jiefu, et autres
Publié: (2026)
par: Zhang, Jiefu, et autres
Publié: (2026)
Emergent properties with repeated examples
par: Charton, François, et autres
Publié: (2024)
par: Charton, François, et autres
Publié: (2024)
When Models Don't Collapse: On the Consistency of Iterative MLE
par: Barzilai, Daniel, et autres
Publié: (2025)
par: Barzilai, Daniel, et autres
Publié: (2025)
Position: Don't be Afraid of Over-Smoothing And Over-Squashing
par: Kormann, Niklas, et autres
Publié: (2026)
par: Kormann, Niklas, et autres
Publié: (2026)
Grow, Don't Overwrite: Fine-tuning Without Forgetting
par: Adila, Dyah, et autres
Publié: (2026)
par: Adila, Dyah, et autres
Publié: (2026)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
par: Roux, Christophe, et autres
Publié: (2025)
par: Roux, Christophe, et autres
Publié: (2025)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
par: Roy, Debjyoti Saha, et autres
Publié: (2024)
par: Roy, Debjyoti Saha, et autres
Publié: (2024)
Don't Stop Me Yet: Sampling Loss Minima via Dissipative Riemannian Mechanics
par: Jacobsen, Albert Kjøller, et autres
Publié: (2026)
par: Jacobsen, Albert Kjøller, et autres
Publié: (2026)
Class Confidence Aware Reweighting for Long Tailed Learning
par: Jagati, Brainard Philemon, et autres
Publié: (2026)
par: Jagati, Brainard Philemon, et autres
Publié: (2026)
We Still Don't Understand High-Dimensional Bayesian Optimization
par: Doumont, Colin, et autres
Publié: (2025)
par: Doumont, Colin, et autres
Publié: (2025)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
par: Shahout, Rana, et autres
Publié: (2024)
par: Shahout, Rana, et autres
Publié: (2024)
Don't Forget Imagination!
par: Vityaev, Evgenii E., et autres
Publié: (2025)
par: Vityaev, Evgenii E., et autres
Publié: (2025)
Transformers Don't In-Context Learn Least Squares Regression
par: Hill, Joshua, et autres
Publié: (2025)
par: Hill, Joshua, et autres
Publié: (2025)
Don't Walk the Line: Boundary Guidance for Filtered Generation
par: Ball, Sarah, et autres
Publié: (2025)
par: Ball, Sarah, et autres
Publié: (2025)
Don't Explain Noise: Robust Counterfactuals for Randomized Ensembles
par: Forel, Alexandre, et autres
Publié: (2022)
par: Forel, Alexandre, et autres
Publié: (2022)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
par: Cassidy, Oliver, et autres
Publié: (2024)
par: Cassidy, Oliver, et autres
Publié: (2024)
Position: Measure Dataset Diversity, Don't Just Claim It
par: Zhao, Dora, et autres
Publié: (2024)
par: Zhao, Dora, et autres
Publié: (2024)
Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning
par: Yin, Bo, et autres
Publié: (2025)
par: Yin, Bo, et autres
Publié: (2025)
Documents similaires
-
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
par: Feng, Yunzhen, et autres
Publié: (2025) -
PILAF: Optimal Human Preference Sampling for Reward Modeling
par: Feng, Yunzhen, et autres
Publié: (2025) -
Model Collapse Demystified: The Case of Regression
par: Dohmatob, Elvis, et autres
Publié: (2024) -
Strong Model Collapse
par: Dohmatob, Elvis, et autres
Publié: (2024) -
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification
par: Feng, Yunzhen, et autres
Publié: (2024)