Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
Fuente:
arXiv
Salvato in:
| Autori principali: | Yao, Wei, Yang, Wenkai, Wang, Ziqiao, Lin, Yankai, Liu, Yong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
di: Yao, Wei, et al.
Pubblicazione: (2025)
di: Yao, Wei, et al.
Pubblicazione: (2025)
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
di: Zhang, Yixian, et al.
Pubblicazione: (2025)
di: Zhang, Yixian, et al.
Pubblicazione: (2025)
On Weak-to-Strong Generalization and f-Divergence
di: Yao, Wei, et al.
Pubblicazione: (2025)
di: Yao, Wei, et al.
Pubblicazione: (2025)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
di: Xiong, Wei, et al.
Pubblicazione: (2023)
di: Xiong, Wei, et al.
Pubblicazione: (2023)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
di: Kong, Jason, et al.
Pubblicazione: (2026)
di: Kong, Jason, et al.
Pubblicazione: (2026)
On Flow Matching KL Divergence
di: Su, Maojiang, et al.
Pubblicazione: (2025)
di: Su, Maojiang, et al.
Pubblicazione: (2025)
KL for a KL: On-Policy Distillation with Control Variate Baseline
di: Oh, Minjae, et al.
Pubblicazione: (2026)
di: Oh, Minjae, et al.
Pubblicazione: (2026)
On the Emergence of Weak-to-Strong Generalization: A Bias-Variance Perspective
di: Xu, Gengze, et al.
Pubblicazione: (2025)
di: Xu, Gengze, et al.
Pubblicazione: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
di: Sun, Ruotong, et al.
Pubblicazione: (2026)
di: Sun, Ruotong, et al.
Pubblicazione: (2026)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
di: Zhu, Lingwei, et al.
Pubblicazione: (2023)
di: Zhu, Lingwei, et al.
Pubblicazione: (2023)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
di: Shu, Yao, et al.
Pubblicazione: (2026)
di: Shu, Yao, et al.
Pubblicazione: (2026)
Optimal Stability of KL Divergence under Gaussian Perturbations
di: Pan, Jialu, et al.
Pubblicazione: (2026)
di: Pan, Jialu, et al.
Pubblicazione: (2026)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
di: Liu, Kezhao, et al.
Pubblicazione: (2025)
di: Liu, Kezhao, et al.
Pubblicazione: (2025)
DeepCritic: Deliberate Critique with Large Language Models
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
KL-regularization Itself is Differentially Private in Bandits and RLHF
di: Zhang, Yizhou, et al.
Pubblicazione: (2025)
di: Zhang, Yizhou, et al.
Pubblicazione: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
di: Zhao, Qingyue, et al.
Pubblicazione: (2026)
di: Zhao, Qingyue, et al.
Pubblicazione: (2026)
KL Penalty Control via Perturbation for Direct Preference Optimization
di: Lee, Sangkyu, et al.
Pubblicazione: (2025)
di: Lee, Sangkyu, et al.
Pubblicazione: (2025)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
di: Shah, Vedant, et al.
Pubblicazione: (2025)
di: Shah, Vedant, et al.
Pubblicazione: (2025)
Offline and Online KL-Regularized RLHF under Differential Privacy
di: Wu, Yulian, et al.
Pubblicazione: (2025)
di: Wu, Yulian, et al.
Pubblicazione: (2025)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
di: Tang, Kenton, et al.
Pubblicazione: (2026)
di: Tang, Kenton, et al.
Pubblicazione: (2026)
Quantifying the Gain in Weak-to-Strong Generalization
di: Charikar, Moses, et al.
Pubblicazione: (2024)
di: Charikar, Moses, et al.
Pubblicazione: (2024)
On the Blessing of Pre-training in Weak-to-Strong Generalization
di: Yao, Wei, et al.
Pubblicazione: (2026)
di: Yao, Wei, et al.
Pubblicazione: (2026)
A KL-regularization Framework for Learning to Plan with Adaptive Priors
di: Serra-Gomez, Álvaro, et al.
Pubblicazione: (2025)
di: Serra-Gomez, Álvaro, et al.
Pubblicazione: (2025)
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
di: Yuan, Rui, et al.
Pubblicazione: (2026)
di: Yuan, Rui, et al.
Pubblicazione: (2026)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
di: Yu, Xin, et al.
Pubblicazione: (2026)
di: Yu, Xin, et al.
Pubblicazione: (2026)
Weak-to-Strong Generalization under Distribution Shifts
di: Jeon, Myeongho, et al.
Pubblicazione: (2025)
di: Jeon, Myeongho, et al.
Pubblicazione: (2025)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
di: Pawelczyk, Martin, et al.
Pubblicazione: (2024)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2024)
Weak-to-Strong Generalization Through the Data-Centric Lens
di: Shin, Changho, et al.
Pubblicazione: (2024)
di: Shin, Changho, et al.
Pubblicazione: (2024)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
di: Zhang, Lunjun, et al.
Pubblicazione: (2026)
di: Zhang, Lunjun, et al.
Pubblicazione: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One
di: Song, Yiwen, et al.
Pubblicazione: (2025)
di: Song, Yiwen, et al.
Pubblicazione: (2025)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
di: Xue, Yihao, et al.
Pubblicazione: (2025)
di: Xue, Yihao, et al.
Pubblicazione: (2025)
Mixture of Weak & Strong Experts on Graphs
di: Zeng, Hanqing, et al.
Pubblicazione: (2023)
di: Zeng, Hanqing, et al.
Pubblicazione: (2023)
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
di: Yao, Yihang, et al.
Pubblicazione: (2026)
di: Yao, Yihang, et al.
Pubblicazione: (2026)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
di: Deng, Wei
Pubblicazione: (2026)
di: Deng, Wei
Pubblicazione: (2026)
Documenti analoghi
-
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
di: Yao, Wei, et al.
Pubblicazione: (2025) -
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning
di: Zhang, Yixian, et al.
Pubblicazione: (2025) -
On Weak-to-Strong Generalization and f-Divergence
di: Yao, Wei, et al.
Pubblicazione: (2025) -
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
di: Xiong, Wei, et al.
Pubblicazione: (2023) -
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
di: Yang, Wenkai, et al.
Pubblicazione: (2024)