A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Yu, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
von: Chen, Fan, et al.
Veröffentlicht: (2022)
von: Chen, Fan, et al.
Veröffentlicht: (2022)
Training Implicit Generative Models via an Invariant Statistical Loss
von: de Frutos, José Manuel, et al.
Veröffentlicht: (2024)
von: de Frutos, José Manuel, et al.
Veröffentlicht: (2024)
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
von: Yu, Hao
Veröffentlicht: (2025)
von: Yu, Hao
Veröffentlicht: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
von: Dong, Zihan, et al.
Veröffentlicht: (2026)
von: Dong, Zihan, et al.
Veröffentlicht: (2026)
Online Learning with Unknown Constraints
von: Sridharan, Karthik, et al.
Veröffentlicht: (2024)
von: Sridharan, Karthik, et al.
Veröffentlicht: (2024)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
von: Zhan, Wenhao, et al.
Veröffentlicht: (2023)
von: Zhan, Wenhao, et al.
Veröffentlicht: (2023)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024)
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024)
A note on the impossibility of conditional PAC-efficient reasoning in large language models
von: Zeng, Hao
Veröffentlicht: (2025)
von: Zeng, Hao
Veröffentlicht: (2025)
RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
von: Hao, Sai, et al.
Veröffentlicht: (2026)
von: Hao, Sai, et al.
Veröffentlicht: (2026)
Compression, Generalization and Learning
von: Campi, Marco C., et al.
Veröffentlicht: (2023)
von: Campi, Marco C., et al.
Veröffentlicht: (2023)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
von: Hazard, Christopher J., et al.
Veröffentlicht: (2025)
von: Hazard, Christopher J., et al.
Veröffentlicht: (2025)
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
von: Kothapalli, Vignesh, et al.
Veröffentlicht: (2024)
On the Statistical Capacity of Deep Generative Models
von: Tam, Edric, et al.
Veröffentlicht: (2025)
von: Tam, Edric, et al.
Veröffentlicht: (2025)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
von: Mayaki, Mansour Zoubeirou a
Veröffentlicht: (2026)
von: Mayaki, Mansour Zoubeirou a
Veröffentlicht: (2026)
Counterfactual Generative Modeling with Variational Causal Inference
von: Wu, Yulun, et al.
Veröffentlicht: (2024)
von: Wu, Yulun, et al.
Veröffentlicht: (2024)
Neural Networks Generalize on Low Complexity Data
von: Chatterjee, Sourav, et al.
Veröffentlicht: (2024)
von: Chatterjee, Sourav, et al.
Veröffentlicht: (2024)
FraPPE: Fast and Efficient Preference-based Pure Exploration
von: Das, Udvas, et al.
Veröffentlicht: (2025)
von: Das, Udvas, et al.
Veröffentlicht: (2025)
On the Provable Performance Guarantee of Efficient Reasoning Models
von: Zeng, Hao, et al.
Veröffentlicht: (2025)
von: Zeng, Hao, et al.
Veröffentlicht: (2025)
On the Statistical Properties of Generative Adversarial Models for Low Intrinsic Data Dimension
von: Chakraborty, Saptarshi, et al.
Veröffentlicht: (2024)
von: Chakraborty, Saptarshi, et al.
Veröffentlicht: (2024)
Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data
von: Chakraborty, Saptarshi, et al.
Veröffentlicht: (2026)
von: Chakraborty, Saptarshi, et al.
Veröffentlicht: (2026)
U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
von: Mei, Song
Veröffentlicht: (2024)
von: Mei, Song
Veröffentlicht: (2024)
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
von: Wang, Kevin, et al.
Veröffentlicht: (2026)
von: Wang, Kevin, et al.
Veröffentlicht: (2026)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
von: Kitaoka, Akira
Veröffentlicht: (2025)
von: Kitaoka, Akira
Veröffentlicht: (2025)
Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability
von: Yu, Lijia, et al.
Veröffentlicht: (2025)
von: Yu, Lijia, et al.
Veröffentlicht: (2025)
A Score-Based Density Formula, with Applications in Diffusion Generative Models
von: Li, Gen, et al.
Veröffentlicht: (2024)
von: Li, Gen, et al.
Veröffentlicht: (2024)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
von: Hellström, Fredrik, et al.
Veröffentlicht: (2023)
Foundations of Top-$k$ Decoding For Language Models
von: Noarov, Georgy, et al.
Veröffentlicht: (2025)
von: Noarov, Georgy, et al.
Veröffentlicht: (2025)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
A Likelihood Based Approach to Distribution Regression Using Conditional Deep Generative Models
von: Kumar, Shivam, et al.
Veröffentlicht: (2024)
von: Kumar, Shivam, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
von: Guo, Yang, et al.
Veröffentlicht: (2025)
von: Guo, Yang, et al.
Veröffentlicht: (2025)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
Statistical inference with belief functions: A survey
von: Cuzzolin, Fabio
Veröffentlicht: (2026)
von: Cuzzolin, Fabio
Veröffentlicht: (2026)
A Quantitative Characterization of Forgetting in Post-Training
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2026)
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2026)
A Fine-Grained Understanding of Uniform Convergence for Halfspaces
von: Kontorovich, Aryeh, et al.
Veröffentlicht: (2026)
von: Kontorovich, Aryeh, et al.
Veröffentlicht: (2026)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
von: Lattimore, Tor
Veröffentlicht: (2026)
von: Lattimore, Tor
Veröffentlicht: (2026)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
von: Zhang, Huiming, et al.
Veröffentlicht: (2026)
von: Zhang, Huiming, et al.
Veröffentlicht: (2026)
The Geometry of Benchmarks: A New Path Toward AGI
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
Deep Ensembles for Epistemic Uncertainty: A Frequentist Perspective
von: Jain, Anchit, et al.
Veröffentlicht: (2025)
von: Jain, Anchit, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
von: Chen, Fan, et al.
Veröffentlicht: (2022) -
Training Implicit Generative Models via an Invariant Statistical Loss
von: de Frutos, José Manuel, et al.
Veröffentlicht: (2024) -
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
von: Yu, Hao
Veröffentlicht: (2025) -
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026) -
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
von: Dong, Zihan, et al.
Veröffentlicht: (2026)