Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Shijin, Ye, Kai, Zhu, Jin, Zhang, Xinyu, Zhou, Hongyi, Shi, Chengchun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
by: Gong, Shijin, et al.
Published: (2026)
by: Gong, Shijin, et al.
Published: (2026)
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
Detecting LLM-Generated Text with Performance Guarantees
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
by: Zhou, Hongyi, et al.
Published: (2025)
by: Zhou, Hongyi, et al.
Published: (2025)
Statistical Inference in Reinforcement Learning: A Selective Survey
by: Shi, Chengchun
Published: (2025)
by: Shi, Chengchun
Published: (2025)
Doubly Robust Alignment for Large Language Models
by: Xu, Erhan, et al.
Published: (2025)
by: Xu, Erhan, et al.
Published: (2025)
PEARL: Performance-Enhanced Aggregated Representation Learning
by: Li, Wenhui, et al.
Published: (2025)
by: Li, Wenhui, et al.
Published: (2025)
Balancing Interference and Correlation in Spatial Experimental Designs: A Causal Graph Cut Approach
by: Zhu, Jin, et al.
Published: (2025)
by: Zhu, Jin, et al.
Published: (2025)
Deep Generative Demand Learning for Newsvendor and Pricing
by: Gong, Shijin, et al.
Published: (2024)
by: Gong, Shijin, et al.
Published: (2024)
Reinforcement Learning from Human Feedback: A Statistical Perspective
by: Liu, Pangpang, et al.
Published: (2026)
by: Liu, Pangpang, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
From Authors to Reviewers: Leveraging Rankings to Improve Peer Review
by: Wang, Weichen, et al.
Published: (2025)
by: Wang, Weichen, et al.
Published: (2025)
Learning Perturbations to Extrapolate Your LLM
by: Cen, Zetai, et al.
Published: (2026)
by: Cen, Zetai, et al.
Published: (2026)
Distributed Clustering based on Distributional Kernel
by: Zhang, Hang, et al.
Published: (2024)
by: Zhang, Hang, et al.
Published: (2024)
From Weighting to Modeling: A Nonparametric Estimator for Off-Policy Evaluation
by: Zhu, Rong J. B.
Published: (2026)
by: Zhu, Rong J. B.
Published: (2026)
Perturbation is All You Need for Extrapolating Language Models
by: Cen, Zetai, et al.
Published: (2026)
by: Cen, Zetai, et al.
Published: (2026)
Nonparametric Instrumental Regression via Kernel Methods is Minimax Optimal
by: Meunier, Dimitri, et al.
Published: (2024)
by: Meunier, Dimitri, et al.
Published: (2024)
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
by: Gong, Xue, et al.
Published: (2026)
by: Gong, Xue, et al.
Published: (2026)
Nonparametric Kernel Clustering with Bandit Feedback
by: Thuot, Victor, et al.
Published: (2026)
by: Thuot, Victor, et al.
Published: (2026)
Active Learning with Neural Networks: Insights from Nonparametric Statistics
by: Zhu, Yinglun, et al.
Published: (2022)
by: Zhu, Yinglun, et al.
Published: (2022)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
by: Shen, Si, et al.
Published: (2025)
by: Shen, Si, et al.
Published: (2025)
In Search of Quantum Advantage: Estimating the Number of Shots in Quantum Kernel Methods
by: Miroszewski, Artur, et al.
Published: (2024)
by: Miroszewski, Artur, et al.
Published: (2024)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
by: Brantley, Kianté, et al.
Published: (2025)
by: Brantley, Kianté, et al.
Published: (2025)
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
by: Xiong, Jian, et al.
Published: (2025)
by: Xiong, Jian, et al.
Published: (2025)
SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning
by: Ai, Zhengyang, et al.
Published: (2026)
by: Ai, Zhengyang, et al.
Published: (2026)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
by: Zhu, Jin, et al.
Published: (2023)
by: Zhu, Jin, et al.
Published: (2023)
Counterfactually Safe Reinforcement Learning
by: Li, Jingyi, et al.
Published: (2026)
by: Li, Jingyi, et al.
Published: (2026)
Dual Active Learning for Reinforcement Learning from Human Feedback
by: Liu, Pangpang, et al.
Published: (2024)
by: Liu, Pangpang, et al.
Published: (2024)
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
The Impact of Isolation Kernel on Agglomerative Hierarchical Clustering Algorithms
by: Han, Xin, et al.
Published: (2020)
by: Han, Xin, et al.
Published: (2020)
Statistical Optimality of Divide and Conquer Kernel-based Functional Linear Regression
by: Liu, Jiading, et al.
Published: (2022)
by: Liu, Jiading, et al.
Published: (2022)
Unraveling the Interplay between Carryover Effects and Reward Autocorrelations in Switchback Experiments
by: Wen, Qianglin, et al.
Published: (2024)
by: Wen, Qianglin, et al.
Published: (2024)
ADORA: Training Reasoning Models with Dynamic Advantage Estimation on Reinforcement Learning
by: Ren, Qingnan, et al.
Published: (2026)
by: Ren, Qingnan, et al.
Published: (2026)
An Online Automatic Modulation Classification Scheme Based on Isolation Distributional Kernel
by: Li, Xinpeng, et al.
Published: (2024)
by: Li, Xinpeng, et al.
Published: (2024)
ATFNet: Adaptive Time-Frequency Ensembled Network for Long-term Time Series Forecasting
by: Ye, Hengyu, et al.
Published: (2024)
by: Ye, Hengyu, et al.
Published: (2024)
Towards Trustworthy Web Attack Detection: An Uncertainty-Aware Ensemble Deep Kernel Learning Model
by: Zhou, Yonghang, et al.
Published: (2024)
by: Zhou, Yonghang, et al.
Published: (2024)
Similar Items
-
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
by: Gong, Shijin, et al.
Published: (2026) -
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
by: Zhou, Hongyi, et al.
Published: (2026) -
Detecting LLM-Generated Text with Performance Guarantees
by: Zhou, Hongyi, et al.
Published: (2026) -
Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text
by: Zhou, Hongyi, et al.
Published: (2026) -
Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning
by: Ye, Kai, et al.
Published: (2025)