LAD: Learning Advantage Distribution for Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Wendi, Li, Sharon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
von: Li, Wendi, et al.
Veröffentlicht: (2026)
von: Li, Wendi, et al.
Veröffentlicht: (2026)
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025)
von: Li, Wendi, et al.
Veröffentlicht: (2025)
Nonconvex Penalized LAD Estimation in Partial Linear Models with DNNs: Asymptotic Analysis and Proximal Algorithms
von: Feng, Lechen, et al.
Veröffentlicht: (2025)
von: Feng, Lechen, et al.
Veröffentlicht: (2025)
FedLAD: A Linear Algebra Based Data Poisoning Defence for Federated Learning
von: Xiong, Qi, et al.
Veröffentlicht: (2025)
von: Xiong, Qi, et al.
Veröffentlicht: (2025)
FedLAD: A Modular and Adaptive Testbed for Federated Log Anomaly Detection
von: Liao, Yihan, et al.
Veröffentlicht: (2025)
von: Liao, Yihan, et al.
Veröffentlicht: (2025)
ADORA: Training Reasoning Models with Dynamic Advantage Estimation on Reinforcement Learning
von: Ren, Qingnan, et al.
Veröffentlicht: (2026)
von: Ren, Qingnan, et al.
Veröffentlicht: (2026)
Exponential Quantum Communication Advantage in Distributed Inference and Learning
von: Gilboa, Dar, et al.
Veröffentlicht: (2023)
von: Gilboa, Dar, et al.
Veröffentlicht: (2023)
Provable Privacy Advantages of Decentralized Federated Learning via Distributed Optimization
von: Yu, Wenrui, et al.
Veröffentlicht: (2024)
von: Yu, Wenrui, et al.
Veröffentlicht: (2024)
Amortized Network Intervention to Steer the Excitatory Point Processes
von: Song, Zitao, et al.
Veröffentlicht: (2023)
von: Song, Zitao, et al.
Veröffentlicht: (2023)
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
von: Wiltzer, Harley, et al.
Veröffentlicht: (2024)
von: Wiltzer, Harley, et al.
Veröffentlicht: (2024)
ECoLAD: Deployment-Oriented Evaluation for Automotive Time-Series Anomaly Detection
von: Özer, Kadir-Kaan, et al.
Veröffentlicht: (2026)
von: Özer, Kadir-Kaan, et al.
Veröffentlicht: (2026)
LAD-BNet: Lag-Aware Dual-Branch Networks for Real-Time Energy Forecasting on Edge Devices
von: Lignier, Jean-Philippe
Veröffentlicht: (2025)
von: Lignier, Jean-Philippe
Veröffentlicht: (2025)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
von: Li, Ziheng, et al.
Veröffentlicht: (2026)
von: Li, Ziheng, et al.
Veröffentlicht: (2026)
Generalized Advantage Estimation for Distributional Policy Gradients
von: Shaik, Shahil, et al.
Veröffentlicht: (2025)
von: Shaik, Shahil, et al.
Veröffentlicht: (2025)
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning
von: Gong, Shijin, et al.
Veröffentlicht: (2026)
von: Gong, Shijin, et al.
Veröffentlicht: (2026)
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
von: Im, Shawn, et al.
Veröffentlicht: (2024)
von: Im, Shawn, et al.
Veröffentlicht: (2024)
Path Learning with Trajectory Advantage Regression
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
von: Miyaguchi, Kohei
Veröffentlicht: (2025)
Prospects of Privacy Advantage in Quantum Machine Learning
von: Heredge, Jamie, et al.
Veröffentlicht: (2024)
von: Heredge, Jamie, et al.
Veröffentlicht: (2024)
Stabilizing Efficient Reasoning with Step-Level Advantage Selection
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
von: Brantley, Kianté, et al.
Veröffentlicht: (2025)
von: Brantley, Kianté, et al.
Veröffentlicht: (2025)
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
von: Xiong, Jian, et al.
Veröffentlicht: (2025)
von: Xiong, Jian, et al.
Veröffentlicht: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
Advantage-based Temporal Attack in Reinforcement Learning
von: He, Shenghong
Veröffentlicht: (2026)
von: He, Shenghong
Veröffentlicht: (2026)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
von: Ghosh, Anurag, et al.
Veröffentlicht: (2026)
von: Ghosh, Anurag, et al.
Veröffentlicht: (2026)
Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
von: Chen, Xinzhu, et al.
Veröffentlicht: (2025)
von: Chen, Xinzhu, et al.
Veröffentlicht: (2025)
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
von: Gong, Shijin, et al.
Veröffentlicht: (2026)
von: Gong, Shijin, et al.
Veröffentlicht: (2026)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
von: Shen, Si, et al.
Veröffentlicht: (2025)
von: Shen, Si, et al.
Veröffentlicht: (2025)
Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models
von: Li, Muxing, et al.
Veröffentlicht: (2025)
von: Li, Muxing, et al.
Veröffentlicht: (2025)
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
von: Peng, Keqin, et al.
Veröffentlicht: (2026)
von: Peng, Keqin, et al.
Veröffentlicht: (2026)
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
von: Huang, Yiming, et al.
Veröffentlicht: (2026)
Your Group-Relative Advantage Is Biased
von: Yang, Fengkai, et al.
Veröffentlicht: (2026)
von: Yang, Fengkai, et al.
Veröffentlicht: (2026)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
von: Liu, Tenglong, et al.
Veröffentlicht: (2024)
von: Liu, Tenglong, et al.
Veröffentlicht: (2024)
Sampling Complexity of TD and PPO in RKHS
von: Zou, Lu, et al.
Veröffentlicht: (2025)
von: Zou, Lu, et al.
Veröffentlicht: (2025)
Advantage Alignment Algorithms
von: Duque, Juan Agustin, et al.
Veröffentlicht: (2024)
von: Duque, Juan Agustin, et al.
Veröffentlicht: (2024)
Active Advantage-Aligned Online Reinforcement Learning with Offline Data
von: Liu, Xuefeng, et al.
Veröffentlicht: (2025)
von: Liu, Xuefeng, et al.
Veröffentlicht: (2025)
Mirror Descent Actor Critic via Bounded Advantage Learning
von: Iwaki, Ryo
Veröffentlicht: (2025)
von: Iwaki, Ryo
Veröffentlicht: (2025)
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
von: Xue, Shuchen, et al.
Veröffentlicht: (2025)
von: Xue, Shuchen, et al.
Veröffentlicht: (2025)
How Well Can Preference Optimization Generalize Under Noisy Feedback?
von: Im, Shawn, et al.
Veröffentlicht: (2025)
von: Im, Shawn, et al.
Veröffentlicht: (2025)
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
von: Li, Junbo, et al.
Veröffentlicht: (2025)
von: Li, Junbo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
von: Li, Wendi, et al.
Veröffentlicht: (2026) -
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025) -
Nonconvex Penalized LAD Estimation in Partial Linear Models with DNNs: Asymptotic Analysis and Proximal Algorithms
von: Feng, Lechen, et al.
Veröffentlicht: (2025) -
FedLAD: A Linear Algebra Based Data Poisoning Defence for Federated Learning
von: Xiong, Qi, et al.
Veröffentlicht: (2025) -
FedLAD: A Modular and Adaptive Testbed for Federated Log Anomaly Detection
von: Liao, Yihan, et al.
Veröffentlicht: (2025)