RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Ren, Tao, Jiang, Jinyang, Yang, Hui, Tian, Wan, Zou, Minhao, Li, Guanghao, Zhang, Zishi, Wang, Qinghao, Qin, Shentao, Zhao, Yanjun, Tao, Rui, Shao, Hui, Peng, Yijie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Nonparametric Bayesian Optimization for General Rewards
por: Zhang, Zishi, et al.
Publicado: (2026)
por: Zhang, Zishi, et al.
Publicado: (2026)
RiskMiner: Discovering Formulaic Alphas via Risk Seeking Monte Carlo Tree Search
por: Ren, Tao, et al.
Publicado: (2024)
por: Ren, Tao, et al.
Publicado: (2024)
FLOPS: Forward Learning with OPtimal Sampling
por: Ren, Tao, et al.
Publicado: (2024)
por: Ren, Tao, et al.
Publicado: (2024)
Optimal low-rank stochastic gradient estimation for LLM training
por: Li, Zehao, et al.
Publicado: (2026)
por: Li, Zehao, et al.
Publicado: (2026)
Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
por: Ren, Tao, et al.
Publicado: (2025)
por: Ren, Tao, et al.
Publicado: (2025)
Omni-Masked Gradient Descent: Memory-Efficient Optimization via Mask Traversal with Improved Convergence
por: Yang, Hui, et al.
Publicado: (2026)
por: Yang, Hui, et al.
Publicado: (2026)
Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection
por: Zhang, Zishi, et al.
Publicado: (2024)
por: Zhang, Zishi, et al.
Publicado: (2024)
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
por: Xu, Ran, et al.
Publicado: (2026)
por: Xu, Ran, et al.
Publicado: (2026)
Risk-Controlled Post-Processing of Decision Policies
por: Joshi, Sunay, et al.
Publicado: (2026)
por: Joshi, Sunay, et al.
Publicado: (2026)
LLM-Inspired Pretrain-Then-Finetune for Small-Data, Large-Scale Optimization
por: Zhang, Zishi, et al.
Publicado: (2026)
por: Zhang, Zishi, et al.
Publicado: (2026)
Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards
por: Normann, Philipp, et al.
Publicado: (2026)
por: Normann, Philipp, et al.
Publicado: (2026)
Deep Reinforcement Learning for Solving Management Problems: Towards A Large Management Mode
por: Jiang, Jinyang, et al.
Publicado: (2024)
por: Jiang, Jinyang, et al.
Publicado: (2024)
Stochastic Approximation Methods for Distortion Risk Measure Optimization
por: Jiang, Jinyang, et al.
Publicado: (2025)
por: Jiang, Jinyang, et al.
Publicado: (2025)
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
por: Cao, Yuan, et al.
Publicado: (2026)
por: Cao, Yuan, et al.
Publicado: (2026)
Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
por: Dupuis, Nicolas, et al.
Publicado: (2025)
por: Dupuis, Nicolas, et al.
Publicado: (2025)
Task-Centric Policy Optimization from Misaligned Motion Priors
por: Zheng, Ziang, et al.
Publicado: (2026)
por: Zheng, Ziang, et al.
Publicado: (2026)
A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
por: Yang, Shentao, et al.
Publicado: (2024)
por: Yang, Shentao, et al.
Publicado: (2024)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
por: Liu, Yuexiao, et al.
Publicado: (2025)
por: Liu, Yuexiao, et al.
Publicado: (2025)
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
por: Ghimire, Mukesh, et al.
Publicado: (2026)
por: Ghimire, Mukesh, et al.
Publicado: (2026)
Consolidating Rewarded Perturbations for LLM Post-Training
por: Zhang, Zheyu, et al.
Publicado: (2026)
por: Zhang, Zheyu, et al.
Publicado: (2026)
AlphaPO: Reward Shape Matters for LLM Alignment
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
por: Le, Minh-Quan, et al.
Publicado: (2025)
por: Le, Minh-Quan, et al.
Publicado: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
por: Pan, Chengjun, et al.
Publicado: (2026)
por: Pan, Chengjun, et al.
Publicado: (2026)
Copper Current Collector: The Cornerstones of Practical Lithium Metal and Anode–Free Batteries
por: Jinyang Zhou, et al.
Publicado: (2024)
por: Jinyang Zhou, et al.
Publicado: (2024)
Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
por: Li, Zhongyi, et al.
Publicado: (2026)
por: Li, Zhongyi, et al.
Publicado: (2026)
Convergence of Neural Network Policies for Risk--Reward Optimization
por: Chen, Chang, et al.
Publicado: (2026)
por: Chen, Chang, et al.
Publicado: (2026)
Navigating the Terrain of Post‐Polypectomy Surveillance: Charting the Complex Landscape of Guidelines and Recurrence Risks
por: Xiaojin Tao, et al.
Publicado: (2026)
por: Xiaojin Tao, et al.
Publicado: (2026)
Hypertensive Uveitis After Intravitreal Faricimab: Clinical and Methodological Considerations
por: Guanghao Qin
Publicado: (2026)
por: Guanghao Qin
Publicado: (2026)
Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
por: Ren, Mengjie, et al.
Publicado: (2026)
por: Ren, Mengjie, et al.
Publicado: (2026)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
por: Liu, Yixin, et al.
Publicado: (2026)
por: Liu, Yixin, et al.
Publicado: (2026)
Cold-Start Forecasting of New Product Life-Cycles via Conditional Diffusion Models
por: Zhou, Ruihan, et al.
Publicado: (2026)
por: Zhou, Ruihan, et al.
Publicado: (2026)
Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation
por: Zou, Qingyun, et al.
Publicado: (2026)
por: Zou, Qingyun, et al.
Publicado: (2026)
Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
por: Ren, Jie, et al.
Publicado: (2025)
por: Ren, Jie, et al.
Publicado: (2025)
Risk Capacity and Optimal Monetary Policy
por: Sun, Rui
Publicado: (2026)
por: Sun, Rui
Publicado: (2026)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
por: Jiang, Yuxin, et al.
Publicado: (2026)
por: Jiang, Yuxin, et al.
Publicado: (2026)
Comment on “Research on the Integrated Training Mode of Higher Art Education for the Deaf”
por: Hui Shao
Publicado: (2023)
por: Hui Shao
Publicado: (2023)
Navigating Ecological Risk: Firm‐Level Innovation Responses to Biodiversity Risks in the United States
por: Miaomiao Tao
Publicado: (2026)
por: Miaomiao Tao
Publicado: (2026)
A Motif-Based Framework for Decomposing Risk Spillovers
por: Shao, Ying-Hui, et al.
Publicado: (2026)
por: Shao, Ying-Hui, et al.
Publicado: (2026)
"Think First, Verify Always": Training Humans to Face AI Risks
por: Aydin, Yuksel
Publicado: (2025)
por: Aydin, Yuksel
Publicado: (2025)
Ejemplares similares
-
Nonparametric Bayesian Optimization for General Rewards
por: Zhang, Zishi, et al.
Publicado: (2026) -
RiskMiner: Discovering Formulaic Alphas via Risk Seeking Monte Carlo Tree Search
por: Ren, Tao, et al.
Publicado: (2024) -
FLOPS: Forward Learning with OPtimal Sampling
por: Ren, Tao, et al.
Publicado: (2024) -
Optimal low-rank stochastic gradient estimation for LLM training
por: Li, Zehao, et al.
Publicado: (2026) -
Half-order Fine-Tuning for Diffusion Model: A Recursive Likelihood Ratio Optimizer
por: Ren, Tao, et al.
Publicado: (2025)