RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Kaichen, Gao, Shenghao, Hong, Yuzhong, Sun, Haipeng, Bao, Junwei, Jiang, Hongfei, Song, Yang, Dingqian, Hong, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
von: Hong, Yuzhong, et al.
Veröffentlicht: (2024)
von: Hong, Yuzhong, et al.
Veröffentlicht: (2024)
Multi-Turn Interactions for Text-to-SQL with Large Language Models
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
Efficient Prediction of Pass@k Scaling in Large Language Models
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2025)
Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
von: Yu, Yang
Veröffentlicht: (2025)
von: Yu, Yang
Veröffentlicht: (2025)
Elo-Evolve: A Co-evolutionary Framework for Language Model Alignment
von: Zhao, Jing, et al.
Veröffentlicht: (2026)
von: Zhao, Jing, et al.
Veröffentlicht: (2026)
Relationships Between the Maximum Principle and Dynamic Programming for Infinite Dimensional Non-Markovian Stochastic Control Systems
von: Gao, Dingqian, et al.
Veröffentlicht: (2025)
von: Gao, Dingqian, et al.
Veröffentlicht: (2025)
Dynamic Programming Principle for Stochastic Control Problems on Riemannian Manifolds
von: Gao, Dingqian, et al.
Veröffentlicht: (2025)
von: Gao, Dingqian, et al.
Veröffentlicht: (2025)
FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning
von: Pan, Haihui, et al.
Veröffentlicht: (2026)
von: Pan, Haihui, et al.
Veröffentlicht: (2026)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
von: Hariri, Mohsen, et al.
Veröffentlicht: (2025)
von: Hariri, Mohsen, et al.
Veröffentlicht: (2025)
Optimized Cost Per Click in Online Advertising: A Theoretical Analysis
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Characterization of the structure of $k$-edge-maximal graphs
von: Xia, Zheng-Jiang, et al.
Veröffentlicht: (2026)
von: Xia, Zheng-Jiang, et al.
Veröffentlicht: (2026)
BoRA: Bi-dimensional Weight-Decomposed Low-Rank Adaptation
von: Wang, Qiushi, et al.
Veröffentlicht: (2024)
von: Wang, Qiushi, et al.
Veröffentlicht: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
On Equivalence of Parameterized Inapproximability of k-Median, k-Max-Coverage, and 2-CSP
von: S., Karthik C., et al.
Veröffentlicht: (2024)
von: S., Karthik C., et al.
Veröffentlicht: (2024)
Interactive-KBQA: Multi-Turn Interactions for Knowledge Base Question Answering with Large Language Models
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
von: Xiong, Guanming, et al.
Veröffentlicht: (2024)
An Improved Greedy Approximation for (Metric) $k$-Means
von: Charikar, Moses, et al.
Veröffentlicht: (2026)
von: Charikar, Moses, et al.
Veröffentlicht: (2026)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
Top Pass: Improve Code Generation by Pass@k-Maximized Code Ranking
von: Lyu, Zhi-Cun, et al.
Veröffentlicht: (2024)
von: Lyu, Zhi-Cun, et al.
Veröffentlicht: (2024)
Discovering Top-k Structural Hole Spanners in Dynamic Networks
von: Goel, Diksha, et al.
Veröffentlicht: (2023)
von: Goel, Diksha, et al.
Veröffentlicht: (2023)
Leveraging LLM Inconsistency to Boost Pass@k Performance
von: Dalal, Uri, et al.
Veröffentlicht: (2025)
von: Dalal, Uri, et al.
Veröffentlicht: (2025)
Tunable Dual-Type Weyl Points in Dirac-Weyl Semimetal CaAgBi
von: Huang, Shenghao, et al.
Veröffentlicht: (2026)
von: Huang, Shenghao, et al.
Veröffentlicht: (2026)
Evaluating Rater Effects of Large Language Models in Automated Essay Scoring: GPT, Claude, Gemini, and DeepSeek
von: Hong Jiao, et al.
Veröffentlicht: (2026)
von: Hong Jiao, et al.
Veröffentlicht: (2026)
Free Lunch for Pass@$k$? Low Cost Diverse Sampling for Diffusion Language Models
von: Lamont, Sean, et al.
Veröffentlicht: (2026)
von: Lamont, Sean, et al.
Veröffentlicht: (2026)
Seeking Nash Equilibrium in Non-cooperative Quadratic Games Under Delayed Information Exchange
von: Jiang, Kaichen, et al.
Veröffentlicht: (2026)
von: Jiang, Kaichen, et al.
Veröffentlicht: (2026)
The Impact of Generative Artificial Intelligence on Market Equilibrium: Evidence from a Natural Experiment
von: Zhang, Kaichen, et al.
Veröffentlicht: (2023)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2023)
Extremum-Seeking Action Selection for Accelerating Policy Optimization
von: Chang, Ya-Chien, et al.
Veröffentlicht: (2024)
von: Chang, Ya-Chien, et al.
Veröffentlicht: (2024)
TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models
von: Choo, Jinho, et al.
Veröffentlicht: (2026)
von: Choo, Jinho, et al.
Veröffentlicht: (2026)
ROS: A GNN-based Relax-Optimize-and-Sample Framework for Max-k-Cut Problems
von: Qiu, Yeqing, et al.
Veröffentlicht: (2024)
von: Qiu, Yeqing, et al.
Veröffentlicht: (2024)
Quantum Approximate Optimization of Integer Graph Problems and Surpassing Semidefinite Programming for Max-k-Cut
von: Apte, Anuj, et al.
Veröffentlicht: (2026)
von: Apte, Anuj, et al.
Veröffentlicht: (2026)
Data Clustering and Visualization with Recursive Max k-Cut Algorithm
von: Ly, An, et al.
Veröffentlicht: (2024)
von: Ly, An, et al.
Veröffentlicht: (2024)
A $(k+1)$-partite entanglement measure of $N$-partite quantum states
von: Hong, Yan, et al.
Veröffentlicht: (2022)
von: Hong, Yan, et al.
Veröffentlicht: (2022)
Associated varieties of simple affine VOAs $L_k(sl_3)$ and $W$-algebras $W_k(sl_3,f)$
von: Jiang, Cuipo, et al.
Veröffentlicht: (2024)
von: Jiang, Cuipo, et al.
Veröffentlicht: (2024)
Block-transitive $t$-($k^2,k,λ$) designs with $PSL(n,q)$ as socle
von: Xiong, Guoqiang, et al.
Veröffentlicht: (2025)
von: Xiong, Guoqiang, et al.
Veröffentlicht: (2025)
A Geometric Approach to $k$-means
von: Hong, Jiazhen, et al.
Veröffentlicht: (2022)
von: Hong, Jiazhen, et al.
Veröffentlicht: (2022)
Antidirected hamiltonian paths in $k$-hypertournaments
von: Yang, Hong, et al.
Veröffentlicht: (2024)
von: Yang, Hong, et al.
Veröffentlicht: (2024)
Metrics of positive Ricci curvature on simply‐connected manifolds of dimension 6k$6k$
von: Philipp Reiser
Veröffentlicht: (2024)
von: Philipp Reiser
Veröffentlicht: (2024)
Ähnliche Einträge
-
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025) -
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models
von: Fan, Yuchen, et al.
Veröffentlicht: (2024) -
Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model
von: Hong, Yuzhong, et al.
Veröffentlicht: (2024) -
Multi-Turn Interactions for Text-to-SQL with Large Language Models
von: Xiong, Guanming, et al.
Veröffentlicht: (2024) -
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)