Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Yu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
von: Fan, Chongyu, et al.
Veröffentlicht: (2026)
von: Fan, Chongyu, et al.
Veröffentlicht: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach
von: Hairi, Fnu, et al.
Veröffentlicht: (2025)
von: Hairi, Fnu, et al.
Veröffentlicht: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
von: Xu, Donglai, et al.
Veröffentlicht: (2025)
von: Xu, Donglai, et al.
Veröffentlicht: (2025)
Evaluating Parameter Efficient Methods for RLVR
von: Yin, Qingyu, et al.
Veröffentlicht: (2025)
von: Yin, Qingyu, et al.
Veröffentlicht: (2025)
VL Norm: Rethink Loss Aggregation in RLVR
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
Efficient Prediction of Pass@k Scaling in Large Language Models
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
von: Kazdan, Joshua, et al.
Veröffentlicht: (2025)
Best-of-Majority: Minimax-Optimal Strategy for Pass@$k$ Inference Scaling
von: Di, Qiwei, et al.
Veröffentlicht: (2025)
von: Di, Qiwei, et al.
Veröffentlicht: (2025)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
von: Liu, Zhanyu, et al.
Veröffentlicht: (2026)
von: Liu, Zhanyu, et al.
Veröffentlicht: (2026)
Not only where, But when: Temporal Scheduling for RLVR
von: Zhang, Jinghao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2026)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors
von: Yuan, Chaohao, et al.
Veröffentlicht: (2026)
von: Yuan, Chaohao, et al.
Veröffentlicht: (2026)
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR
von: Mou, Chaoli, et al.
Veröffentlicht: (2026)
von: Mou, Chaoli, et al.
Veröffentlicht: (2026)
Network Embedding Exploration Tool (NEExT)
von: Dehghan, Ashkan, et al.
Veröffentlicht: (2025)
von: Dehghan, Ashkan, et al.
Veröffentlicht: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
von: Hao, Zhezheng, et al.
Veröffentlicht: (2025)
The Unlearnability Phenomenon in RLVR for Language Models
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
Improving Intrinsic Exploration by Creating Stationary Objectives
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2023)
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2023)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Li, Kaiyuan, et al.
Veröffentlicht: (2026)
Leveraging LLM Inconsistency to Boost Pass@k Performance
von: Dalal, Uri, et al.
Veröffentlicht: (2025)
von: Dalal, Uri, et al.
Veröffentlicht: (2025)
Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning
von: Alam, Md Tanvirul, et al.
Veröffentlicht: (2025)
von: Alam, Md Tanvirul, et al.
Veröffentlicht: (2025)
Data-Efficient RLVR via Off-Policy Influence Guidance
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
von: Zhu, Erle, et al.
Veröffentlicht: (2025)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
von: Min, Zijun, et al.
Veröffentlicht: (2026)
von: Min, Zijun, et al.
Veröffentlicht: (2026)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Spurious Rewards: Rethinking Training Signals in RLVR
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
von: Mitsuhashi, Ryo, et al.
Veröffentlicht: (2026)
von: Mitsuhashi, Ryo, et al.
Veröffentlicht: (2026)
LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
von: Kim, Chanyoung, et al.
Veröffentlicht: (2026)
von: Kim, Chanyoung, et al.
Veröffentlicht: (2026)
PubSwap: Public-Data Off-Policy Coordination for Federated RLVR
von: Nayak, Anupam, et al.
Veröffentlicht: (2026)
von: Nayak, Anupam, et al.
Veröffentlicht: (2026)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
The Path Not Taken: RLVR Provably Learns Off the Principals
von: Zhu, Hanqing, et al.
Veröffentlicht: (2025)
von: Zhu, Hanqing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025) -
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
von: Fan, Chongyu, et al.
Veröffentlicht: (2026) -
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025) -
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026) -
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)