Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
Fuente:
arXiv
Saved in:
| Main Authors: | Barakat, Anas, Chakraborty, Souradip, Pahwa, Khushbu, Bedi, Amrit Singh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025)
by: Trivedi, Prashant, et al.
Published: (2025)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
PROPS: Progressively Private Self-alignment of Large Language Models
by: Teku, Noel, et al.
Published: (2025)
by: Teku, Noel, et al.
Published: (2025)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
Leveraging LLM Inconsistency to Boost Pass@k Performance
by: Dalal, Uri, et al.
Published: (2025)
by: Dalal, Uri, et al.
Published: (2025)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
by: Lepel, Olivier, et al.
Published: (2024)
by: Lepel, Olivier, et al.
Published: (2024)
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)
by: Protopapas, Kimon, et al.
Published: (2024)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)
by: Bai, Qinbo, et al.
Published: (2022)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
by: Gaur, Mudit, et al.
Published: (2024)
by: Gaur, Mudit, et al.
Published: (2024)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
by: Sun, Xingpeng, et al.
Published: (2024)
by: Sun, Xingpeng, et al.
Published: (2024)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
by: Reddy, Avinash, et al.
Published: (2026)
by: Reddy, Avinash, et al.
Published: (2026)
Graph Unitary Message Passing
by: Qiu, Haiquan, et al.
Published: (2024)
by: Qiu, Haiquan, et al.
Published: (2024)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
by: Liu, Zehao, et al.
Published: (2026)
by: Liu, Zehao, et al.
Published: (2026)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025)
by: Chen, Zhipeng, et al.
Published: (2025)
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
by: Zhang, Huishuai, et al.
Published: (2025)
by: Zhang, Huishuai, et al.
Published: (2025)
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
by: Oota, Subba Reddy, et al.
Published: (2026)
by: Oota, Subba Reddy, et al.
Published: (2026)
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
by: Liu, Yiqun, et al.
Published: (2026)
by: Liu, Yiqun, et al.
Published: (2026)
Active Preference Optimization for Sample Efficient RLHF
by: Das, Nirjhar, et al.
Published: (2024)
by: Das, Nirjhar, et al.
Published: (2024)
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
by: Chen, Yuyang, et al.
Published: (2024)
by: Chen, Yuyang, et al.
Published: (2024)
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
by: Adeseye, Aisvarya, et al.
Published: (2026)
by: Adeseye, Aisvarya, et al.
Published: (2026)
Distributed Conformal Prediction via Message Passing
by: Wen, Haifeng, et al.
Published: (2025)
by: Wen, Haifeng, et al.
Published: (2025)
Neural-Symbolic Message Passing with Dynamic Pruning
by: Zhang, Chongzhi, et al.
Published: (2025)
by: Zhang, Chongzhi, et al.
Published: (2025)
Link Prediction with Untrained Message Passing Layers
by: Qarkaxhija, Lisi, et al.
Published: (2024)
by: Qarkaxhija, Lisi, et al.
Published: (2024)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
by: Chakraborty, Souradip, et al.
Published: (2023)
by: Chakraborty, Souradip, et al.
Published: (2023)
Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
by: Liu, Xiangyu, et al.
Published: (2023)
by: Liu, Xiangyu, et al.
Published: (2023)
Could Chemical LLMs benefit from Message Passing
by: Xie, Jiaqing, et al.
Published: (2024)
by: Xie, Jiaqing, et al.
Published: (2024)
Similar Items
-
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024) -
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
by: Trivedi, Prashant, et al.
Published: (2025) -
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024) -
PROPS: Progressively Private Self-alignment of Large Language Models
by: Teku, Noel, et al.
Published: (2025) -
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)