Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Heckel, Reinhard, Soltanolkotabi, Mahdi, Thramboulidis, Christos |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Theoretical Insights into Overparameterized Models in Multi-Task and Replay-Based Continual Learning
par: Banayeeanzade, Amin, et autres
Publié: (2024)
par: Banayeeanzade, Amin, et autres
Publié: (2024)
Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks
par: Wang, Kun, et autres
Publié: (2026)
par: Wang, Kun, et autres
Publié: (2026)
Measuring Fingerprints of Web-filtered Text Datasets and Fingerprint Propagation Through Training
par: Mansour, Youssef, et autres
Publié: (2024)
par: Mansour, Youssef, et autres
Publié: (2024)
A Deep Learning Method for Simultaneous Denoising and Missing Wedge Reconstruction in Cryogenic Electron Tomography
par: Wiedemann, Simon, et autres
Publié: (2023)
par: Wiedemann, Simon, et autres
Publié: (2023)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
par: Goel, Gautam, et autres
Publié: (2026)
par: Goel, Gautam, et autres
Publié: (2026)
Robustness of Deep Learning for Accelerated MRI: Benefits of Diverse Training Data
par: Lin, Kang, et autres
Publié: (2023)
par: Lin, Kang, et autres
Publié: (2023)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
par: Vural, Nuri Mert, et autres
Publié: (2026)
par: Vural, Nuri Mert, et autres
Publié: (2026)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
par: Cai, Xin-Qiang, et autres
Publié: (2025)
par: Cai, Xin-Qiang, et autres
Publié: (2025)
Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
par: Zeng, Guanning, et autres
Publié: (2025)
par: Zeng, Guanning, et autres
Publié: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
par: Lu, Xiaodong, et autres
Publié: (2026)
par: Lu, Xiaodong, et autres
Publié: (2026)
Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
par: Collins, Liam, et autres
Publié: (2023)
par: Collins, Liam, et autres
Publié: (2023)
Minerva: Reinforcement Learning with Verifiable Rewards for Cyber Threat Intelligence LLMs
par: Alam, Md Tanvirul, et autres
Publié: (2026)
par: Alam, Md Tanvirul, et autres
Publié: (2026)
Trace Reconstruction with Language Models
par: Weindel, Franziska, et autres
Publié: (2025)
par: Weindel, Franziska, et autres
Publié: (2025)
Transformer-Based Decoding in Concatenated Coding Schemes Under Synchronization Errors
par: Streit, Julian, et autres
Publié: (2025)
par: Streit, Julian, et autres
Publié: (2025)
Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards
par: Zhu, Yuxuan, et autres
Publié: (2026)
par: Zhu, Yuxuan, et autres
Publié: (2026)
ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization
par: Gan, Haosheng, et autres
Publié: (2025)
par: Gan, Haosheng, et autres
Publié: (2025)
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
par: Tang, Xinyu, et autres
Publié: (2025)
par: Tang, Xinyu, et autres
Publié: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
par: Gunjal, Anisha, et autres
Publié: (2025)
par: Gunjal, Anisha, et autres
Publié: (2025)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
par: Tang, Yunhao, et autres
Publié: (2025)
par: Tang, Yunhao, et autres
Publié: (2025)
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
par: Mroueh, Youssef
Publié: (2025)
par: Mroueh, Youssef
Publié: (2025)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
par: Kovačević, Filip, et autres
Publié: (2026)
par: Kovačević, Filip, et autres
Publié: (2026)
MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI
par: Arguello, Paula, et autres
Publié: (2026)
par: Arguello, Paula, et autres
Publié: (2026)
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
par: Yoon, Deokgyu, et autres
Publié: (2026)
par: Yoon, Deokgyu, et autres
Publié: (2026)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
par: Hu, Haoyu, et autres
Publié: (2026)
par: Hu, Haoyu, et autres
Publié: (2026)
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
par: Zhang, Feng, et autres
Publié: (2026)
par: Zhang, Feng, et autres
Publié: (2026)
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
par: Huang, Guanhua, et autres
Publié: (2025)
par: Huang, Guanhua, et autres
Publié: (2025)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
par: Nguyen, Hieu Trung, et autres
Publié: (2026)
par: Nguyen, Hieu Trung, et autres
Publié: (2026)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
par: Stojanovski, Zafir, et autres
Publié: (2025)
par: Stojanovski, Zafir, et autres
Publié: (2025)
Deep Learning for Accelerated and Robust MRI Reconstruction: a Review
par: Heckel, Reinhard, et autres
Publié: (2024)
par: Heckel, Reinhard, et autres
Publié: (2024)
Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning
par: Nguyen, Viet Bac, et autres
Publié: (2026)
par: Nguyen, Viet Bac, et autres
Publié: (2026)
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
par: Fabian, Zalan, et autres
Publié: (2023)
par: Fabian, Zalan, et autres
Publié: (2023)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
par: Zhang, Xin, et autres
Publié: (2026)
par: Zhang, Xin, et autres
Publié: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
par: Vasudeva, Bhavya, et autres
Publié: (2025)
par: Vasudeva, Bhavya, et autres
Publié: (2025)
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
par: Zhu, Speed, et autres
Publié: (2025)
par: Zhu, Speed, et autres
Publié: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
par: Ackermann, Johannes, et autres
Publié: (2026)
par: Ackermann, Johannes, et autres
Publié: (2026)
Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
par: Sinha, Sanchit, et autres
Publié: (2025)
par: Sinha, Sanchit, et autres
Publié: (2025)
Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
par: Wang, Zhen, et autres
Publié: (2025)
par: Wang, Zhen, et autres
Publié: (2025)
Reducing the Representation Error of GAN Image Priors Using the Deep Decoder
par: Daniels, Mara, et autres
Publié: (2020)
par: Daniels, Mara, et autres
Publié: (2020)
Resolution-Robust 3D MRI Reconstruction with 2D Diffusion Priors: Diverse-Resolution Training Outperforms Interpolation
par: Krainovic, Anselm, et autres
Publié: (2024)
par: Krainovic, Anselm, et autres
Publié: (2024)
Selector-Guided Autonomous Curriculum for One-Shot Reinforcement Learning from Verifiable Rewards
par: Dave, Rudray, et autres
Publié: (2026)
par: Dave, Rudray, et autres
Publié: (2026)
Documents similaires
-
Theoretical Insights into Overparameterized Models in Multi-Task and Replay-Based Continual Learning
par: Banayeeanzade, Amin, et autres
Publié: (2024) -
Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks
par: Wang, Kun, et autres
Publié: (2026) -
Measuring Fingerprints of Web-filtered Text Datasets and Fingerprint Propagation Through Training
par: Mansour, Youssef, et autres
Publié: (2024) -
A Deep Learning Method for Simultaneous Denoising and Missing Wedge Reconstruction in Cryogenic Electron Tomography
par: Wiedemann, Simon, et autres
Publié: (2023) -
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
par: Goel, Gautam, et autres
Publié: (2026)