When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Jiakang, Liu, Runze, Cai, Qingpeng, Lin, Lei, Hu, Wenping, Li, Xiu, Zhang, Fuzheng, Zhou, Guorui, Gai, Kun, Pan, Ling |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
por: Liu, Runze, et al.
Publicado: (2025)
por: Liu, Runze, et al.
Publicado: (2025)
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
por: Wang, Jiakang, et al.
Publicado: (2025)
por: Wang, Jiakang, et al.
Publicado: (2025)
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
por: Su, Zhenpeng, et al.
Publicado: (2025)
por: Su, Zhenpeng, et al.
Publicado: (2025)
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
por: Su, Zhenpeng, et al.
Publicado: (2025)
por: Su, Zhenpeng, et al.
Publicado: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
por: Su, Zhenpeng, et al.
Publicado: (2025)
por: Su, Zhenpeng, et al.
Publicado: (2025)
Misallocation in the Chinese land market
por: Xuan Fei, et al.
Publicado: (2024)
por: Xuan Fei, et al.
Publicado: (2024)
AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender Systems
por: Xue, Zhenghai, et al.
Publicado: (2023)
por: Xue, Zhenghai, et al.
Publicado: (2023)
Leanabell-Prover: Posttraining Scaling in Formal Reasoning
por: Zhang, Jingyuan, et al.
Publicado: (2025)
por: Zhang, Jingyuan, et al.
Publicado: (2025)
Random Policy Evaluation Uncovers Policies of Generative Flow Networks
por: He, Haoran, et al.
Publicado: (2024)
por: He, Haoran, et al.
Publicado: (2024)
Hierarchical Semantic RL: Tackling the Problem of Dynamic Action Space for RL-based Recommendations
por: Wang, Minmao, et al.
Publicado: (2025)
por: Wang, Minmao, et al.
Publicado: (2025)
State Regularized Policy Optimization on Data with Dynamics Shift
por: Xue, Zhenghai, et al.
Publicado: (2023)
por: Xue, Zhenghai, et al.
Publicado: (2023)
AIS: Adaptive Importance Sampling for Quantized RL
por: Zhou, Jiajun, et al.
Publicado: (2026)
por: Zhou, Jiajun, et al.
Publicado: (2026)
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
por: Wang, Shaobo, et al.
Publicado: (2026)
por: Wang, Shaobo, et al.
Publicado: (2026)
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
por: Ji, Xingguang, et al.
Publicado: (2025)
por: Ji, Xingguang, et al.
Publicado: (2025)
Chaos and Misallocation under Price Controls
por: Albrecht, Brian C., et al.
Publicado: (2026)
por: Albrecht, Brian C., et al.
Publicado: (2026)
Symposium on Misallocation and Structural Transformation: Introduction
por: Tasso Adamopoulos, et al.
Publicado: (2024)
por: Tasso Adamopoulos, et al.
Publicado: (2024)
Production Function Estimation With Resource Misallocation
por: Shigang Li, et al.
Publicado: (2026)
por: Shigang Li, et al.
Publicado: (2026)
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
por: Ou, Jiao, et al.
Publicado: (2024)
por: Ou, Jiao, et al.
Publicado: (2024)
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
por: Tang, Yihong, et al.
Publicado: (2024)
por: Tang, Yihong, et al.
Publicado: (2024)
Enhancing Role-playing Systems through Aggressive Queries: Evaluation and Improvement
por: Tang, Yihong, et al.
Publicado: (2024)
por: Tang, Yihong, et al.
Publicado: (2024)
HoME: Hierarchy of Multi-Gate Experts for Multi-Task Learning at Kuaishou
por: Wang, Xu, et al.
Publicado: (2024)
por: Wang, Xu, et al.
Publicado: (2024)
The Impact of New Digital Infrastructure on Resource Misallocation
por: Qunli Wang, et al.
Publicado: (2026)
por: Qunli Wang, et al.
Publicado: (2026)
MISS: Multi-Modal Tree Indexing and Searching with Lifelong Sequential Behavior for Retrieval Recommendation
por: Guo, Chengcheng, et al.
Publicado: (2025)
por: Guo, Chengcheng, et al.
Publicado: (2025)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
por: Cheng, Tianhao, et al.
Publicado: (2026)
por: Cheng, Tianhao, et al.
Publicado: (2026)
Future Impact Decomposition in Request-level Recommendations
por: Wang, Xiaobei, et al.
Publicado: (2024)
por: Wang, Xiaobei, et al.
Publicado: (2024)
Sequential Recommendation for Optimizing Both Immediate Feedback and Long-term Retention
por: Liu, Ziru, et al.
Publicado: (2024)
por: Liu, Ziru, et al.
Publicado: (2024)
Video Object Segmentation with Dynamic Query Modulation
por: Zhou, Hantao, et al.
Publicado: (2024)
por: Zhou, Hantao, et al.
Publicado: (2024)
Tournament-Based Performance Evaluation and Systematic Misallocation: Why Forced Ranking Systems Produce Random Outcomes
por: McEntire, Jeremy
Publicado: (2025)
por: McEntire, Jeremy
Publicado: (2025)
Bifurcated Generative Flow Networks
por: Li, Chunhui, et al.
Publicado: (2024)
por: Li, Chunhui, et al.
Publicado: (2024)
How Metro Expansion Influences Enterprise Labor Misallocation
por: Mengting Zhang, et al.
Publicado: (2026)
por: Mengting Zhang, et al.
Publicado: (2026)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
FIM: Frequency-Aware Multi-View Interest Modeling for Local-Life Service Recommendation
por: Wang, Guoquan, et al.
Publicado: (2025)
por: Wang, Guoquan, et al.
Publicado: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
por: Ou, Jiao, et al.
Publicado: (2023)
por: Ou, Jiao, et al.
Publicado: (2023)
Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
por: Lin, Lei, et al.
Publicado: (2023)
por: Lin, Lei, et al.
Publicado: (2023)
Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling
por: Tang, Yunhao, et al.
Publicado: (2024)
por: Tang, Yunhao, et al.
Publicado: (2024)
CRM: Retrieval Model with Controllable Condition
por: Liu, Chi, et al.
Publicado: (2024)
por: Liu, Chi, et al.
Publicado: (2024)
PROMISE: Process Reward Models Unlock Test-Time Scaling Laws in Generative Recommendations
por: Guo, Chengcheng, et al.
Publicado: (2026)
por: Guo, Chengcheng, et al.
Publicado: (2026)
From Principles to Applications: A Comprehensive Survey of Discrete Tokenizers in Generation, Comprehension, Recommendation, and Information Retrieval
por: Jia, Jian, et al.
Publicado: (2025)
por: Jia, Jian, et al.
Publicado: (2025)
The RNA-seq raw data
por: Pan, Wenping
Publicado: (2030)
por: Pan, Wenping
Publicado: (2030)
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
por: Lei, Shiye, et al.
Publicado: (2026)
por: Lei, Shiye, et al.
Publicado: (2026)
Ejemplares similares
-
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
por: Liu, Runze, et al.
Publicado: (2025) -
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
por: Wang, Jiakang, et al.
Publicado: (2025) -
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
por: Su, Zhenpeng, et al.
Publicado: (2025) -
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
por: Su, Zhenpeng, et al.
Publicado: (2025) -
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
por: Su, Zhenpeng, et al.
Publicado: (2025)