Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoon, Deokgyu, Kang, Hyungkyu, Lee, Joongkyu, Kim, Byeongchan, Shin, Gyungin, Park, Sungrae, Oh, Min-hwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!