Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: He, Haoran, Ye, Yuxiao, Cai, Qingpeng, Hu, Chen, Jiao, Binxing, Jiang, Daxin, Pan, Ling
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!