Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Ximing, Acuna, David, Jung, Jaehun, Hu, Jian, Zhang, Di, Diao, Shizhe, Zou, Yunheng, Zhang, Shaokun, Cui, Brandon, Liu, Mingjie, Kim, Hyunwoo, Ammanabrolu, Prithviraj, Kautz, Jan, Dong, Yi, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
by: Cui, Brandon, et al.
Published: (2026)
by: Cui, Brandon, et al.
Published: (2026)
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
by: Kim, Bosung, et al.
Published: (2026)
by: Kim, Bosung, et al.
Published: (2026)
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
by: Wang, Ruiyi, et al.
Published: (2025)
by: Wang, Ruiyi, et al.
Published: (2025)
Polar: Agentic RL on Any Harness at Scale
by: Xu, Binfeng, et al.
Published: (2026)
by: Xu, Binfeng, et al.
Published: (2026)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
by: Wu, Fang, et al.
Published: (2025)
by: Wu, Fang, et al.
Published: (2025)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
by: Dionisopoulos, Lucas, et al.
Published: (2026)
by: Dionisopoulos, Lucas, et al.
Published: (2026)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Critique-out-Loud Reward Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)
by: Lu, Ximing, et al.
Published: (2025)
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
by: Fisher, Jillian, et al.
Published: (2024)
by: Fisher, Jillian, et al.
Published: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Preference-Based Learning in Audio Applications: A Systematic Analysis
by: Broukhim, Aaron, et al.
Published: (2025)
by: Broukhim, Aaron, et al.
Published: (2025)
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
by: Cui, Christopher Z., et al.
Published: (2026)
by: Cui, Christopher Z., et al.
Published: (2026)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
by: Shen, Yiran, et al.
Published: (2025)
by: Shen, Yiran, et al.
Published: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
by: Hallinan, Skyler, et al.
Published: (2025)
by: Hallinan, Skyler, et al.
Published: (2025)
Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
by: Kim, Hyunwoo, et al.
Published: (2026)
by: Kim, Hyunwoo, et al.
Published: (2026)
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
by: Liu, Shih-Yang, et al.
Published: (2025)
by: Liu, Shih-Yang, et al.
Published: (2025)
Information-Theoretic Distillation for Reference-less Summarization
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
by: Jung, Jaehun, et al.
Published: (2025)
by: Jung, Jaehun, et al.
Published: (2025)
TALES: Text Adventure Learning Environment Suite
by: Cui, Christopher Zhang, et al.
Published: (2025)
by: Cui, Christopher Zhang, et al.
Published: (2025)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
by: Jung, Jaehun, et al.
Published: (2023)
by: Jung, Jaehun, et al.
Published: (2023)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
by: Su, Hongjin, et al.
Published: (2025)
by: Su, Hongjin, et al.
Published: (2025)
CPS-TaskForge: Generating Collaborative Problem Solving Environments for Diverse Communication Tasks
by: Haduong, Nikita, et al.
Published: (2024)
by: Haduong, Nikita, et al.
Published: (2024)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
by: Seo, Wooseok, et al.
Published: (2025)
by: Seo, Wooseok, et al.
Published: (2025)
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
by: Shum, KaShun, et al.
Published: (2023)
by: Shum, KaShun, et al.
Published: (2023)
First Report of a Novel Goose Adenovirus Outbreak in Lion Head Gooses in China
by: Rongchang Liu, et al.
Published: (2024)
by: Rongchang Liu, et al.
Published: (2024)
The General's Goose
by: Robertson, Robbie
Published: (2017)
by: Robertson, Robbie
Published: (2017)
The Wild Goose
by: Mori, Ogai
Published: (2021)
by: Mori, Ogai
Published: (2021)
Write Better Conclusions With This One Simple Trick!
by: Abbott, Miriam
Published: (2025)
by: Abbott, Miriam
Published: (2025)
Experimental Study on the Geometrical and Mechanical Properties of Goose Eggshells
by: J Zhang
Published: (2017)
by: J Zhang
Published: (2017)
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
by: Saporta, Adriel, et al.
Published: (2024)
by: Saporta, Adriel, et al.
Published: (2024)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
Similar Items
-
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025) -
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
by: Cui, Brandon, et al.
Published: (2026) -
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
by: Kim, Bosung, et al.
Published: (2026) -
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
by: Acuna, David, et al.
Published: (2025) -
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025)