Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Jiasheng, Cao, Boxi, Yu, Boxi, Zhang, Yuzhong, Cao, Jialun, Lu, Yaojie, Lin, Hongyu, Han, Xianpei, Sun, Le
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916064863977472
author Zheng, Jiasheng
Cao, Boxi
Yu, Boxi
Zhang, Yuzhong
Cao, Jialun
Lu, Yaojie
Lin, Hongyu
Han, Xianpei
Sun, Le
author_facet Zheng, Jiasheng
Cao, Boxi
Yu, Boxi
Zhang, Yuzhong
Cao, Jialun
Lu, Yaojie
Lin, Hongyu
Han, Xianpei
Sun, Le
contents Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However, the scalability of RLVR is severely constrained by the scarcity of sufficiently challenging verifiable code tasks that target near the model's edge of competence. Prior studies often rely on heuristic seed expansions for data synthesis, which severely limits both novelty and difficulty. Consequently, the training value of such data fails to scale proportionally with the size of its synthesis. To this end, we propose Atomic Decomposition and Recombination (ADR), a novel framework that generates verifiable code tasks via decomposition into atomic elements and controlled recombination, thereby enabling the generation of genuinely novel and challenging verifiable code tasks. Experiments and analysis demonstrate that ADR achieves superior originality, difficulty, diversity, and test quality over existing baselines, and consistently delivers greater improvements in code ability across RLVR in diverse downstream domains, including algorithmic programming, tool usage, and data science. Our work sheds light on a new paradigm for novel code task synthesis and scalable RLVR training.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31058
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
Zheng, Jiasheng
Cao, Boxi
Yu, Boxi
Zhang, Yuzhong
Cao, Jialun
Lu, Yaojie
Lin, Hongyu
Han, Xianpei
Sun, Le
Computation and Language
Software Engineering
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However, the scalability of RLVR is severely constrained by the scarcity of sufficiently challenging verifiable code tasks that target near the model's edge of competence. Prior studies often rely on heuristic seed expansions for data synthesis, which severely limits both novelty and difficulty. Consequently, the training value of such data fails to scale proportionally with the size of its synthesis. To this end, we propose Atomic Decomposition and Recombination (ADR), a novel framework that generates verifiable code tasks via decomposition into atomic elements and controlled recombination, thereby enabling the generation of genuinely novel and challenging verifiable code tasks. Experiments and analysis demonstrate that ADR achieves superior originality, difficulty, diversity, and test quality over existing baselines, and consistently delivers greater improvements in code ability across RLVR in diverse downstream domains, including algorithmic programming, tool usage, and data science. Our work sheds light on a new paradigm for novel code task synthesis and scalable RLVR training.
title Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2605.31058