SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Junteng, Fan, Yuanxiang, Jiang, Zhuo, Ding, Han, Hu, Yongyi, Zhang, Chi, Shi, Yiqi, Weng, Shitong, Chen, Aili, Chen, Shiqi, Huang, Yunan, Zhang, Mozhi, Zhao, Pengyu, Yan, Junjie, He, Junxian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909636343365632
author Liu, Junteng
Fan, Yuanxiang
Jiang, Zhuo
Ding, Han
Hu, Yongyi
Zhang, Chi
Shi, Yiqi
Weng, Shitong
Chen, Aili
Chen, Shiqi
Huang, Yunan
Zhang, Mozhi
Zhao, Pengyu
Yan, Junjie
He, Junxian
author_facet Liu, Junteng
Fan, Yuanxiang
Jiang, Zhuo
Ding, Han
Hu, Yongyi
Zhang, Chi
Shi, Yiqi
Weng, Shitong
Chen, Aili
Chen, Shiqi
Huang, Yunan
Zhang, Mozhi
Zhao, Pengyu
Yan, Junjie
He, Junxian
contents Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplored. This gap is partly due to the challenge of collecting diverse and verifiable reasoning data suitable for RL. We hypothesize that logical reasoning is critical for developing general reasoning capabilities, as logic forms a fundamental building block of reasoning. In this work, we present SynLogic, a data synthesis framework and dataset that generates diverse logical reasoning data at scale, encompassing 35 diverse logical reasoning tasks. The SynLogic approach enables controlled synthesis of data with adjustable difficulty and quantity. Importantly, all examples can be verified by simple rules, making them ideally suited for RL with verifiable rewards. In our experiments, we validate the effectiveness of RL training on the SynLogic dataset based on 7B and 32B models. SynLogic leads to state-of-the-art logical reasoning performance among open-source datasets, surpassing DeepSeek-R1-Distill-Qwen-32B by 6 points on BBEH. Furthermore, mixing SynLogic data with mathematical and coding tasks improves the training efficiency of these domains and significantly enhances reasoning generalization. Notably, our mixed training model outperforms DeepSeek-R1-Zero-Qwen-32B across multiple benchmarks. These findings position SynLogic as a valuable resource for advancing the broader reasoning capabilities of LLMs. We open-source both the data synthesis pipeline and the SynLogic dataset at https://github.com/MiniMax-AI/SynLogic.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19641
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
Liu, Junteng
Fan, Yuanxiang
Jiang, Zhuo
Ding, Han
Hu, Yongyi
Zhang, Chi
Shi, Yiqi
Weng, Shitong
Chen, Aili
Chen, Shiqi
Huang, Yunan
Zhang, Mozhi
Zhao, Pengyu
Yan, Junjie
He, Junxian
Artificial Intelligence
Computation and Language
Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplored. This gap is partly due to the challenge of collecting diverse and verifiable reasoning data suitable for RL. We hypothesize that logical reasoning is critical for developing general reasoning capabilities, as logic forms a fundamental building block of reasoning. In this work, we present SynLogic, a data synthesis framework and dataset that generates diverse logical reasoning data at scale, encompassing 35 diverse logical reasoning tasks. The SynLogic approach enables controlled synthesis of data with adjustable difficulty and quantity. Importantly, all examples can be verified by simple rules, making them ideally suited for RL with verifiable rewards. In our experiments, we validate the effectiveness of RL training on the SynLogic dataset based on 7B and 32B models. SynLogic leads to state-of-the-art logical reasoning performance among open-source datasets, surpassing DeepSeek-R1-Distill-Qwen-32B by 6 points on BBEH. Furthermore, mixing SynLogic data with mathematical and coding tasks improves the training efficiency of these domains and significantly enhances reasoning generalization. Notably, our mixed training model outperforms DeepSeek-R1-Zero-Qwen-32B across multiple benchmarks. These findings position SynLogic as a valuable resource for advancing the broader reasoning capabilities of LLMs. We open-source both the data synthesis pipeline and the SynLogic dataset at https://github.com/MiniMax-AI/SynLogic.
title SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.19641