WildSci: Advancing Scientific Reasoning from In-the-Wild Literature

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Tengxiao, Nathani, Deepak, Li, Zekun, Yang, Kevin, Wang, William Yang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911363050242048
author Liu, Tengxiao
Nathani, Deepak
Li, Zekun
Yang, Kevin
Wang, William Yang
author_facet Liu, Tengxiao
Nathani, Deepak
Li, Zekun
Yang, Kevin
Wang, William Yang
contents Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are readily available. In contrast, progress in LLM reasoning models remains limited in scientific domains such as medicine and materials science due to limited dataset coverage and the inherent complexity of open-ended scientific questions. To address these challenges, we introduce WildSci, a new dataset of domain-specific science questions automatically synthesized from peer-reviewed literature, covering 9 scientific disciplines and 26 subdomains. By framing complex scientific reasoning tasks in a multiple-choice format, we enable scalable training with well-defined reward signals. We further apply reinforcement learning to finetune models on these data and analyze the resulting training dynamics, including domain-specific performance changes, response behaviors, and generalization trends. Experiments on a suite of scientific benchmarks demonstrate the effectiveness of our dataset and approach. We release WildSci to enable scalable and sustainable research in scientific reasoning, available at https://huggingface.co/datasets/JustinTX/WildSci.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05567
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
Liu, Tengxiao
Nathani, Deepak
Li, Zekun
Yang, Kevin
Wang, William Yang
Artificial Intelligence
Computation and Language
Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are readily available. In contrast, progress in LLM reasoning models remains limited in scientific domains such as medicine and materials science due to limited dataset coverage and the inherent complexity of open-ended scientific questions. To address these challenges, we introduce WildSci, a new dataset of domain-specific science questions automatically synthesized from peer-reviewed literature, covering 9 scientific disciplines and 26 subdomains. By framing complex scientific reasoning tasks in a multiple-choice format, we enable scalable training with well-defined reward signals. We further apply reinforcement learning to finetune models on these data and analyze the resulting training dynamics, including domain-specific performance changes, response behaviors, and generalization trends. Experiments on a suite of scientific benchmarks demonstrate the effectiveness of our dataset and approach. We release WildSci to enable scalable and sustainable research in scientific reasoning, available at https://huggingface.co/datasets/JustinTX/WildSci.
title WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.05567