Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yeqi, Chen, Yue, Ye, Yanwei, Su, Guanhao, Mai, Luo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910276730748928
author Huang, Yeqi
Chen, Yue
Ye, Yanwei
Su, Guanhao
Mai, Luo
author_facet Huang, Yeqi
Chen, Yue
Ye, Yanwei
Su, Guanhao
Mai, Luo
contents General-purpose VLMs remain unreliable for biomedical research because valid answers in scientific papers depend on evidence split across figures, tables, charts, captions, and referring text. Existing post-training pipelines are bottlenecked by costly expert annotation and by synthetic data that drops this evidence structure. We present Ryze, a fully automated system that converts raw biomedical papers into an evidence-enriched training set and a domain-specialized VLM. Ryze synthesizes QA pairs with complete supporting evidence (visual element, caption, extracted structure, and referring paragraphs), reduces layout and OCR errors via chart/table-aware extraction and LLM-based cleansing, and applies a progress-gated post-training strategy combining supervised fine-tuning with reinforcement learning. Starting from Qwen3-VL-8B, Ryze produces BioVLM-8B at under USD 200, achieving 48.0% weighted accuracy on LAB-Bench, outperforming the base model by +12.6 percentage points (pp) and surpassing GPT-5.2 by +3.8 pp. We release Ryze as open source together with the trained BioVLM-8B model.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00902
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers
Huang, Yeqi
Chen, Yue
Ye, Yanwei
Su, Guanhao
Mai, Luo
Artificial Intelligence
General-purpose VLMs remain unreliable for biomedical research because valid answers in scientific papers depend on evidence split across figures, tables, charts, captions, and referring text. Existing post-training pipelines are bottlenecked by costly expert annotation and by synthetic data that drops this evidence structure. We present Ryze, a fully automated system that converts raw biomedical papers into an evidence-enriched training set and a domain-specialized VLM. Ryze synthesizes QA pairs with complete supporting evidence (visual element, caption, extracted structure, and referring paragraphs), reduces layout and OCR errors via chart/table-aware extraction and LLM-based cleansing, and applies a progress-gated post-training strategy combining supervised fine-tuning with reinforcement learning. Starting from Qwen3-VL-8B, Ryze produces BioVLM-8B at under USD 200, achieving 48.0% weighted accuracy on LAB-Bench, outperforming the base model by +12.6 percentage points (pp) and surpassing GPT-5.2 by +3.8 pp. We release Ryze as open source together with the trained BioVLM-8B model.
title Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers
topic Artificial Intelligence
url https://arxiv.org/abs/2606.00902