MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Rongsheng, Chen, Junying, Ji, Ke, Cai, Zhenyang, Chen, Shunian, Yang, Yunjin, Wang, Benyou
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918086259507200
author Wang, Rongsheng
Chen, Junying
Ji, Ke
Cai, Zhenyang
Chen, Shunian
Yang, Yunjin
Wang, Benyou
author_facet Wang, Rongsheng
Chen, Junying
Ji, Ke
Cai, Zhenyang
Chen, Shunian
Yang, Yunjin
Wang, Benyou
contents Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce MedVideoCap-55K, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop MedGen, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy. We hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation. Our code and data is available at https://github.com/FreedomIntelligence/MedGen
format Preprint
id arxiv_https___arxiv_org_abs_2507_05675
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
Wang, Rongsheng
Chen, Junying
Ji, Ke
Cai, Zhenyang
Chen, Shunian
Yang, Yunjin
Wang, Benyou
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce MedVideoCap-55K, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop MedGen, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy. We hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation. Our code and data is available at https://github.com/FreedomIntelligence/MedGen
title MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.05675