CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jianyou, Zheng, Youze, Bao, Longtian, Zhang, Hanyuan, Zheng, Qirui, Chen, Yuhan, Zhang, Yang, Feng, Matthew, Khan, Maxim, Sehgal, Aditya K., Rosin, Christopher D., Paturi, Ramamohan, Dube, Umber, Bergen, Leon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914485837496320
author Wang, Jianyou
Zheng, Youze
Bao, Longtian
Zhang, Hanyuan
Zheng, Qirui
Chen, Yuhan
Zhang, Yang
Feng, Matthew
Khan, Maxim
Sehgal, Aditya K.
Rosin, Christopher D.
Paturi, Ramamohan
Dube, Umber
Bergen, Leon
author_facet Wang, Jianyou
Zheng, Youze
Bao, Longtian
Zhang, Hanyuan
Zheng, Qirui
Chen, Yuhan
Zhang, Yang
Feng, Matthew
Khan, Maxim
Sehgal, Aditya K.
Rosin, Christopher D.
Paturi, Ramamohan
Dube, Umber
Bergen, Leon
contents Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We study this question through clinical trial outcome prediction, a high-stakes open challenge even for domain experts. We introduce CT Open, an open-access, live platform that will run four challenge every year. Anyone can submit predictions for each challenge. CT Open evaluates those submissions on trials whose outcomes were not yet public at the time of submission but were made public afterwards. Determining if a trial's outcome is public on the internet before a certain date is surprisingly difficult. Outcomes posted on official registries may lag behind by years, while the first mention may appear in obscure articles. To address this, we propose a novel, fully automated decontamination pipeline that uses iterative LLM-powered web search to identify the earliest mention of trial outcomes. We validate the pipeline's quality and accuracy by human expert's annotations. Since CT Open's pipeline ensures that every evaluated trial had no publicly reported outcome when the prediction was made, it allows participants to use any methodology and any data source. In this paper, we release a training set and two time-stamped test benchmarks, Winter 2025 and Summer 2025. We believe CT Open can serve as a central hub for advancing AI research on forecasting real-world outcomes before they occur, while also informing biomedical research and improving clinical trial design. CT Open Platform is hosted at $\href{https://ct-open.net/}{https://ct-open.net/}$
format Preprint
id arxiv_https___arxiv_org_abs_2604_16742
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction
Wang, Jianyou
Zheng, Youze
Bao, Longtian
Zhang, Hanyuan
Zheng, Qirui
Chen, Yuhan
Zhang, Yang
Feng, Matthew
Khan, Maxim
Sehgal, Aditya K.
Rosin, Christopher D.
Paturi, Ramamohan
Dube, Umber
Bergen, Leon
Artificial Intelligence
Computation and Language
Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We study this question through clinical trial outcome prediction, a high-stakes open challenge even for domain experts. We introduce CT Open, an open-access, live platform that will run four challenge every year. Anyone can submit predictions for each challenge. CT Open evaluates those submissions on trials whose outcomes were not yet public at the time of submission but were made public afterwards. Determining if a trial's outcome is public on the internet before a certain date is surprisingly difficult. Outcomes posted on official registries may lag behind by years, while the first mention may appear in obscure articles. To address this, we propose a novel, fully automated decontamination pipeline that uses iterative LLM-powered web search to identify the earliest mention of trial outcomes. We validate the pipeline's quality and accuracy by human expert's annotations. Since CT Open's pipeline ensures that every evaluated trial had no publicly reported outcome when the prediction was made, it allows participants to use any methodology and any data source. In this paper, we release a training set and two time-stamped test benchmarks, Winter 2025 and Summer 2025. We believe CT Open can serve as a central hub for advancing AI research on forecasting real-world outcomes before they occur, while also informing biomedical research and improving clinical trial design. CT Open Platform is hosted at $\href{https://ct-open.net/}{https://ct-open.net/}$
title CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.16742