TwinWeaver: An LLM-Based Foundation Model Framework for Pan-Cancer Digital Twins
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915787952881664 |
|---|---|
| author | Makarov, Nikita Bordukova, Maria von Voithenberg, Lena Voith Pivel-Villanueva, Estrella Mielke, Sabrina Wickes, Jonathan Wang, Hanchen Ma, Mingyu Derek Choi, Keunwoo Cho, Kyunghyun Ra, Stephen Rodriguez-Esteban, Raul Schmich, Fabian Menden, Michael |
| author_facet | Makarov, Nikita Bordukova, Maria von Voithenberg, Lena Voith Pivel-Villanueva, Estrella Mielke, Sabrina Wickes, Jonathan Wang, Hanchen Ma, Mingyu Derek Choi, Keunwoo Cho, Kyunghyun Ra, Stephen Rodriguez-Esteban, Raul Schmich, Fabian Menden, Michael |
| contents | Precision oncology requires forecasting clinical events and trajectories, yet modeling sparse, multi-modal clinical time series remains a critical challenge. We introduce TwinWeaver, an open-source framework that serializes longitudinal patient histories into text, enabling unified event prediction as well as forecasting with large language models, and use it to build Genie Digital Twin (GDT) on 93,054 patients across 20 cancer types. In benchmarks, GDT significantly reduces forecasting error, achieving a median Mean Absolute Scaled Error (MASE) of 0.87 compared to 0.97 for the strongest time-series baseline (p<0.001). Furthermore, GDT improves risk stratification, achieving an average concordance index (C-index) of 0.703 across survival, progression, and therapy switching tasks, surpassing the best baseline of 0.662. GDT also generalizes to out-of-distribution clinical trials, matching trained baselines at zero-shot and surpassing them with fine-tuning, achieving a median MASE of 0.75-0.88 and outperforming the strongest baseline in event prediction with an average C-index of 0.672 versus 0.648. Finally, TwinWeaver enables an interpretable clinical reasoning extension, providing a scalable and transparent foundation for longitudinal clinical modeling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_20906 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TwinWeaver: An LLM-Based Foundation Model Framework for Pan-Cancer Digital Twins Makarov, Nikita Bordukova, Maria von Voithenberg, Lena Voith Pivel-Villanueva, Estrella Mielke, Sabrina Wickes, Jonathan Wang, Hanchen Ma, Mingyu Derek Choi, Keunwoo Cho, Kyunghyun Ra, Stephen Rodriguez-Esteban, Raul Schmich, Fabian Menden, Michael Machine Learning Precision oncology requires forecasting clinical events and trajectories, yet modeling sparse, multi-modal clinical time series remains a critical challenge. We introduce TwinWeaver, an open-source framework that serializes longitudinal patient histories into text, enabling unified event prediction as well as forecasting with large language models, and use it to build Genie Digital Twin (GDT) on 93,054 patients across 20 cancer types. In benchmarks, GDT significantly reduces forecasting error, achieving a median Mean Absolute Scaled Error (MASE) of 0.87 compared to 0.97 for the strongest time-series baseline (p<0.001). Furthermore, GDT improves risk stratification, achieving an average concordance index (C-index) of 0.703 across survival, progression, and therapy switching tasks, surpassing the best baseline of 0.662. GDT also generalizes to out-of-distribution clinical trials, matching trained baselines at zero-shot and surpassing them with fine-tuning, achieving a median MASE of 0.75-0.88 and outperforming the strongest baseline in event prediction with an average C-index of 0.672 versus 0.648. Finally, TwinWeaver enables an interpretable clinical reasoning extension, providing a scalable and transparent foundation for longitudinal clinical modeling. |
| title | TwinWeaver: An LLM-Based Foundation Model Framework for Pan-Cancer Digital Twins |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2601.20906 |