Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Götting, Jasper, Medeiros, Pedro, Sanders, Jon G, Li, Nathaniel, Phan, Long, Elabd, Karam, Justen, Lennart, Hendrycks, Dan, Donoughe, Seth
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912352559955968
author Götting, Jasper
Medeiros, Pedro
Sanders, Jon G
Li, Nathaniel
Phan, Long
Elabd, Karam
Justen, Lennart
Hendrycks, Dan
Donoughe, Seth
author_facet Götting, Jasper
Medeiros, Pedro
Sanders, Jon G
Li, Nathaniel
Phan, Long
Elabd, Karam
Justen, Lennart
Hendrycks, Dan
Donoughe, Seth
contents We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Constructed from the inputs of dozens of PhD-level expert virologists, VCT consists of $322$ multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. VCT is difficult: expert virologists with access to the internet score an average of $22.1\%$ on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI's o3, reaches $43.8\%$ accuracy, outperforming $94\%$ of expert virologists even within their sub-areas of specialization. The ability to provide expert-level virology troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations. We propose that the capability of LLMs to provide expert-level troubleshooting of dual-use virology work should be integrated into existing frameworks for handling dual-use technologies in the life sciences.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Götting, Jasper
Medeiros, Pedro
Sanders, Jon G
Li, Nathaniel
Phan, Long
Elabd, Karam
Justen, Lennart
Hendrycks, Dan
Donoughe, Seth
Computers and Society
Machine Learning
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Constructed from the inputs of dozens of PhD-level expert virologists, VCT consists of $322$ multimodal questions covering fundamental, tacit, and visual knowledge that is essential for practical work in virology laboratories. VCT is difficult: expert virologists with access to the internet score an average of $22.1\%$ on questions specifically in their sub-areas of expertise. However, the most performant LLM, OpenAI's o3, reaches $43.8\%$ accuracy, outperforming $94\%$ of expert virologists even within their sub-areas of specialization. The ability to provide expert-level virology troubleshooting is inherently dual-use: it is useful for beneficial research, but it can also be misused. Therefore, the fact that publicly available models outperform virologists on VCT raises pressing governance considerations. We propose that the capability of LLMs to provide expert-level troubleshooting of dual-use virology work should be integrated into existing frameworks for handling dual-use technologies in the life sciences.
title Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
topic Computers and Society
Machine Learning
url https://arxiv.org/abs/2504.16137