Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ackerman, Gary, Behlendorf, Brandon, Kallenborn, Zachary, Almakki, Sheriff, Clifford, Doug, LaTourette, Jenna, Peterson, Hayley, Sheinbaum, Noah, Shoemaker, Olivia, Wetzel, Anna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908699735359488
author Ackerman, Gary
Behlendorf, Brandon
Kallenborn, Zachary
Almakki, Sheriff
Clifford, Doug
LaTourette, Jenna
Peterson, Hayley
Sheinbaum, Noah
Shoemaker, Olivia
Wetzel, Anna
author_facet Ackerman, Gary
Behlendorf, Brandon
Kallenborn, Zachary
Almakki, Sheriff
Clifford, Doug
LaTourette, Jenna
Peterson, Hayley
Sheinbaum, Noah
Shoemaker, Olivia
Wetzel, Anna
contents Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons. An important element of such efforts is the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper describes the first component of a novel Biothreat Benchmark Generation (BBG) Framework. The BBG approach is designed to help model developers and evaluators reliably measure and assess the biosecurity risk uplift and general harm potential of existing and future AI models, while accounting for key aspects of the threat itself that are often overlooked in other benchmarking efforts, including different actor capability levels, and operational (in addition to purely technical) risk factors. As a pilot, the BBG is first being developed to address bacterial biological threats only. The BBG is built upon a hierarchical structure of biothreat categories, elements and tasks, which then serves as the basis for the development of task-aligned queries. This paper outlines the development of this biothreat task-query architecture, which we have named the Bacterial Biothreat Schema, while future papers will describe follow-on efforts to turn queries into model prompts, as well as how the resulting benchmarks can be implemented for model evaluation. Overall, the BBG Framework, including the Bacterial Biothreat Schema, seeks to offer a robust, re-usable structure for evaluating bacterial biological risks arising from LLMs across multiple levels of aggregation, which captures the full scope of technical and operational requirements for biological adversaries, and which accounts for a wide spectrum of biological adversary capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08130
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
Ackerman, Gary
Behlendorf, Brandon
Kallenborn, Zachary
Almakki, Sheriff
Clifford, Doug
LaTourette, Jenna
Peterson, Hayley
Sheinbaum, Noah
Shoemaker, Olivia
Wetzel, Anna
Machine Learning
Artificial Intelligence
Computers and Society
Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons. An important element of such efforts is the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper describes the first component of a novel Biothreat Benchmark Generation (BBG) Framework. The BBG approach is designed to help model developers and evaluators reliably measure and assess the biosecurity risk uplift and general harm potential of existing and future AI models, while accounting for key aspects of the threat itself that are often overlooked in other benchmarking efforts, including different actor capability levels, and operational (in addition to purely technical) risk factors. As a pilot, the BBG is first being developed to address bacterial biological threats only. The BBG is built upon a hierarchical structure of biothreat categories, elements and tasks, which then serves as the basis for the development of task-aligned queries. This paper outlines the development of this biothreat task-query architecture, which we have named the Bacterial Biothreat Schema, while future papers will describe follow-on efforts to turn queries into model prompts, as well as how the resulting benchmarks can be implemented for model evaluation. Overall, the BBG Framework, including the Bacterial Biothreat Schema, seeks to offer a robust, re-usable structure for evaluating bacterial biological risks arising from LLMs across multiple levels of aggregation, which captures the full scope of technical and operational requirements for biological adversaries, and which accounts for a wide spectrum of biological adversary capabilities.
title Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2512.08130