Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Boyi, Che, Zora, Li, Nathaniel, Sehwag, Udari Madhushani, Götting, Jasper, Nedungadi, Samira, Michael, Julian, Yue, Summer, Hendrycks, Dan, Henderson, Peter, Wang, Zifan, Donoughe, Seth, Mazeika, Mantas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aggressive Compression Enables LLM Weight Theft
by: Brown, Davis, et al.
Published: (2026)
by: Brown, Davis, et al.
Published: (2026)
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
by: Götting, Jasper, et al.
Published: (2025)
by: Götting, Jasper, et al.
Published: (2025)
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
by: McCaslin, Tegan, et al.
Published: (2025)
by: McCaslin, Tegan, et al.
Published: (2025)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
Can LLMs be Scammed? A Baseline Measurement Study
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
by: Cook, Thomas, et al.
Published: (2025)
by: Cook, Thomas, et al.
Published: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
The Reality of AI and Biorisk
by: Peppin, Aidan, et al.
Published: (2024)
by: Peppin, Aidan, et al.
Published: (2024)
AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
by: Karthikeyan, Harish, et al.
Published: (2025)
by: Karthikeyan, Harish, et al.
Published: (2025)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
by: Xie, Tinghao, et al.
Published: (2024)
by: Xie, Tinghao, et al.
Published: (2024)
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026)
by: Pu, George, et al.
Published: (2026)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
by: Ren, Richard, et al.
Published: (2025)
by: Ren, Richard, et al.
Published: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Lista preliminar de los caracoles terrestres de la Región Septentrional de Colombia
by: Götting, K.J.
Published: (1978)
by: Götting, K.J.
Published: (1978)
Lista preliminar de los caracoles terrestres de la Región Septentrional de Colombia
by: Götting, K.J.
Published: (1978)
by: Götting, K.J.
Published: (1978)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
by: Xiao, Yuchen, et al.
Published: (2023)
by: Xiao, Yuchen, et al.
Published: (2023)
A Critical Study of Classical Dance Education in Indian Universities
by: Dr Divya Nedungadi
Published: (2018)
by: Dr Divya Nedungadi
Published: (2018)
Representation Engineering: A Top-Down Approach to AI Transparency
by: Zou, Andy, et al.
Published: (2023)
by: Zou, Andy, et al.
Published: (2023)
Superintelligence Strategy: Expert Version
by: Hendrycks, Dan, et al.
Published: (2025)
by: Hendrycks, Dan, et al.
Published: (2025)
Pholoe minuta
by: Meißner, Karin, et al.
Published: (2016)
by: Meißner, Karin, et al.
Published: (2016)
A Heterogeneous Agent Model of Mortgage Servicing: An Income-based Relief Analysis
by: Garg, Deepeka, et al.
Published: (2024)
by: Garg, Deepeka, et al.
Published: (2024)
Continuous Sign Language Recognition with Adapted Conformer via Unsupervised Pretraining
by: Aloysius, Neena, et al.
Published: (2024)
by: Aloysius, Neena, et al.
Published: (2024)
AI Risk Management Should Incorporate Both Safety and Security
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
ATCAT: Astronomical Timeseries CAusal Transformer
by: Tung, Zora
Published: (2025)
by: Tung, Zora
Published: (2025)
Debajo del árbol florido (In Xochicuáhuitl Itzintlan) Nezahualcóyotl desde la tradición de la poesía prehispánica
by: Zora Rohousová
Published: (2005)
by: Zora Rohousová
Published: (2005)
Task Shifting, eHealth and Shared Decision‐Making—Preference Heterogeneity in the Adult Population for Developments in Outpatient Primary Healthcare
by: Zora Föhn
Published: (2025)
by: Zora Föhn
Published: (2025)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
Similar Items
-
Aggressive Compression Enables LLM Weight Theft
by: Brown, Davis, et al.
Published: (2026) -
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
by: Götting, Jasper, et al.
Published: (2025) -
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025) -
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024) -
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
by: McCaslin, Tegan, et al.
Published: (2025)