AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Bowen, Dillon, Dombrowski, Ann-Kathrin, Gleave, Adam, Cundy, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025)
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
by: Cundy, Chris, et al.
Published: (2025)
by: Cundy, Chris, et al.
Published: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
by: Taufeeque, Mohammad, et al.
Published: (2026)
by: Taufeeque, Mohammad, et al.
Published: (2026)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)
by: Jo, Nathanael, et al.
Published: (2025)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Planning in a recurrent neural network that plays Sokoban
by: Taufeeque, Mohammad, et al.
Published: (2024)
by: Taufeeque, Mohammad, et al.
Published: (2024)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
Defining and Evaluating Physical Safety for Large Language Models
by: Tang, Yung-Chen, et al.
Published: (2024)
by: Tang, Yung-Chen, et al.
Published: (2024)
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
Scaling Trends for Data Poisoning in LLMs
by: Bowen, Dillon, et al.
Published: (2024)
by: Bowen, Dillon, et al.
Published: (2024)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
by: Cundy, Chris, et al.
Published: (2023)
by: Cundy, Chris, et al.
Published: (2023)
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025)
by: Kim, HyunJin, et al.
Published: (2025)
Understanding and Mitigating Risks of Generative AI in Financial Services
by: Gehrmann, Sebastian, et al.
Published: (2025)
by: Gehrmann, Sebastian, et al.
Published: (2025)
Towards Environmentally Equitable AI
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data
by: Nolte, Henrik, et al.
Published: (2025)
by: Nolte, Henrik, et al.
Published: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
by: Tang, Zeyu, et al.
Published: (2025)
by: Tang, Zeyu, et al.
Published: (2025)
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
by: Jurenka, Irina, et al.
Published: (2024)
by: Jurenka, Irina, et al.
Published: (2024)
Testing autonomous vehicles and AI: perspectives and challenges from cybersecurity, transparency, robustness and fairness
by: Llorca, David Fernández, et al.
Published: (2024)
by: Llorca, David Fernández, et al.
Published: (2024)
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
by: Reuel, Anka, et al.
Published: (2025)
by: Reuel, Anka, et al.
Published: (2025)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
by: Potham, Ram
Published: (2025)
by: Potham, Ram
Published: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
Evaluating AI Group Fairness: a Fuzzy Logic Perspective
by: Krasanakis, Emmanouil, et al.
Published: (2024)
by: Krasanakis, Emmanouil, et al.
Published: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
by: Patwardhan, Tejal, et al.
Published: (2025)
by: Patwardhan, Tejal, et al.
Published: (2025)
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
by: Wang, Miles, et al.
Published: (2026)
by: Wang, Miles, et al.
Published: (2026)
Watermarking Should Be Treated as a Monitoring Primitive
by: Aremu, Toluwani, et al.
Published: (2026)
by: Aremu, Toluwani, et al.
Published: (2026)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
Forecasting and Mitigating Disruptions in Public Bus Transit Services
by: Han, Chaeeun, et al.
Published: (2024)
by: Han, Chaeeun, et al.
Published: (2024)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Similar Items
-
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
by: Dombrowski, Ann-Kathrin, et al.
Published: (2025) -
Preference Learning with Lie Detectors can Induce Honesty or Evasion
by: Cundy, Chris, et al.
Published: (2025) -
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
by: Taufeeque, Mohammad, et al.
Published: (2026) -
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025) -
Position: AI Evaluations Should be Grounded on a Theory of Capability
by: Jo, Nathanael, et al.
Published: (2025)