Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Grey, Markov, Segerie, Charbel-Raphaël |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats
par: Grey, Markov, et autres
Publié: (2025)
par: Grey, Markov, et autres
Publié: (2025)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
par: Dorn, Diego, et autres
Publié: (2024)
par: Dorn, Diego, et autres
Publié: (2024)
The bitter lesson of misuse detection
par: Mariaccia, Hadrien, et autres
Publié: (2025)
par: Mariaccia, Hadrien, et autres
Publié: (2025)
Unpacking Human-AI Interaction in Safety-Critical Industries: A Systematic Literature Review
par: Bach, Tita A., et autres
Publié: (2023)
par: Bach, Tita A., et autres
Publié: (2023)
Continuous Time Continuous Space Homeostatic Reinforcement Learning (CTCS-HRRL) : Towards Biological Self-Autonomous Agent
par: Laurencon, Hugo, et autres
Publié: (2024)
par: Laurencon, Hugo, et autres
Publié: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
par: Röttger, Paul, et autres
Publié: (2024)
par: Röttger, Paul, et autres
Publié: (2024)
Mechanistic Interpretability for AI Safety -- A Review
par: Bereska, Leonard, et autres
Publié: (2024)
par: Bereska, Leonard, et autres
Publié: (2024)
Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering
par: Sammour, Farouq, et autres
Publié: (2024)
par: Sammour, Farouq, et autres
Publié: (2024)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
par: Vijayvargiya, Sanidhya, et autres
Publié: (2025)
par: Vijayvargiya, Sanidhya, et autres
Publié: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
par: Scholefield, Rebecca, et autres
Publié: (2025)
par: Scholefield, Rebecca, et autres
Publié: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
par: Siu, Vincent, et autres
Publié: (2025)
par: Siu, Vincent, et autres
Publié: (2025)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
par: Yueh-Han, Chen, et autres
Publié: (2025)
par: Yueh-Han, Chen, et autres
Publié: (2025)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
par: Xie, Tinghao, et autres
Publié: (2024)
par: Xie, Tinghao, et autres
Publié: (2024)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
par: Vaccaro, Michelle, et autres
Publié: (2026)
par: Vaccaro, Michelle, et autres
Publié: (2026)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
par: Dobbe, Roel
Publié: (2025)
par: Dobbe, Roel
Publié: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
par: Ren, Richard, et autres
Publié: (2024)
par: Ren, Richard, et autres
Publié: (2024)
SafePro: Evaluating the Safety of Professional-Level AI Agents
par: Zhou, Kaiwen, et autres
Publié: (2026)
par: Zhou, Kaiwen, et autres
Publié: (2026)
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
par: Clark, Hannah-Beth, et autres
Publié: (2025)
par: Clark, Hannah-Beth, et autres
Publié: (2025)
Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems
par: Fan, Yihe, et autres
Publié: (2025)
par: Fan, Yihe, et autres
Publié: (2025)
Generative AI for Requirements Engineering: A Systematic Literature Review
par: Cheng, Haowei, et autres
Publié: (2024)
par: Cheng, Haowei, et autres
Publié: (2024)
AI Safety: A Climb To Armageddon?
par: Cappelen, Herman, et autres
Publié: (2024)
par: Cappelen, Herman, et autres
Publié: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
par: Hilton, Benjamin, et autres
Publié: (2025)
par: Hilton, Benjamin, et autres
Publié: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
par: Costa, Mariana Lins
Publié: (2026)
par: Costa, Mariana Lins
Publié: (2026)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
par: Weidinger, Laura, et autres
Publié: (2024)
par: Weidinger, Laura, et autres
Publié: (2024)
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
par: François, Camille, et autres
Publié: (2025)
par: François, Camille, et autres
Publié: (2025)
NeuroAI for AI Safety
par: Mineault, Patrick, et autres
Publié: (2024)
par: Mineault, Patrick, et autres
Publié: (2024)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
par: Griffin, Charlie, et autres
Publié: (2024)
par: Griffin, Charlie, et autres
Publié: (2024)
A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care
par: Normand, Oliver, et autres
Publié: (2025)
par: Normand, Oliver, et autres
Publié: (2025)
Systematic Literature Review: Explainable AI Definitions and Challenges in Education
par: Altukhi, Zaid M., et autres
Publié: (2025)
par: Altukhi, Zaid M., et autres
Publié: (2025)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
par: Han, Tessa, et autres
Publié: (2024)
par: Han, Tessa, et autres
Publié: (2024)
AI Adoption in NGOs: A Systematic Literature Review
par: Rotter, Janne, et autres
Publié: (2025)
par: Rotter, Janne, et autres
Publié: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
par: Zhang, Zhexin, et autres
Publié: (2025)
par: Zhang, Zhexin, et autres
Publié: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
par: Clymer, Joshua, et autres
Publié: (2024)
par: Clymer, Joshua, et autres
Publié: (2024)
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops
par: Ghafoor, Zainab, et autres
Publié: (2026)
par: Ghafoor, Zainab, et autres
Publié: (2026)
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
par: Gringras, David
Publié: (2026)
par: Gringras, David
Publié: (2026)
Data-Driven Methods and AI in Engineering Design: A Systematic Literature Review Focusing on Challenges and Opportunities
par: Afifi, Nehal, et autres
Publié: (2025)
par: Afifi, Nehal, et autres
Publié: (2025)
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis
par: Holzner, Niklas, et autres
Publié: (2025)
par: Holzner, Niklas, et autres
Publié: (2025)
AI2-Active Safety: AI-enabled Interaction-aware Active Safety Analysis with Vehicle Dynamics
par: Wu, Keshu, et autres
Publié: (2025)
par: Wu, Keshu, et autres
Publié: (2025)
AI in Computational Thinking Education in Higher Education: A Systematic Literature Review
par: Rahimi, Ebrahim, et autres
Publié: (2025)
par: Rahimi, Ebrahim, et autres
Publié: (2025)
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
par: Li, Wenkai, et autres
Publié: (2026)
par: Li, Wenkai, et autres
Publié: (2026)
Documents similaires
-
The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats
par: Grey, Markov, et autres
Publié: (2025) -
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
par: Dorn, Diego, et autres
Publié: (2024) -
The bitter lesson of misuse detection
par: Mariaccia, Hadrien, et autres
Publié: (2025) -
Unpacking Human-AI Interaction in Safety-Critical Industries: A Systematic Literature Review
par: Bach, Tita A., et autres
Publié: (2023) -
Continuous Time Continuous Space Homeostatic Reinforcement Learning (CTCS-HRRL) : Towards Biological Self-Autonomous Agent
par: Laurencon, Hugo, et autres
Publié: (2024)