Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
Fuente:
arXiv
Guardado en:
| Autores principales: | Grey, Markov, Segerie, Charbel-Raphaël |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats
por: Grey, Markov, et al.
Publicado: (2025)
por: Grey, Markov, et al.
Publicado: (2025)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
por: Dorn, Diego, et al.
Publicado: (2024)
por: Dorn, Diego, et al.
Publicado: (2024)
The bitter lesson of misuse detection
por: Mariaccia, Hadrien, et al.
Publicado: (2025)
por: Mariaccia, Hadrien, et al.
Publicado: (2025)
Unpacking Human-AI Interaction in Safety-Critical Industries: A Systematic Literature Review
por: Bach, Tita A., et al.
Publicado: (2023)
por: Bach, Tita A., et al.
Publicado: (2023)
Continuous Time Continuous Space Homeostatic Reinforcement Learning (CTCS-HRRL) : Towards Biological Self-Autonomous Agent
por: Laurencon, Hugo, et al.
Publicado: (2024)
por: Laurencon, Hugo, et al.
Publicado: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
por: Röttger, Paul, et al.
Publicado: (2024)
por: Röttger, Paul, et al.
Publicado: (2024)
Mechanistic Interpretability for AI Safety -- A Review
por: Bereska, Leonard, et al.
Publicado: (2024)
por: Bereska, Leonard, et al.
Publicado: (2024)
Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering
por: Sammour, Farouq, et al.
Publicado: (2024)
por: Sammour, Farouq, et al.
Publicado: (2024)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
por: Scholefield, Rebecca, et al.
Publicado: (2025)
por: Scholefield, Rebecca, et al.
Publicado: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
por: Yueh-Han, Chen, et al.
Publicado: (2025)
por: Yueh-Han, Chen, et al.
Publicado: (2025)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
por: Xie, Tinghao, et al.
Publicado: (2024)
por: Xie, Tinghao, et al.
Publicado: (2024)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
por: Vaccaro, Michelle, et al.
Publicado: (2026)
por: Vaccaro, Michelle, et al.
Publicado: (2026)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
por: Dobbe, Roel
Publicado: (2025)
por: Dobbe, Roel
Publicado: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
por: Ren, Richard, et al.
Publicado: (2024)
por: Ren, Richard, et al.
Publicado: (2024)
SafePro: Evaluating the Safety of Professional-Level AI Agents
por: Zhou, Kaiwen, et al.
Publicado: (2026)
por: Zhou, Kaiwen, et al.
Publicado: (2026)
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems
por: Fan, Yihe, et al.
Publicado: (2025)
por: Fan, Yihe, et al.
Publicado: (2025)
Generative AI for Requirements Engineering: A Systematic Literature Review
por: Cheng, Haowei, et al.
Publicado: (2024)
por: Cheng, Haowei, et al.
Publicado: (2024)
AI Safety: A Climb To Armageddon?
por: Cappelen, Herman, et al.
Publicado: (2024)
por: Cappelen, Herman, et al.
Publicado: (2024)
Safety Cases: A Scalable Approach to Frontier AI Safety
por: Hilton, Benjamin, et al.
Publicado: (2025)
por: Hilton, Benjamin, et al.
Publicado: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
por: Costa, Mariana Lins
Publicado: (2026)
por: Costa, Mariana Lins
Publicado: (2026)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
por: Weidinger, Laura, et al.
Publicado: (2024)
por: Weidinger, Laura, et al.
Publicado: (2024)
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
por: François, Camille, et al.
Publicado: (2025)
por: François, Camille, et al.
Publicado: (2025)
NeuroAI for AI Safety
por: Mineault, Patrick, et al.
Publicado: (2024)
por: Mineault, Patrick, et al.
Publicado: (2024)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
por: Griffin, Charlie, et al.
Publicado: (2024)
por: Griffin, Charlie, et al.
Publicado: (2024)
A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care
por: Normand, Oliver, et al.
Publicado: (2025)
por: Normand, Oliver, et al.
Publicado: (2025)
Systematic Literature Review: Explainable AI Definitions and Challenges in Education
por: Altukhi, Zaid M., et al.
Publicado: (2025)
por: Altukhi, Zaid M., et al.
Publicado: (2025)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
por: Han, Tessa, et al.
Publicado: (2024)
por: Han, Tessa, et al.
Publicado: (2024)
AI Adoption in NGOs: A Systematic Literature Review
por: Rotter, Janne, et al.
Publicado: (2025)
por: Rotter, Janne, et al.
Publicado: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
por: Zhang, Zhexin, et al.
Publicado: (2025)
por: Zhang, Zhexin, et al.
Publicado: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
por: Clymer, Joshua, et al.
Publicado: (2024)
por: Clymer, Joshua, et al.
Publicado: (2024)
Improving the Safety and Trustworthiness of Medical AI via Multi-Agent Evaluation Loops
por: Ghafoor, Zainab, et al.
Publicado: (2026)
por: Ghafoor, Zainab, et al.
Publicado: (2026)
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
por: Gringras, David
Publicado: (2026)
por: Gringras, David
Publicado: (2026)
Data-Driven Methods and AI in Engineering Design: A Systematic Literature Review Focusing on Challenges and Opportunities
por: Afifi, Nehal, et al.
Publicado: (2025)
por: Afifi, Nehal, et al.
Publicado: (2025)
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis
por: Holzner, Niklas, et al.
Publicado: (2025)
por: Holzner, Niklas, et al.
Publicado: (2025)
AI2-Active Safety: AI-enabled Interaction-aware Active Safety Analysis with Vehicle Dynamics
por: Wu, Keshu, et al.
Publicado: (2025)
por: Wu, Keshu, et al.
Publicado: (2025)
AI in Computational Thinking Education in Higher Education: A Systematic Literature Review
por: Rahimi, Ebrahim, et al.
Publicado: (2025)
por: Rahimi, Ebrahim, et al.
Publicado: (2025)
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
por: Li, Wenkai, et al.
Publicado: (2026)
por: Li, Wenkai, et al.
Publicado: (2026)
Ejemplares similares
-
The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats
por: Grey, Markov, et al.
Publicado: (2025) -
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
por: Dorn, Diego, et al.
Publicado: (2024) -
The bitter lesson of misuse detection
por: Mariaccia, Hadrien, et al.
Publicado: (2025) -
Unpacking Human-AI Interaction in Safety-Critical Industries: A Systematic Literature Review
por: Bach, Tita A., et al.
Publicado: (2023) -
Continuous Time Continuous Space Homeostatic Reinforcement Learning (CTCS-HRRL) : Towards Biological Self-Autonomous Agent
por: Laurencon, Hugo, et al.
Publicado: (2024)