How Should AI Safety Benchmarks Benchmark Safety?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Cheng, Engelmann, Severin, Cao, Ruoxuan, Ali, Dalia, Papakyriakopoulos, Orestis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Information Retrieval Induced Safety Degradation in AI Agents
von: Yu, Cheng, et al.
Veröffentlicht: (2025)
von: Yu, Cheng, et al.
Veröffentlicht: (2025)
Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
von: Ali, Dalia, et al.
Veröffentlicht: (2025)
von: Ali, Dalia, et al.
Veröffentlicht: (2025)
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
von: Hutiri, Wiebke, et al.
Veröffentlicht: (2024)
von: Hutiri, Wiebke, et al.
Veröffentlicht: (2024)
Position: Measure Dataset Diversity, Don't Just Claim It
von: Zhao, Dora, et al.
Veröffentlicht: (2024)
von: Zhao, Dora, et al.
Veröffentlicht: (2024)
AI Adoption Across Mission-Driven Organizations
von: Ali, Dalia, et al.
Veröffentlicht: (2025)
von: Ali, Dalia, et al.
Veröffentlicht: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024)
von: Ren, Richard, et al.
Veröffentlicht: (2024)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
Benchmarking and Understanding Safety Risks in AI Character Platforms
von: Wei, Yiluo, et al.
Veröffentlicht: (2025)
von: Wei, Yiluo, et al.
Veröffentlicht: (2025)
AI Safety Frameworks Should Include Procedures for Model Access Decisions
von: Kembery, Edward, et al.
Veröffentlicht: (2024)
von: Kembery, Edward, et al.
Veröffentlicht: (2024)
AI Safety Should Prioritize the Future of Work
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
von: Clymer, Joshua, et al.
Veröffentlicht: (2024)
von: Clymer, Joshua, et al.
Veröffentlicht: (2024)
Countering Privacy Nihilism
von: Engelmann, Severin, et al.
Veröffentlicht: (2025)
von: Engelmann, Severin, et al.
Veröffentlicht: (2025)
Visions of a Discipline: Analyzing Introductory AI Courses on YouTube
von: Engelmann, Severin, et al.
Veröffentlicht: (2024)
von: Engelmann, Severin, et al.
Veröffentlicht: (2024)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
von: Yew, Rui-Jie, et al.
Veröffentlicht: (2025)
von: Yew, Rui-Jie, et al.
Veröffentlicht: (2025)
Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
Safety First: Psychological Safety as the Key to AI Transformation
von: Reich, Aaron, et al.
Veröffentlicht: (2026)
von: Reich, Aaron, et al.
Veröffentlicht: (2026)
Engaged AI Governance: Addressing the Last Mile Challenge Through Internal Expert Collaboration
von: Jarvers, Simon, et al.
Veröffentlicht: (2026)
von: Jarvers, Simon, et al.
Veröffentlicht: (2026)
AI Safety for Everyone
von: Gyevnar, Balint, et al.
Veröffentlicht: (2025)
von: Gyevnar, Balint, et al.
Veröffentlicht: (2025)
Generating Robot Constitutions & Benchmarks for Semantic Safety
von: Sermanet, Pierre, et al.
Veröffentlicht: (2025)
von: Sermanet, Pierre, et al.
Veröffentlicht: (2025)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
von: Fort, Kristina
Veröffentlicht: (2024)
von: Fort, Kristina
Veröffentlicht: (2024)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
von: Cobben, Pepijn, et al.
Veröffentlicht: (2026)
von: Cobben, Pepijn, et al.
Veröffentlicht: (2026)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
von: Vishwarupe, Varad, et al.
Veröffentlicht: (2026)
Safety cases for frontier AI
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2024)
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2024)
Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions
von: Jiao, Junfeng, et al.
Veröffentlicht: (2025)
von: Jiao, Junfeng, et al.
Veröffentlicht: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
von: Dobbe, Roel
Veröffentlicht: (2025)
von: Dobbe, Roel
Veröffentlicht: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
von: Nakao, Mahiro, et al.
Veröffentlicht: (2026)
von: Nakao, Mahiro, et al.
Veröffentlicht: (2026)
AI Safety, Alignment, and Ethics (AI SAE)
von: Waldner, Dylan
Veröffentlicht: (2025)
von: Waldner, Dylan
Veröffentlicht: (2025)
AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
von: Zeng, Yi, et al.
Veröffentlicht: (2024)
International AI Safety Report 2026
von: Bengio, Yoshua, et al.
Veröffentlicht: (2026)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2026)
The BIG Argument for AI Safety Cases
von: Habli, Ibrahim, et al.
Veröffentlicht: (2025)
von: Habli, Ibrahim, et al.
Veröffentlicht: (2025)
Persuasion and Safety in the Era of Generative AI
von: Kong, Haein
Veröffentlicht: (2025)
von: Kong, Haein
Veröffentlicht: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
von: Hilton, Benjamin, et al.
Veröffentlicht: (2025)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
von: Nair, Variath Madhupal Gautham, et al.
Veröffentlicht: (2025)
von: Nair, Variath Madhupal Gautham, et al.
Veröffentlicht: (2025)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
von: Akiri, Charankumar, et al.
Veröffentlicht: (2025)
von: Akiri, Charankumar, et al.
Veröffentlicht: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
von: Xiao, Yang, et al.
Veröffentlicht: (2023)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Information Retrieval Induced Safety Degradation in AI Agents
von: Yu, Cheng, et al.
Veröffentlicht: (2025) -
Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
von: Ali, Dalia, et al.
Veröffentlicht: (2025) -
Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
von: Hutiri, Wiebke, et al.
Veröffentlicht: (2024) -
Position: Measure Dataset Diversity, Don't Just Claim It
von: Zhao, Dora, et al.
Veröffentlicht: (2024) -
AI Adoption Across Mission-Driven Organizations
von: Ali, Dalia, et al.
Veröffentlicht: (2025)