Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Richard, Basart, Steven, Khoja, Adam, Gatti, Alice, Phan, Long, Yin, Xuwang, Mazeika, Mantas, Pan, Alexander, Mukobi, Gabriel, Kim, Ryan H., Fitz, Stephen, Hendrycks, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TextQuests: How Good are LLMs at Text-Based Video Games?
von: Phan, Long, et al.
Veröffentlicht: (2025)
von: Phan, Long, et al.
Veröffentlicht: (2025)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025)
von: Ren, Richard, et al.
Veröffentlicht: (2025)
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
von: Fitz, Stephen, et al.
Veröffentlicht: (2025)
von: Fitz, Stephen, et al.
Veröffentlicht: (2025)
Reducing Political Manipulation with Consistency Training
von: Phan, Long, et al.
Veröffentlicht: (2026)
von: Phan, Long, et al.
Veröffentlicht: (2026)
Introduction to AI Safety, Ethics, and Society
von: Hendrycks, Dan
Veröffentlicht: (2024)
von: Hendrycks, Dan
Veröffentlicht: (2024)
Introduction to AI Safety, Ethics, and Society
von: Hendrycks, Dan
Veröffentlicht: (2024)
von: Hendrycks, Dan
Veröffentlicht: (2024)
Representation Engineering: A Top-Down Approach to AI Transparency
von: Zou, Andy, et al.
Veröffentlicht: (2023)
von: Zou, Andy, et al.
Veröffentlicht: (2023)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
Aggressive Compression Enables LLM Weight Theft
von: Brown, Davis, et al.
Veröffentlicht: (2026)
von: Brown, Davis, et al.
Veröffentlicht: (2026)
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
Reasons to Doubt the Impact of AI Risk Evaluations
von: Mukobi, Gabriel
Veröffentlicht: (2024)
von: Mukobi, Gabriel
Veröffentlicht: (2024)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
von: Wang, Clinton J., et al.
Veröffentlicht: (2025)
Remote Labor Index: Measuring AI Automation of Remote Work
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025)
Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics
von: Romero, Peter, et al.
Veröffentlicht: (2024)
von: Romero, Peter, et al.
Veröffentlicht: (2024)
How Should AI Safety Benchmarks Benchmark Safety?
von: Yu, Cheng, et al.
Veröffentlicht: (2026)
von: Yu, Cheng, et al.
Veröffentlicht: (2026)
Testing the Machine Consciousness Hypothesis
von: Fitz, Stephen
Veröffentlicht: (2025)
von: Fitz, Stephen
Veröffentlicht: (2025)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
von: Meek, Austin, et al.
Veröffentlicht: (2025)
von: Meek, Austin, et al.
Veröffentlicht: (2025)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2023)
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2023)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
Superintelligence Strategy: Expert Version
von: Hendrycks, Dan, et al.
Veröffentlicht: (2025)
von: Hendrycks, Dan, et al.
Veröffentlicht: (2025)
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
von: Götting, Jasper, et al.
Veröffentlicht: (2025)
von: Götting, Jasper, et al.
Veröffentlicht: (2025)
Safety First: Psychological Safety as the Key to AI Transformation
von: Reich, Aaron, et al.
Veröffentlicht: (2026)
von: Reich, Aaron, et al.
Veröffentlicht: (2026)
When Names Disappear: Revealing What LLMs Actually Understand About Code
von: Le, Cuong Chi, et al.
Veröffentlicht: (2025)
von: Le, Cuong Chi, et al.
Veröffentlicht: (2025)
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiangyang, et al.
Veröffentlicht: (2025)
Generating Robot Constitutions & Benchmarks for Semantic Safety
von: Sermanet, Pierre, et al.
Veröffentlicht: (2025)
von: Sermanet, Pierre, et al.
Veröffentlicht: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
von: Hellrigel-Holderbaum, Max, et al.
Veröffentlicht: (2026)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
von: Zhao, Haochen, et al.
Veröffentlicht: (2024)
von: Zhao, Haochen, et al.
Veröffentlicht: (2024)
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2026)
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2026)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
von: Lee, Chia-Hsuan, et al.
Veröffentlicht: (2026)
von: Lee, Chia-Hsuan, et al.
Veröffentlicht: (2026)
Benchmarking and Understanding Safety Risks in AI Character Platforms
von: Wei, Yiluo, et al.
Veröffentlicht: (2025)
von: Wei, Yiluo, et al.
Veröffentlicht: (2025)
International Scientific Report on the Safety of Advanced AI (Interim Report)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2024)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
von: Shi, Enyi, et al.
Veröffentlicht: (2026)
von: Shi, Enyi, et al.
Veröffentlicht: (2026)
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
von: Liu, Xin, et al.
Veröffentlicht: (2023)
von: Liu, Xin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
TextQuests: How Good are LLMs at Text-Based Video Games?
von: Phan, Long, et al.
Veröffentlicht: (2025) -
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
von: Mazeika, Mantas, et al.
Veröffentlicht: (2025) -
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024) -
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
von: Ren, Richard, et al.
Veröffentlicht: (2025) -
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
von: Fitz, Stephen, et al.
Veröffentlicht: (2025)