AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
Fuente:
arXiv
Salvato in:
| Autori principali: | Zeng, Yi, Yang, Yu, Zhou, Andy, Tan, Jeffrey Ziwei, Tu, Yuheng, Mai, Yifan, Klyman, Kevin, Pan, Minzhou, Jia, Ruoxi, Song, Dawn, Liang, Percy, Li, Bo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies
di: Zeng, Yi, et al.
Pubblicazione: (2024)
di: Zeng, Yi, et al.
Pubblicazione: (2024)
Acceptable Use Policies for Foundation Models
di: Klyman, Kevin
Pubblicazione: (2024)
di: Klyman, Kevin
Pubblicazione: (2024)
Language model developers should report train-test overlap
di: Zhang, Andy K, et al.
Pubblicazione: (2024)
di: Zhang, Andy K, et al.
Pubblicazione: (2024)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
di: Li, Xu, et al.
Pubblicazione: (2026)
di: Li, Xu, et al.
Pubblicazione: (2026)
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
SpecEval: Evaluating Model Adherence to Behavior Specifications
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating Disorders
di: Winecoff, Amy, et al.
Pubblicazione: (2025)
di: Winecoff, Amy, et al.
Pubblicazione: (2025)
The 2024 Foundation Model Transparency Index
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
di: Chen, Jianlyu, et al.
Pubblicazione: (2024)
di: Chen, Jianlyu, et al.
Pubblicazione: (2024)
JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
di: Pan, Minzhou, et al.
Pubblicazione: (2024)
di: Pan, Minzhou, et al.
Pubblicazione: (2024)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation
di: Alanoca, Sacha, et al.
Pubblicazione: (2025)
di: Alanoca, Sacha, et al.
Pubblicazione: (2025)
Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset
di: Kang, Mintong, et al.
Pubblicazione: (2025)
di: Kang, Mintong, et al.
Pubblicazione: (2025)
Point Objects and Derived Equivalences of Twisted Derived Categories of Abelian Varieties
di: Li, Ruoxi
Pubblicazione: (2025)
di: Li, Ruoxi
Pubblicazione: (2025)
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study
di: Wang, Chenguang, et al.
Pubblicazione: (2024)
di: Wang, Chenguang, et al.
Pubblicazione: (2024)
BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
di: Zeng, Yi, et al.
Pubblicazione: (2024)
di: Zeng, Yi, et al.
Pubblicazione: (2024)
User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
di: King, Jennifer, et al.
Pubblicazione: (2025)
di: King, Jennifer, et al.
Pubblicazione: (2025)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
di: Xu, Peiyang, et al.
Pubblicazione: (2025)
di: Xu, Peiyang, et al.
Pubblicazione: (2025)
AutoBencher: Towards Declarative Benchmark Construction
di: Li, Xiang Lisa, et al.
Pubblicazione: (2024)
di: Li, Xiang Lisa, et al.
Pubblicazione: (2024)
Reliable and Efficient Amortized Model-based Evaluation
di: Truong, Sang, et al.
Pubblicazione: (2025)
di: Truong, Sang, et al.
Pubblicazione: (2025)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
di: Cobben, Pepijn, et al.
Pubblicazione: (2026)
di: Cobben, Pepijn, et al.
Pubblicazione: (2026)
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
di: Roberts, Josselin Somerville, et al.
Pubblicazione: (2024)
di: Roberts, Josselin Somerville, et al.
Pubblicazione: (2024)
Foundation Model Transparency Reports
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
di: Bommasani, Rishi, et al.
Pubblicazione: (2024)
The 2025 Foundation Model Transparency Index
di: Wan, Alexander, et al.
Pubblicazione: (2025)
di: Wan, Alexander, et al.
Pubblicazione: (2025)
AIR: Improving Agent Safety through Incident Response
di: Xiao, Zibo, et al.
Pubblicazione: (2026)
di: Xiao, Zibo, et al.
Pubblicazione: (2026)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
di: Zhao, Haochen, et al.
Pubblicazione: (2024)
di: Zhao, Haochen, et al.
Pubblicazione: (2024)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)
di: Zhao, Yunhan, et al.
Pubblicazione: (2026)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
di: Pan, Leyi, et al.
Pubblicazione: (2025)
di: Pan, Leyi, et al.
Pubblicazione: (2025)
Do AI Companies Make Good on Voluntary Commitments to the White House?
di: Wang, Jennifer, et al.
Pubblicazione: (2025)
di: Wang, Jennifer, et al.
Pubblicazione: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
di: Zhou, Yujun, et al.
Pubblicazione: (2024)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
di: Liu, Xuannan, et al.
Pubblicazione: (2025)
The Limits of AI Data Transparency Policy: Three Disclosure Fallacies
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2026)
di: Shen, Judy Hanwen, et al.
Pubblicazione: (2026)
Evaluating and Mitigating IP Infringement in Visual Generative AI
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
NAS-Bench-Graph: Benchmarking Graph Neural Architecture Search
di: Qin, Yijian, et al.
Pubblicazione: (2022)
di: Qin, Yijian, et al.
Pubblicazione: (2022)
PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
di: Yang, Zhiwen, et al.
Pubblicazione: (2026)
di: Yang, Zhiwen, et al.
Pubblicazione: (2026)
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
di: McGregor, Sean, et al.
Pubblicazione: (2025)
di: McGregor, Sean, et al.
Pubblicazione: (2025)
RiskBench: A Scenario-based Benchmark for Risk Identification
di: Kung, Chi-Hsi, et al.
Pubblicazione: (2023)
di: Kung, Chi-Hsi, et al.
Pubblicazione: (2023)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
di: Xie, Tinghao, et al.
Pubblicazione: (2024)
di: Xie, Tinghao, et al.
Pubblicazione: (2024)
AI Risk Management Should Incorporate Both Safety and Security
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
di: Qi, Xiangyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies
di: Zeng, Yi, et al.
Pubblicazione: (2024) -
Acceptable Use Policies for Foundation Models
di: Klyman, Kevin
Pubblicazione: (2024) -
Language model developers should report train-test overlap
di: Zhang, Andy K, et al.
Pubblicazione: (2024) -
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
di: Li, Xu, et al.
Pubblicazione: (2026) -
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)