Defining and Evaluating Physical Safety for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Yung-Chen, Chen, Pin-Yu, Ho, Tsung-Yi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917282406465536
author Tang, Yung-Chen
Chen, Pin-Yu
Ho, Tsung-Yi
author_facet Tang, Yung-Chen
Chen, Pin-Yu
Ho, Tsung-Yi
contents Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain unexplored. Our study addresses the critical gap in evaluating LLM physical safety by developing a comprehensive benchmark for drone control. We classify the physical safety risks of drones into four categories: (1) human-targeted threats, (2) object-targeted threats, (3) infrastructure attacks, and (4) regulatory violations. Our evaluation of mainstream LLMs reveals an undesirable trade-off between utility and safety, with models that excel in code generation often performing poorly in crucial safety aspects. Furthermore, while incorporating advanced prompt engineering techniques such as In-Context Learning and Chain-of-Thought can improve safety, these methods still struggle to identify unintentional attacks. In addition, larger models demonstrate better safety capabilities, particularly in refusing dangerous commands. Our findings and benchmark can facilitate the design and evaluation of physical safety for LLMs. The project page is available at huggingface.co/spaces/TrustSafeAI/LLM-physical-safety.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02317
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Defining and Evaluating Physical Safety for Large Language Models
Tang, Yung-Chen
Chen, Pin-Yu
Ho, Tsung-Yi
Machine Learning
Artificial Intelligence
Computers and Society
Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain unexplored. Our study addresses the critical gap in evaluating LLM physical safety by developing a comprehensive benchmark for drone control. We classify the physical safety risks of drones into four categories: (1) human-targeted threats, (2) object-targeted threats, (3) infrastructure attacks, and (4) regulatory violations. Our evaluation of mainstream LLMs reveals an undesirable trade-off between utility and safety, with models that excel in code generation often performing poorly in crucial safety aspects. Furthermore, while incorporating advanced prompt engineering techniques such as In-Context Learning and Chain-of-Thought can improve safety, these methods still struggle to identify unintentional attacks. In addition, larger models demonstrate better safety capabilities, particularly in refusing dangerous commands. Our findings and benchmark can facilitate the design and evaluation of physical safety for LLMs. The project page is available at huggingface.co/spaces/TrustSafeAI/LLM-physical-safety.
title Defining and Evaluating Physical Safety for Large Language Models
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2411.02317