Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arrieta, Aitor, Ugarte, Miriam, Valle, Pablo, Parejo, José Antonio, Segura, Sergio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912209457643520
author Arrieta, Aitor
Ugarte, Miriam
Valle, Pablo
Parejo, José Antonio
Segura, Sergio
author_facet Arrieta, Aitor
Ugarte, Miriam
Valle, Pablo
Parejo, José Antonio
Segura, Sergio
contents Large Language Models (LLMs) have become an integral part of our daily lives. However, they impose certain risks, including those that can harm individuals' privacy, perpetuate biases and spread misinformation. These risks highlight the need for robust safety mechanisms, ethical guidelines, and thorough testing to ensure their responsible deployment. Safety of LLMs is a key property that needs to be thoroughly tested prior the model to be deployed and accessible to the general users. This paper reports the external safety testing experience conducted by researchers from Mondragon University and University of Seville on OpenAI's new o3-mini LLM as part of OpenAI's early access for safety testing program. In particular, we apply our tool, ASTRAL, to automatically and systematically generate up to date unsafe test inputs (i.e., prompts) that helps us test and assess different safety categories of LLMs. We automatically generate and execute a total of 10,080 unsafe test input on a early o3-mini beta version. After manually verifying the test cases classified as unsafe by ASTRAL, we identify a total of 87 actual instances of unsafe LLM behavior. We highlight key insights and findings uncovered during the pre-deployment external testing phase of OpenAI's latest LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
Arrieta, Aitor
Ugarte, Miriam
Valle, Pablo
Parejo, José Antonio
Segura, Sergio
Software Engineering
Artificial Intelligence
Large Language Models (LLMs) have become an integral part of our daily lives. However, they impose certain risks, including those that can harm individuals' privacy, perpetuate biases and spread misinformation. These risks highlight the need for robust safety mechanisms, ethical guidelines, and thorough testing to ensure their responsible deployment. Safety of LLMs is a key property that needs to be thoroughly tested prior the model to be deployed and accessible to the general users. This paper reports the external safety testing experience conducted by researchers from Mondragon University and University of Seville on OpenAI's new o3-mini LLM as part of OpenAI's early access for safety testing program. In particular, we apply our tool, ASTRAL, to automatically and systematically generate up to date unsafe test inputs (i.e., prompts) that helps us test and assess different safety categories of LLMs. We automatically generate and execute a total of 10,080 unsafe test input on a early o3-mini beta version. After manually verifying the test cases classified as unsafe by ASTRAL, we identify a total of 87 actual instances of unsafe LLM behavior. We highlight key insights and findings uncovered during the pre-deployment external testing phase of OpenAI's latest LLM.
title Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2501.17749