Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Arrieta, Aitor, Ugarte, Miriam, Valle, Pablo, Parejo, José Antonio, Segura, Sergio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
o3-mini vs DeepSeek-R1: Which One is Safer?
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
ASTRAL: Automated Safety Testing of Large Language Models
by: Ugarte, Miriam, et al.
Published: (2025)
by: Ugarte, Miriam, et al.
Published: (2025)
Metamorphic Testing of Vision-Language Action-Enabled Robots
by: Valle, Pablo, et al.
Published: (2026)
by: Valle, Pablo, et al.
Published: (2026)
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
by: Romero-Arjona, Miguel, et al.
Published: (2025)
by: Romero-Arjona, Miguel, et al.
Published: (2025)
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
by: Romero-Arjona, Miguel, et al.
Published: (2025)
by: Romero-Arjona, Miguel, et al.
Published: (2025)
An Empirical Evaluation of White-box and Black-box Test Case Prioritization Techniques in CPSs Modeled in Simulink
by: Arrieta, Aitor
Published: (2025)
by: Arrieta, Aitor
Published: (2025)
Search-based Automated Program Repair of CPS Controllers Modeled in Simulink-Stateflow
by: Arrieta, Aitor, et al.
Published: (2024)
by: Arrieta, Aitor, et al.
Published: (2024)
Exploring the Potential of Large Language Models in Simulink-Stateflow Mutant Generation
by: Valle, Pablo, et al.
Published: (2026)
by: Valle, Pablo, et al.
Published: (2026)
Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
by: Valle, Pablo, et al.
Published: (2025)
by: Valle, Pablo, et al.
Published: (2025)
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots
by: Saurabh, Prasun, et al.
Published: (2026)
by: Saurabh, Prasun, et al.
Published: (2026)
Evaluating the Effectiveness of OpenAI's Parental Control System
by: Ersoz, Kerem, et al.
Published: (2026)
by: Ersoz, Kerem, et al.
Published: (2026)
OpenAI for OpenAPI: Automated generation of REST API specification via LLMs
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
MarMot: Metamorphic Runtime Monitoring of Autonomous Driving Systems
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
Voices from the Frontier: A Comprehensive Analysis of the OpenAI Developer Forum
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
An Empirical Study of OpenAI API Discussions on Stack Overflow
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
by: Lu, Chengjie, et al.
Published: (2025)
by: Lu, Chengjie, et al.
Published: (2025)
Foundation Models for the Digital Twin Creation of Cyber-Physical Systems
by: Ali, Shaukat, et al.
Published: (2024)
by: Ali, Shaukat, et al.
Published: (2024)
Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
by: Wu, Jiahui, et al.
Published: (2025)
by: Wu, Jiahui, et al.
Published: (2025)
Reinforcement Learning for Testing Interdependent Requirements in Autonomous Vehicles: An Empirical Study
by: Wu, Jiahui, et al.
Published: (2025)
by: Wu, Jiahui, et al.
Published: (2025)
Search-based Generation of Waypoints for Triggering Self-Adaptations in Maritime Autonomous Vessels
by: Nylænder, Karoline, et al.
Published: (2025)
by: Nylænder, Karoline, et al.
Published: (2025)
A Tool for Benchmarking Large Language Models' Robustness in Assessing the Realism of Driving Scenarios
by: Wu, Jiahui, et al.
Published: (2025)
by: Wu, Jiahui, et al.
Published: (2025)
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
Pricing4SaaS: a suite of software libraries for pricing-driven feature toggling
by: García-Fernández, Alejandro, et al.
Published: (2024)
by: García-Fernández, Alejandro, et al.
Published: (2024)
Automated Analysis of Pricings in SaaS-based Information Systems
by: García-Fernández, Alejandro, et al.
Published: (2025)
by: García-Fernández, Alejandro, et al.
Published: (2025)
Reality Bites: Assessing the Realism of Driving Scenarios with Large Language Models
by: Wu, Jiahui, et al.
Published: (2024)
by: Wu, Jiahui, et al.
Published: (2024)
GenMorph: Automatically Generating Metamorphic Relations via Genetic Programming
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
Application of Quantum Extreme Learning Machines for QoS Prediction of Elevators' Software in an Industrial Context
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
SATORI: Static Test Oracle Generation for REST APIs
by: Alonso, Juan C., et al.
Published: (2025)
by: Alonso, Juan C., et al.
Published: (2025)
Assessing Quantum Extreme Learning Machines for Software Testing in Practice
by: Muqeet, Asmar, et al.
Published: (2024)
by: Muqeet, Asmar, et al.
Published: (2024)
Pricing4SaaS: Towards a pricing model to drive the operation of SaaS
by: García-Fernández, Alejandro, et al.
Published: (2024)
by: García-Fernández, Alejandro, et al.
Published: (2024)
Pricing-driven Development and Operation of SaaS : Challenges and Opportunities
by: García-Fernández, Alejandro, et al.
Published: (2024)
by: García-Fernández, Alejandro, et al.
Published: (2024)
HORIZON: a Classification and Comparison Framework for Pricing-driven Feature Toggling
by: García-Fernández, Alejandro, et al.
Published: (2025)
by: García-Fernández, Alejandro, et al.
Published: (2025)
Assessing Vision-Language Models for Perception in Autonomous Underwater Robotic Software
by: Yousaf, Muhammad, et al.
Published: (2026)
by: Yousaf, Muhammad, et al.
Published: (2026)
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
by: Irugalbandara, Chandra, et al.
Published: (2023)
by: Irugalbandara, Chandra, et al.
Published: (2023)
Automated Repair of Cyber-Physical Systems
by: Valle, Pablo
Published: (2025)
by: Valle, Pablo
Published: (2025)
Racing the Market: An Industry Support Analysis for Pricing-Driven DevOps in SaaS
by: Garcia-Fernández, Alejandro, et al.
Published: (2024)
by: Garcia-Fernández, Alejandro, et al.
Published: (2024)
Using Large Language Models to Develop Requirements Elicitation Skills
by: Lojo, Nelson, et al.
Published: (2025)
by: Lojo, Nelson, et al.
Published: (2025)
Agentic AI in Industry: Adoption Level and Deployment Barriers
by: Apostolou, Spyridon Alvanakis, et al.
Published: (2026)
by: Apostolou, Spyridon Alvanakis, et al.
Published: (2026)
HerAgent: Rethinking the Automated Environment Deployment via Hierarchical Test Pyramid
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Quantum software experiments: A reporting and laboratory package structure guidelines
by: Moguel, Enrique, et al.
Published: (2024)
by: Moguel, Enrique, et al.
Published: (2024)
Similar Items
-
o3-mini vs DeepSeek-R1: Which One is Safer?
by: Arrieta, Aitor, et al.
Published: (2025) -
ASTRAL: Automated Safety Testing of Large Language Models
by: Ugarte, Miriam, et al.
Published: (2025) -
Metamorphic Testing of Vision-Language Action-Enabled Robots
by: Valle, Pablo, et al.
Published: (2026) -
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
by: Romero-Arjona, Miguel, et al.
Published: (2025) -
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
by: Romero-Arjona, Miguel, et al.
Published: (2025)