o3-mini vs DeepSeek-R1: Which One is Safer?
Fuente:
arXiv
Saved in:
| Main Authors: | Arrieta, Aitor, Ugarte, Miriam, Valle, Pablo, Parejo, José Antonio, Segura, Sergio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
ASTRAL: Automated Safety Testing of Large Language Models
by: Ugarte, Miriam, et al.
Published: (2025)
by: Ugarte, Miriam, et al.
Published: (2025)
Metamorphic Testing of Vision-Language Action-Enabled Robots
by: Valle, Pablo, et al.
Published: (2026)
by: Valle, Pablo, et al.
Published: (2026)
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
by: Romero-Arjona, Miguel, et al.
Published: (2025)
by: Romero-Arjona, Miguel, et al.
Published: (2025)
Search-based Automated Program Repair of CPS Controllers Modeled in Simulink-Stateflow
by: Arrieta, Aitor, et al.
Published: (2024)
by: Arrieta, Aitor, et al.
Published: (2024)
Exploring the Potential of Large Language Models in Simulink-Stateflow Mutant Generation
by: Valle, Pablo, et al.
Published: (2026)
by: Valle, Pablo, et al.
Published: (2026)
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
by: Romero-Arjona, Miguel, et al.
Published: (2025)
by: Romero-Arjona, Miguel, et al.
Published: (2025)
An Empirical Evaluation of White-box and Black-box Test Case Prioritization Techniques in CPSs Modeled in Simulink
by: Arrieta, Aitor
Published: (2025)
by: Arrieta, Aitor
Published: (2025)
Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
by: Valle, Pablo, et al.
Published: (2025)
by: Valle, Pablo, et al.
Published: (2025)
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
by: Manik, Md Motaleb Hossen
Published: (2025)
by: Manik, Md Motaleb Hossen
Published: (2025)
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots
by: Saurabh, Prasun, et al.
Published: (2026)
by: Saurabh, Prasun, et al.
Published: (2026)
MarMot: Metamorphic Runtime Monitoring of Autonomous Driving Systems
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
C2SaferRust: Transforming C Projects into Safer Rust with NeuroSymbolic Techniques
by: Nitin, Vikram, et al.
Published: (2025)
by: Nitin, Vikram, et al.
Published: (2025)
Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
by: Lu, Chengjie, et al.
Published: (2025)
by: Lu, Chengjie, et al.
Published: (2025)
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
by: Sadik, Ahmed R., et al.
Published: (2025)
by: Sadik, Ahmed R., et al.
Published: (2025)
Foundation Models for the Digital Twin Creation of Cyber-Physical Systems
by: Ali, Shaukat, et al.
Published: (2024)
by: Ali, Shaukat, et al.
Published: (2024)
Pricing4SaaS: a suite of software libraries for pricing-driven feature toggling
by: García-Fernández, Alejandro, et al.
Published: (2024)
by: García-Fernández, Alejandro, et al.
Published: (2024)
Automated Analysis of Pricings in SaaS-based Information Systems
by: García-Fernández, Alejandro, et al.
Published: (2025)
by: García-Fernández, Alejandro, et al.
Published: (2025)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
by: Guo, Daya, et al.
Published: (2024)
by: Guo, Daya, et al.
Published: (2024)
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Search-based Generation of Waypoints for Triggering Self-Adaptations in Maritime Autonomous Vessels
by: Nylænder, Karoline, et al.
Published: (2025)
by: Nylænder, Karoline, et al.
Published: (2025)
A Tool for Benchmarking Large Language Models' Robustness in Assessing the Realism of Driving Scenarios
by: Wu, Jiahui, et al.
Published: (2025)
by: Wu, Jiahui, et al.
Published: (2025)
Pricing4SaaS: Towards a pricing model to drive the operation of SaaS
by: García-Fernández, Alejandro, et al.
Published: (2024)
by: García-Fernández, Alejandro, et al.
Published: (2024)
Pricing-driven Development and Operation of SaaS : Challenges and Opportunities
by: García-Fernández, Alejandro, et al.
Published: (2024)
by: García-Fernández, Alejandro, et al.
Published: (2024)
HORIZON: a Classification and Comparison Framework for Pricing-driven Feature Toggling
by: García-Fernández, Alejandro, et al.
Published: (2025)
by: García-Fernández, Alejandro, et al.
Published: (2025)
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Reality Bites: Assessing the Realism of Driving Scenarios with Large Language Models
by: Wu, Jiahui, et al.
Published: (2024)
by: Wu, Jiahui, et al.
Published: (2024)
GenMorph: Automatically Generating Metamorphic Relations via Genetic Programming
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
Application of Quantum Extreme Learning Machines for QoS Prediction of Elevators' Software in an Industrial Context
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Racing the Market: An Industry Support Analysis for Pricing-Driven DevOps in SaaS
by: Garcia-Fernández, Alejandro, et al.
Published: (2024)
by: Garcia-Fernández, Alejandro, et al.
Published: (2024)
Assessing Vision-Language Models for Perception in Autonomous Underwater Robotic Software
by: Yousaf, Muhammad, et al.
Published: (2026)
by: Yousaf, Muhammad, et al.
Published: (2026)
Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
by: Wu, Jiahui, et al.
Published: (2025)
by: Wu, Jiahui, et al.
Published: (2025)
Quantum software experiments: A reporting and laboratory package structure guidelines
by: Moguel, Enrique, et al.
Published: (2024)
by: Moguel, Enrique, et al.
Published: (2024)
Automated Repair of Cyber-Physical Systems
by: Valle, Pablo
Published: (2025)
by: Valle, Pablo
Published: (2025)
Using Large Language Models to Develop Requirements Elicitation Skills
by: Lojo, Nelson, et al.
Published: (2025)
by: Lojo, Nelson, et al.
Published: (2025)
Reinforcement Learning for Testing Interdependent Requirements in Autonomous Vehicles: An Empirical Study
by: Wu, Jiahui, et al.
Published: (2025)
by: Wu, Jiahui, et al.
Published: (2025)
Safer Builders, Risky Maintainers: A Comparative Study of Breaking Changes in Human vs Agentic PRs
by: Ferdous, K M, et al.
Published: (2026)
by: Ferdous, K M, et al.
Published: (2026)
Pricing-Driven Resource Allocation in the Computing Continuum
by: García-Fernández, Alejandro, et al.
Published: (2026)
by: García-Fernández, Alejandro, et al.
Published: (2026)
Towards a Transpiler for C/C++ to Safer Rust
by: Tripuramallu, Dhiren, et al.
Published: (2024)
by: Tripuramallu, Dhiren, et al.
Published: (2024)
SATORI: Static Test Oracle Generation for REST APIs
by: Alonso, Juan C., et al.
Published: (2025)
by: Alonso, Juan C., et al.
Published: (2025)
Similar Items
-
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
by: Arrieta, Aitor, et al.
Published: (2025) -
ASTRAL: Automated Safety Testing of Large Language Models
by: Ugarte, Miriam, et al.
Published: (2025) -
Metamorphic Testing of Vision-Language Action-Enabled Robots
by: Valle, Pablo, et al.
Published: (2026) -
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
by: Romero-Arjona, Miguel, et al.
Published: (2025) -
Search-based Automated Program Repair of CPS Controllers Modeled in Simulink-Stateflow
by: Arrieta, Aitor, et al.
Published: (2024)