A Tool for Benchmarking Large Language Models' Robustness in Assessing the Realism of Driving Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jiahui, Lu, Chengjie, Arrieta, Aitor, Ali, Shaukat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reality Bites: Assessing the Realism of Driving Scenarios with Large Language Models
von: Wu, Jiahui, et al.
Veröffentlicht: (2024)
von: Wu, Jiahui, et al.
Veröffentlicht: (2024)
Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
von: Wu, Jiahui, et al.
Veröffentlicht: (2025)
von: Wu, Jiahui, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Testing Interdependent Requirements in Autonomous Vehicles: An Empirical Study
von: Wu, Jiahui, et al.
Veröffentlicht: (2025)
von: Wu, Jiahui, et al.
Veröffentlicht: (2025)
Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
von: Valle, Pablo, et al.
Veröffentlicht: (2025)
von: Valle, Pablo, et al.
Veröffentlicht: (2025)
Exploring the Potential of Large Language Models in Simulink-Stateflow Mutant Generation
von: Valle, Pablo, et al.
Veröffentlicht: (2026)
von: Valle, Pablo, et al.
Veröffentlicht: (2026)
Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
von: Lu, Chengjie, et al.
Veröffentlicht: (2025)
von: Lu, Chengjie, et al.
Veröffentlicht: (2025)
Assessing Vision-Language Models for Perception in Autonomous Underwater Robotic Software
von: Yousaf, Muhammad, et al.
Veröffentlicht: (2026)
von: Yousaf, Muhammad, et al.
Veröffentlicht: (2026)
Foundation Models for the Digital Twin Creation of Cyber-Physical Systems
von: Ali, Shaukat, et al.
Veröffentlicht: (2024)
von: Ali, Shaukat, et al.
Veröffentlicht: (2024)
Search-based Automated Program Repair of CPS Controllers Modeled in Simulink-Stateflow
von: Arrieta, Aitor, et al.
Veröffentlicht: (2024)
von: Arrieta, Aitor, et al.
Veröffentlicht: (2024)
Metamorphic Testing of Vision-Language Action-Enabled Robots
von: Valle, Pablo, et al.
Veröffentlicht: (2026)
von: Valle, Pablo, et al.
Veröffentlicht: (2026)
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots
von: Saurabh, Prasun, et al.
Veröffentlicht: (2026)
von: Saurabh, Prasun, et al.
Veröffentlicht: (2026)
Search-based Generation of Waypoints for Triggering Self-Adaptations in Maritime Autonomous Vessels
von: Nylænder, Karoline, et al.
Veröffentlicht: (2025)
von: Nylænder, Karoline, et al.
Veröffentlicht: (2025)
Application of Quantum Extreme Learning Machines for QoS Prediction of Elevators' Software in an Industrial Context
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
An Empirical Evaluation of White-box and Black-box Test Case Prioritization Techniques in CPSs Modeled in Simulink
von: Arrieta, Aitor
Veröffentlicht: (2025)
von: Arrieta, Aitor
Veröffentlicht: (2025)
Assessing Quantum Extreme Learning Machines for Software Testing in Practice
von: Muqeet, Asmar, et al.
Veröffentlicht: (2024)
von: Muqeet, Asmar, et al.
Veröffentlicht: (2024)
Search-based Selection of Metamorphic Relations for Optimized Robustness Testing of Large Language Models
von: Hyun, Sangwon, et al.
Veröffentlicht: (2025)
von: Hyun, Sangwon, et al.
Veröffentlicht: (2025)
ASTRAL: Automated Safety Testing of Large Language Models
von: Ugarte, Miriam, et al.
Veröffentlicht: (2025)
von: Ugarte, Miriam, et al.
Veröffentlicht: (2025)
A Tool for Test Case Scenarios Generation Using Large Language Models
von: Sami, Abdul Malik, et al.
Veröffentlicht: (2024)
von: Sami, Abdul Malik, et al.
Veröffentlicht: (2024)
MarMot: Metamorphic Runtime Monitoring of Autonomous Driving Systems
von: Ayerdi, Jon, et al.
Veröffentlicht: (2023)
von: Ayerdi, Jon, et al.
Veröffentlicht: (2023)
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
von: Romero-Arjona, Miguel, et al.
Veröffentlicht: (2025)
von: Romero-Arjona, Miguel, et al.
Veröffentlicht: (2025)
UAMTERS: Uncertainty-Aware Mutation Analysis for DL-enabled Robotic Software
von: Lu, Chengjie, et al.
Veröffentlicht: (2026)
von: Lu, Chengjie, et al.
Veröffentlicht: (2026)
A Survey on the Application of Large Language Models in Scenario-Based Testing of Automated Driving Systems
von: Zhao, Yongqi, et al.
Veröffentlicht: (2025)
von: Zhao, Yongqi, et al.
Veröffentlicht: (2025)
Identifying Uncertainty in Self-Adaptive Robotics with Large Language Models
von: Sartaj, Hassan, et al.
Veröffentlicht: (2025)
von: Sartaj, Hassan, et al.
Veröffentlicht: (2025)
Modeling Language for Scenario Development of Autonomous Driving Systems
von: Aoki, Toshiaki, et al.
Veröffentlicht: (2025)
von: Aoki, Toshiaki, et al.
Veröffentlicht: (2025)
A Large Scale Empirical Analysis on the Adherence Gap between Standards and Tools in SBOM
von: Wang, Chengjie, et al.
Veröffentlicht: (2026)
von: Wang, Chengjie, et al.
Veröffentlicht: (2026)
LeGEND: A Top-Down Approach to Scenario Generation of Autonomous Driving Systems Assisted by Large Language Models
von: Tang, Shuncheng, et al.
Veröffentlicht: (2024)
von: Tang, Shuncheng, et al.
Veröffentlicht: (2024)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
von: Hu, Xing, et al.
Veröffentlicht: (2025)
von: Hu, Xing, et al.
Veröffentlicht: (2025)
Assessing Large Language Models for Stabilizing Numerical Expressions in Scientific Software
von: Nguyen, Tien, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien, et al.
Veröffentlicht: (2026)
A Software Engineering Perspective on Testing Large Language Models: Research, Practice, Tools and Benchmarks
von: Hudson, Sinclair, et al.
Veröffentlicht: (2024)
von: Hudson, Sinclair, et al.
Veröffentlicht: (2024)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
Large Language Models Versus Static Code Analysis Tools: A Systematic Benchmark for Vulnerability Detection
von: Gnieciak, Damian, et al.
Veröffentlicht: (2025)
von: Gnieciak, Damian, et al.
Veröffentlicht: (2025)
Robust Mutation Analysis of Quantum Programs Under Noise
von: Fortz, Sophie, et al.
Veröffentlicht: (2026)
von: Fortz, Sophie, et al.
Veröffentlicht: (2026)
Model-based Digital Twins of Medicine Dispensers for Healthcare IoT Applications
von: Sartaj, Hassan, et al.
Veröffentlicht: (2023)
von: Sartaj, Hassan, et al.
Veröffentlicht: (2023)
Quantum Artificial Intelligence for Software Engineering: the Road Ahead
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Evaluating Search-Based Software Microbenchmark Prioritization
von: Laaber, Christoph, et al.
Veröffentlicht: (2022)
von: Laaber, Christoph, et al.
Veröffentlicht: (2022)
Quantum Program Testing Through Commuting Pauli Strings on IBM's Quantum Computers
von: Muqeet, Asmar, et al.
Veröffentlicht: (2024)
von: Muqeet, Asmar, et al.
Veröffentlicht: (2024)
Envisioning Responsible Quantum Software Engineering and Quantum Artificial Intelligence
von: Bano, Muneera, et al.
Veröffentlicht: (2024)
von: Bano, Muneera, et al.
Veröffentlicht: (2024)
Search-Based Quantum Program Testing via Commuting Pauli String
von: Muqeet, Asmar, et al.
Veröffentlicht: (2026)
von: Muqeet, Asmar, et al.
Veröffentlicht: (2026)
REST API Testing in DevOps: A Study on an Evolving Healthcare IoT Application
von: Sartaj, Hassan, et al.
Veröffentlicht: (2024)
von: Sartaj, Hassan, et al.
Veröffentlicht: (2024)
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reality Bites: Assessing the Realism of Driving Scenarios with Large Language Models
von: Wu, Jiahui, et al.
Veröffentlicht: (2024) -
Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
von: Wu, Jiahui, et al.
Veröffentlicht: (2025) -
Reinforcement Learning for Testing Interdependent Requirements in Autonomous Vehicles: An Empirical Study
von: Wu, Jiahui, et al.
Veröffentlicht: (2025) -
Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
von: Valle, Pablo, et al.
Veröffentlicht: (2025) -
Exploring the Potential of Large Language Models in Simulink-Stateflow Mutant Generation
von: Valle, Pablo, et al.
Veröffentlicht: (2026)