Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
Fuente:
arXiv
Saved in:
| Main Authors: | Jahangirova, Gunel, Humbatova, Nargiz, Kim, Jinhan, Yoo, Shin, Tonella, Paolo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024)
by: Humbatova, Nargiz, et al.
Published: (2024)
Fault Localisation and Repair for DL Systems: An Empirical Study with LLMs
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
MuFF: Stable and Sensitive Post-training Mutation Testing for Deep Learning
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
New Formulation of DNN Statistical Mutation Killing for Ensuring Monotonicity: A Technical Report
by: Kim, Jinhan, et al.
Published: (2025)
by: Kim, Jinhan, et al.
Published: (2025)
Revisiting "Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion": A Critical Review and Implications on DNN Coverage Testing
by: Kim, Jinhan, et al.
Published: (2026)
by: Kim, Jinhan, et al.
Published: (2026)
muPRL: A Mutation Testing Pipeline for Deep Reinforcement Learning based on Real Faults
by: Thomas, Deepak-George, et al.
Published: (2024)
by: Thomas, Deepak-George, et al.
Published: (2024)
TopoMap: A Feature-based Semantic Discriminator of the Topographical Regions in the Test Input Space
by: De Vita, Gianmarco, et al.
Published: (2025)
by: De Vita, Gianmarco, et al.
Published: (2025)
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026)
by: Bodicoat, Adam, et al.
Published: (2026)
A Taxonomy of Real Faults in Hybrid Quantum-Classical Architectures
by: Bensoussan, Avner, et al.
Published: (2025)
by: Bensoussan, Avner, et al.
Published: (2025)
Testing of Deep Reinforcement Learning Agents with Surrogate Models
by: Biagiola, Matteo, et al.
Published: (2023)
by: Biagiola, Matteo, et al.
Published: (2023)
Detecting Trojaned DNNs via Spectral Regression Analysis
by: Pasini, Samuele, et al.
Published: (2026)
by: Pasini, Samuele, et al.
Published: (2026)
Cross-site scripting adversarial attacks based on deep reinforcement learning: Evaluation and extension study
by: Pasini, Samuele, et al.
Published: (2025)
by: Pasini, Samuele, et al.
Published: (2025)
Reinforcement Learning for Online Testing of Autonomous Driving Systems: a Replication and Extension Study
by: Giamattei, Luca, et al.
Published: (2024)
by: Giamattei, Luca, et al.
Published: (2024)
GenMorph: Automatically Generating Metamorphic Relations via Genetic Programming
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
SETA: Statistical Fault Attribution for Compound AI Systems
by: Chowdhury, Sayak, et al.
Published: (2026)
by: Chowdhury, Sayak, et al.
Published: (2026)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
by: Rafi, Md Nakhla, et al.
Published: (2024)
by: Rafi, Md Nakhla, et al.
Published: (2024)
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
by: Haroon, Sabaat, et al.
Published: (2025)
by: Haroon, Sabaat, et al.
Published: (2025)
COSMosFL: Ensemble of Small Language Models for Fault Localisation
by: Cho, Hyunjoon, et al.
Published: (2025)
by: Cho, Hyunjoon, et al.
Published: (2025)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
by: Feng, Yunfei, et al.
Published: (2026)
by: Feng, Yunfei, et al.
Published: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
FetaFix: Automatic Fault Localization and Repair of Deep Learning Model Conversions
by: Louloudakis, Nikolaos, et al.
Published: (2023)
by: Louloudakis, Nikolaos, et al.
Published: (2023)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
DiTOX: Fault Detection and Localization in the ONNX Optimizer
by: Louloudakis, Nikolaos, et al.
Published: (2025)
by: Louloudakis, Nikolaos, et al.
Published: (2025)
High-Dimensional Fault Tolerance Testing of Highly Automated Vehicles Based on Low-Rank Models
by: Mei, Yuewen, et al.
Published: (2024)
by: Mei, Yuewen, et al.
Published: (2024)
Boundary State Generation for Testing and Improvement of Autonomous Driving Systems
by: Biagiola, Matteo, et al.
Published: (2023)
by: Biagiola, Matteo, et al.
Published: (2023)
Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs
by: Pasini, Samuele, et al.
Published: (2024)
by: Pasini, Samuele, et al.
Published: (2024)
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
by: Baresi, Luciano, et al.
Published: (2024)
by: Baresi, Luciano, et al.
Published: (2024)
DANDI: Diffusion as Normative Distribution for Deep Neural Network Input
by: Kim, Somin, et al.
Published: (2025)
by: Kim, Somin, et al.
Published: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
by: Butsko, Christina, et al.
Published: (2025)
by: Butsko, Christina, et al.
Published: (2025)
Investigating Reproducibility in Deep Learning-Based Software Fault Prediction
by: Mukhtar, Adil, et al.
Published: (2024)
by: Mukhtar, Adil, et al.
Published: (2024)
DeepKnowledge: Generalisation-Driven Deep Learning Testing
by: Missaoui, Sondess, et al.
Published: (2024)
by: Missaoui, Sondess, et al.
Published: (2024)
Enhancing Resilience and Scalability in Travel Booking Systems: A Microservices Approach to Fault Tolerance, Load Balancing, and Service Discovery
by: Barua, Biman, et al.
Published: (2024)
by: Barua, Biman, et al.
Published: (2024)
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
On the Replicability and Reproducibility of Deep Learning in Software Engineering
by: Liu, Chao, et al.
Published: (2020)
by: Liu, Chao, et al.
Published: (2020)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
by: Wu, Siyu, et al.
Published: (2026)
by: Wu, Siyu, et al.
Published: (2026)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
by: Bouchoucha, Rached, et al.
Published: (2024)
by: Bouchoucha, Rached, et al.
Published: (2024)
Fault Localization in Deep Learning-based Software: A System-level Approach
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
by: Siddiq, Mohammed Latif, et al.
Published: (2024)
Similar Items
-
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024) -
Fault Localisation and Repair for DL Systems: An Empirical Study with LLMs
by: Kim, Jinhan, et al.
Published: (2025) -
MuFF: Stable and Sensitive Post-training Mutation Testing for Deep Learning
by: Kim, Jinhan, et al.
Published: (2025) -
New Formulation of DNN Statistical Mutation Killing for Ensuring Monotonicity: A Technical Report
by: Kim, Jinhan, et al.
Published: (2025) -
Revisiting "Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion": A Critical Review and Implications on DNN Coverage Testing
by: Kim, Jinhan, et al.
Published: (2026)