Themis: Automatic and Efficient Deep Learning System Testing with Strong Fault Detection Capability
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Dong, Li, Tsz On, Xie, Xiaofei, Cui, Heming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatically Learning a Precise Measurement for Fault Diagnosis Capability of Test Cases
by: Zhao, Yifan, et al.
Published: (2025)
by: Zhao, Yifan, et al.
Published: (2025)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
Measuring the Influence of Incorrect Code on Test Generation
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
MoDitector: Module-Directed Testing for Autonomous Driving Systems
by: Wang, Renzhi, et al.
Published: (2025)
by: Wang, Renzhi, et al.
Published: (2025)
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Testing for Fault Diversity in Reinforcement Learning
by: Mazouni, Quentin, et al.
Published: (2024)
by: Mazouni, Quentin, et al.
Published: (2024)
Evaluation and Improvement of Fault Detection for Large Language Models
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
muPRL: A Mutation Testing Pipeline for Deep Reinforcement Learning based on Real Faults
by: Thomas, Deepak-George, et al.
Published: (2024)
by: Thomas, Deepak-George, et al.
Published: (2024)
Efficient Black-Box Fault Localization for System-Level Test Code Using Large Language Models
by: Yaraghi, Ahmadreza Saboor, et al.
Published: (2025)
by: Yaraghi, Ahmadreza Saboor, et al.
Published: (2025)
Subgraph-Oriented Testing for Deep Learning Libraries
by: Xie, Xiaoyuan, et al.
Published: (2024)
by: Xie, Xiaoyuan, et al.
Published: (2024)
Fuzzing Automatic Differentiation in Deep-Learning Libraries
by: Yang, Chenyuan, et al.
Published: (2023)
by: Yang, Chenyuan, et al.
Published: (2023)
Effective Random Test Generation for Deep Learning Compilers
by: Ren, Luyao, et al.
Published: (2023)
by: Ren, Luyao, et al.
Published: (2023)
Fault-Tolerant Design and Multi-Objective Model Checking for Real-Time Deep Reinforcement Learning Systems
by: Su, Guoxin, et al.
Published: (2026)
by: Su, Guoxin, et al.
Published: (2026)
Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ Bugs
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
LLMCFG-TGen: Using LLM-Generated Control Flow Graphs to Automatically Create Test Cases from Use Cases
by: Yang, Zhenzhen, et al.
Published: (2025)
by: Yang, Zhenzhen, et al.
Published: (2025)
Boundary Value Test Input Generation Using Prompt Engineering with LLMs: Fault Detection and Coverage Analysis
by: Guo, Xiujing, et al.
Published: (2025)
by: Guo, Xiujing, et al.
Published: (2025)
XAMT: Cross-Framework API Matching for Testing Deep Learning Libraries
by: Duan, Bin, et al.
Published: (2025)
by: Duan, Bin, et al.
Published: (2025)
Many-Objective Search-Based Coverage-Guided Automatic Test Generation for Deep Neural Networks
by: Li, Dongcheng, et al.
Published: (2024)
by: Li, Dongcheng, et al.
Published: (2024)
A Comprehensive Study on Static Application Security Testing (SAST) Tools for Android
by: Zhu, Jingyun, et al.
Published: (2024)
by: Zhu, Jingyun, et al.
Published: (2024)
Software Fault Localization Based on Multi-objective Feature Fusion and Deep Learning
by: Hu, Xiaolei, et al.
Published: (2024)
by: Hu, Xiaolei, et al.
Published: (2024)
SBEST: Spectrum-Based Fault Localization Without Fault-Triggering Tests
by: Rafi, Md Nakhla, et al.
Published: (2024)
by: Rafi, Md Nakhla, et al.
Published: (2024)
From Exploration to Specification: LLM-Based Property Generation for Mobile App Testing
by: Xiong, Yiheng, et al.
Published: (2026)
by: Xiong, Yiheng, et al.
Published: (2026)
GenFair: Systematic Test Generation for Fairness Fault Detection in Large Language Models
by: Srinivasan, Madhusudan, et al.
Published: (2025)
by: Srinivasan, Madhusudan, et al.
Published: (2025)
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
by: Paul, Indraneil, et al.
Published: (2026)
by: Paul, Indraneil, et al.
Published: (2026)
Leveraging LLM Agents for Automated Video Game Testing
by: Wang, Chengjia, et al.
Published: (2025)
by: Wang, Chengjia, et al.
Published: (2025)
Introducing Ensemble Machine Learning Algorithms for Automatic Test Case Generation using Learning Based Testing
by: Rahman, Sheikh Md. Mushfiqur, et al.
Published: (2024)
by: Rahman, Sheikh Md. Mushfiqur, et al.
Published: (2024)
DriveTester: A Unified Platform for Simulation-Based Autonomous Driving Testing
by: Cheng, Mingfei, et al.
Published: (2024)
by: Cheng, Mingfei, et al.
Published: (2024)
STCLocker: Deadlock Avoidance Testing for Autonomous Driving Systems
by: Cheng, Mingfei, et al.
Published: (2025)
by: Cheng, Mingfei, et al.
Published: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
DEFT: Differentiable Automatic Test Pattern Generation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Automatically Removing Unnecessary Stubbings from Test Suites
by: Li, Mengzhen, et al.
Published: (2024)
by: Li, Mengzhen, et al.
Published: (2024)
Real-World Fault Detection for C-Extended Python Projects with Automated Unit Test Generation
by: Berg, Lucas, et al.
Published: (2026)
by: Berg, Lucas, et al.
Published: (2026)
Improved Detection and Diagnosis of Faults in Deep Neural Networks Using Hierarchical and Explainable Classification
by: Jahan, Sigma, et al.
Published: (2025)
by: Jahan, Sigma, et al.
Published: (2025)
Hybrid Fault-Driven Mutation Testing for Python
by: Alimadadi, Saba, et al.
Published: (2026)
by: Alimadadi, Saba, et al.
Published: (2026)
Optimization-Aware Test Generation for Deep Learning Compilers
by: Shen, Qingchao, et al.
Published: (2025)
by: Shen, Qingchao, et al.
Published: (2025)
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing
by: Huang, Kai, et al.
Published: (2025)
by: Huang, Kai, et al.
Published: (2025)
Coverage-Guided Testing for Deep Learning Models: A Comprehensive Survey
by: Guo, Hongjing, et al.
Published: (2025)
by: Guo, Hongjing, et al.
Published: (2025)
Fault Localization in Deep Learning-based Software: A System-level Approach
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
Similar Items
-
Automatically Learning a Precise Measurement for Fault Diagnosis Capability of Test Cases
by: Zhao, Yifan, et al.
Published: (2025) -
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023) -
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024) -
Measuring the Influence of Incorrect Code on Test Generation
by: Huang, Dong, et al.
Published: (2024) -
MoDitector: Module-Directed Testing for Autonomous Driving Systems
by: Wang, Renzhi, et al.
Published: (2025)