Automatically Detecting Numerical Instability in Machine Learning Applications via Soft Assertions
Fuente:
arXiv
Saved in:
| Main Authors: | Sharmin, Shaila, Zahid, Anwar Hossain, Bhattacharjee, Subhankar, Igwilo, Chiamaka, Kim, Miryung, Le, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Testing SSD Firmware with State Data-Aware Fuzzing: Accelerating Coverage in Nondeterministic I/O Environments
by: Yoon, Gangho, et al.
Published: (2025)
by: Yoon, Gangho, et al.
Published: (2025)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
by: Yu, Boxi, et al.
Published: (2026)
by: Yu, Boxi, et al.
Published: (2026)
Combined Program Analysis Techniques: A Systematic Mapping Study
by: Braione, Pietro, et al.
Published: (2026)
by: Braione, Pietro, et al.
Published: (2026)
Validating Formal Specifications with LLM-generated Test Cases
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
Synthesizing Test Cases for Narrowing Specification Candidates
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
Comparing Human and LLM Generated Code: The Jury is Still Out!
by: Licorish, Sherlock A., et al.
Published: (2025)
by: Licorish, Sherlock A., et al.
Published: (2025)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
GBM Returns the Best Prediction Performance among Regression Approaches: A Case Study of Stack Overflow Code Quality
by: Licorish, Sherlock A., et al.
Published: (2025)
by: Licorish, Sherlock A., et al.
Published: (2025)
Evaluating Cryptographic API Misuse Detectors for Go
by: Andersson, Vivi, et al.
Published: (2026)
by: Andersson, Vivi, et al.
Published: (2026)
Tractable Verification of Model Transformations: A Cutoff-Theorem Approach for DSLTrans
by: Lucio, Levi
Published: (2026)
by: Lucio, Levi
Published: (2026)
PALM: Path-aware LLM-based Test Generation with Comprehension
by: Wu, Yaoxuan, et al.
Published: (2025)
by: Wu, Yaoxuan, et al.
Published: (2025)
Automated Vulnerability Detection Using Deep Learning Technique
by: Yang, Guan-Yan, et al.
Published: (2024)
by: Yang, Guan-Yan, et al.
Published: (2024)
Trace Validation of Unmodified Concurrent Systems with OmniLink
by: Hackett, Finn, et al.
Published: (2026)
by: Hackett, Finn, et al.
Published: (2026)
Path-optimal symbolic execution of heap-manipulating programs
by: Braione, Pietro, et al.
Published: (2024)
by: Braione, Pietro, et al.
Published: (2024)
Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing
by: Liyanage, Danushka, et al.
Published: (2025)
by: Liyanage, Danushka, et al.
Published: (2025)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Understanding and Reusing Test Suites Across Database Systems
by: Zhong, Suyang, et al.
Published: (2024)
by: Zhong, Suyang, et al.
Published: (2024)
Monitoring Agentic Systems Before They're Reliable
by: Boston, Marisa Ferrara, et al.
Published: (2026)
by: Boston, Marisa Ferrara, et al.
Published: (2026)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
by: Parris, William M.
Published: (2026)
by: Parris, William M.
Published: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
by: Baldonado, Juan Manuel, et al.
Published: (2025)
by: Baldonado, Juan Manuel, et al.
Published: (2025)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
by: Calboreanu, Elias
Published: (2026)
by: Calboreanu, Elias
Published: (2026)
Validating Solidity Code Defects using Symbolic and Concrete Execution powered by Large Language Models
by: Susan, Ştefan-Claudiu, et al.
Published: (2025)
by: Susan, Ştefan-Claudiu, et al.
Published: (2025)
DEFault++: Automated Fault Detection, Categorization, and Diagnosis for Transformer Architectures
by: Jahan, Sigma, et al.
Published: (2026)
by: Jahan, Sigma, et al.
Published: (2026)
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
by: Salva, Sébastien, et al.
Published: (2025)
by: Salva, Sébastien, et al.
Published: (2025)
Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches
by: Wienczkowski, Michael
Published: (2026)
by: Wienczkowski, Michael
Published: (2026)
CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test Generation
by: Kiecker, Tobias, et al.
Published: (2026)
by: Kiecker, Tobias, et al.
Published: (2026)
Good modelling software practices
by: Lemmen, Carsten, et al.
Published: (2024)
by: Lemmen, Carsten, et al.
Published: (2024)
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026)
by: Kohl, Jens, et al.
Published: (2026)
ASSERTIFY: Utilizing Large Language Models to Generate Assertions for Production Code
by: Torkamani, Mohammad Jalili, et al.
Published: (2024)
by: Torkamani, Mohammad Jalili, et al.
Published: (2024)
A History Equivalence Algorithm for Dynamic Process Migration
by: Bakshi, Gargi, et al.
Published: (2024)
by: Bakshi, Gargi, et al.
Published: (2024)
Detecting and Preventing Latent Risk Accumulation in High-Performance Software Systems
by: Arafat, Jahidul, et al.
Published: (2025)
by: Arafat, Jahidul, et al.
Published: (2025)
A Study of Undefined Behavior Across Foreign Function Boundaries in Rust Libraries
by: McCormack, Ian, et al.
Published: (2024)
by: McCormack, Ian, et al.
Published: (2024)
PyPackIT: Automated Research Software Engineering for Scientific Python Applications on GitHub
by: Ariamajd, Armin, et al.
Published: (2025)
by: Ariamajd, Armin, et al.
Published: (2025)
Conflict Essences for Transformation Rules with Nested Application Conditions -- Long Version
by: Lauer, Alexander, et al.
Published: (2026)
by: Lauer, Alexander, et al.
Published: (2026)
Point Intervention: Improving ACVP Test Vector Generation Through Human Assisted Fuzzing
by: Gridin, Iaroslav, et al.
Published: (2024)
by: Gridin, Iaroslav, et al.
Published: (2024)
QUT: A Unit Testing Framework for Quantum Subroutines
by: Klymenko, Mykhailo V., et al.
Published: (2025)
by: Klymenko, Mykhailo V., et al.
Published: (2025)
It's Alive! What a Live Object Environment Changes in Software Engineering Practice
by: Grigera, Julián, et al.
Published: (2026)
by: Grigera, Julián, et al.
Published: (2026)
NOETHER: A Constructive Framework for Metamorphic Pattern Discovery from Operator Algebras
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Proof of Concept as a First-Class Architectural Decision Instrument
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
Similar Items
-
Testing SSD Firmware with State Data-Aware Fuzzing: Accelerating Coverage in Nondeterministic I/O Environments
by: Yoon, Gangho, et al.
Published: (2025) -
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
by: Yu, Boxi, et al.
Published: (2026) -
Combined Program Analysis Techniques: A Systematic Mapping Study
by: Braione, Pietro, et al.
Published: (2026) -
Validating Formal Specifications with LLM-generated Test Cases
by: Cunha, Alcino, et al.
Published: (2025) -
Synthesizing Test Cases for Narrowing Specification Candidates
by: Cunha, Alcino, et al.
Published: (2025)