Validating Formal Specifications with LLM-generated Test Cases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cunha, Alcino, Macedo, Nuno |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Synthesizing Test Cases for Narrowing Specification Candidates
von: Cunha, Alcino, et al.
Veröffentlicht: (2025)
von: Cunha, Alcino, et al.
Veröffentlicht: (2025)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
von: Yu, Boxi, et al.
Veröffentlicht: (2026)
von: Yu, Boxi, et al.
Veröffentlicht: (2026)
Testing SSD Firmware with State Data-Aware Fuzzing: Accelerating Coverage in Nondeterministic I/O Environments
von: Yoon, Gangho, et al.
Veröffentlicht: (2025)
von: Yoon, Gangho, et al.
Veröffentlicht: (2025)
PICKLES: a Natural Language Framework for Requirement Specification and Model-Based Testing
von: Rodríguez, María Belén, et al.
Veröffentlicht: (2026)
von: Rodríguez, María Belén, et al.
Veröffentlicht: (2026)
Model checking of hyperproperties for high-level relational models
von: Macedo, Nuno, et al.
Veröffentlicht: (2025)
von: Macedo, Nuno, et al.
Veröffentlicht: (2025)
Comparing Human and LLM Generated Code: The Jury is Still Out!
von: Licorish, Sherlock A., et al.
Veröffentlicht: (2025)
von: Licorish, Sherlock A., et al.
Veröffentlicht: (2025)
Automatically Detecting Numerical Instability in Machine Learning Applications via Soft Assertions
von: Sharmin, Shaila, et al.
Veröffentlicht: (2025)
von: Sharmin, Shaila, et al.
Veröffentlicht: (2025)
Combined Program Analysis Techniques: A Systematic Mapping Study
von: Braione, Pietro, et al.
Veröffentlicht: (2026)
von: Braione, Pietro, et al.
Veröffentlicht: (2026)
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
von: Salva, Sébastien, et al.
Veröffentlicht: (2025)
von: Salva, Sébastien, et al.
Veröffentlicht: (2025)
The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review
von: Zietsman, Christo
Veröffentlicht: (2026)
von: Zietsman, Christo
Veröffentlicht: (2026)
Talk is Cheap, Logic is Hard: Benchmarking LLMs on Post-Condition Formalization
von: Prasetya, I. S. W. B., et al.
Veröffentlicht: (2026)
von: Prasetya, I. S. W. B., et al.
Veröffentlicht: (2026)
GBM Returns the Best Prediction Performance among Regression Approaches: A Case Study of Stack Overflow Code Quality
von: Licorish, Sherlock A., et al.
Veröffentlicht: (2025)
von: Licorish, Sherlock A., et al.
Veröffentlicht: (2025)
Trace Validation of Unmodified Concurrent Systems with OmniLink
von: Hackett, Finn, et al.
Veröffentlicht: (2026)
von: Hackett, Finn, et al.
Veröffentlicht: (2026)
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
von: Ravi, Ravin, et al.
Veröffentlicht: (2026)
von: Ravi, Ravin, et al.
Veröffentlicht: (2026)
Evaluating Cryptographic API Misuse Detectors for Go
von: Andersson, Vivi, et al.
Veröffentlicht: (2026)
von: Andersson, Vivi, et al.
Veröffentlicht: (2026)
Tractable Verification of Model Transformations: A Cutoff-Theorem Approach for DSLTrans
von: Lucio, Levi
Veröffentlicht: (2026)
von: Lucio, Levi
Veröffentlicht: (2026)
Validating Solidity Code Defects using Symbolic and Concrete Execution powered by Large Language Models
von: Susan, Ştefan-Claudiu, et al.
Veröffentlicht: (2025)
von: Susan, Ştefan-Claudiu, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Formal Software Requirements -- Challenges and Prospects
von: Beg, Arshad, et al.
Veröffentlicht: (2025)
von: Beg, Arshad, et al.
Veröffentlicht: (2025)
Understanding and Reusing Test Suites Across Database Systems
von: Zhong, Suyang, et al.
Veröffentlicht: (2024)
von: Zhong, Suyang, et al.
Veröffentlicht: (2024)
Evaluating LLM-Generated ACSL Annotations for Formal Verification
von: Beg, Arshad, et al.
Veröffentlicht: (2026)
von: Beg, Arshad, et al.
Veröffentlicht: (2026)
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
von: Calboreanu, Elias
Veröffentlicht: (2026)
von: Calboreanu, Elias
Veröffentlicht: (2026)
Short Version of VERIFAI2026 Paper -- Learning Infused Formal Reasoning: Contract Synthesis, Artefact Reuse and Semantic Foundations
von: Beg, Arshad, et al.
Veröffentlicht: (2026)
von: Beg, Arshad, et al.
Veröffentlicht: (2026)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
von: Baldonado, Juan Manuel, et al.
Veröffentlicht: (2025)
von: Baldonado, Juan Manuel, et al.
Veröffentlicht: (2025)
Path-optimal symbolic execution of heap-manipulating programs
von: Braione, Pietro, et al.
Veröffentlicht: (2024)
von: Braione, Pietro, et al.
Veröffentlicht: (2024)
Recommended Practices for Spreadsheet Testing
von: Panko, Raymond R.
Veröffentlicht: (2007)
von: Panko, Raymond R.
Veröffentlicht: (2007)
Three Decades of Formal Methods in Business Process Compliance: A Systematic Literature Review
von: López, Hugo A., et al.
Veröffentlicht: (2024)
von: López, Hugo A., et al.
Veröffentlicht: (2024)
Learning-Infused Formal Reasoning: From Contract Synthesis to Artifact Reuse and Formal Semantics
von: Beg, Arshad, et al.
Veröffentlicht: (2026)
von: Beg, Arshad, et al.
Veröffentlicht: (2026)
Assessing Reliability of Statistical Maximum Coverage Estimators in Fuzzing
von: Liyanage, Danushka, et al.
Veröffentlicht: (2025)
von: Liyanage, Danushka, et al.
Veröffentlicht: (2025)
Good modelling software practices
von: Lemmen, Carsten, et al.
Veröffentlicht: (2024)
von: Lemmen, Carsten, et al.
Veröffentlicht: (2024)
Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches
von: Wienczkowski, Michael
Veröffentlicht: (2026)
von: Wienczkowski, Michael
Veröffentlicht: (2026)
Automated Vulnerability Detection Using Deep Learning Technique
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2024)
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2024)
Monitoring Agentic Systems Before They're Reliable
von: Boston, Marisa Ferrara, et al.
Veröffentlicht: (2026)
von: Boston, Marisa Ferrara, et al.
Veröffentlicht: (2026)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
von: Parris, William M.
Veröffentlicht: (2026)
von: Parris, William M.
Veröffentlicht: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Constrained LTL Specification Learning from Examples
von: Zhang, Changjian, et al.
Veröffentlicht: (2024)
von: Zhang, Changjian, et al.
Veröffentlicht: (2024)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
von: Kessel, Marcus
Veröffentlicht: (2024)
von: Kessel, Marcus
Veröffentlicht: (2024)
A Short Survey on Formalising Software Requirements using Large Language Models
von: Beg, Arshad, et al.
Veröffentlicht: (2025)
von: Beg, Arshad, et al.
Veröffentlicht: (2025)
Utilizing Precise and Complete Code Context to Guide LLM in Automatic False Positive Mitigation
von: Chen, Jinbao, et al.
Veröffentlicht: (2024)
von: Chen, Jinbao, et al.
Veröffentlicht: (2024)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
von: Rehan, Tzafrir
Veröffentlicht: (2026)
von: Rehan, Tzafrir
Veröffentlicht: (2026)
Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation
von: Marri, Srinivas Rao
Veröffentlicht: (2026)
von: Marri, Srinivas Rao
Veröffentlicht: (2026)
Ähnliche Einträge
-
Synthesizing Test Cases for Narrowing Specification Candidates
von: Cunha, Alcino, et al.
Veröffentlicht: (2025) -
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
von: Yu, Boxi, et al.
Veröffentlicht: (2026) -
Testing SSD Firmware with State Data-Aware Fuzzing: Accelerating Coverage in Nondeterministic I/O Environments
von: Yoon, Gangho, et al.
Veröffentlicht: (2025) -
PICKLES: a Natural Language Framework for Requirement Specification and Model-Based Testing
von: Rodríguez, María Belén, et al.
Veröffentlicht: (2026) -
Model checking of hyperproperties for high-level relational models
von: Macedo, Nuno, et al.
Veröffentlicht: (2025)