Talk is Cheap, Logic is Hard: Benchmarking LLMs on Post-Condition Formalization
Fuente:
arXiv
Saved in:
| Main Authors: | Prasetya, I. S. W. B., Kifetew, Fitsum, Prandi, Davide |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Validating Formal Specifications with LLM-generated Test Cases
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
PICKLES: a Natural Language Framework for Requirement Specification and Model-Based Testing
by: Rodríguez, María Belén, et al.
Published: (2026)
by: Rodríguez, María Belén, et al.
Published: (2026)
Synthesizing Test Cases for Narrowing Specification Candidates
by: Cunha, Alcino, et al.
Published: (2025)
by: Cunha, Alcino, et al.
Published: (2025)
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
by: Salva, Sébastien, et al.
Published: (2025)
by: Salva, Sébastien, et al.
Published: (2025)
Utilizing Precise and Complete Code Context to Guide LLM in Automatic False Positive Mitigation
by: Chen, Jinbao, et al.
Published: (2024)
by: Chen, Jinbao, et al.
Published: (2024)
Constrained LTL Specification Learning from Examples
by: Zhang, Changjian, et al.
Published: (2024)
by: Zhang, Changjian, et al.
Published: (2024)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents
by: Lahiri, Shuvendu K.
Published: (2026)
by: Lahiri, Shuvendu K.
Published: (2026)
The Future of AI-Driven Software Engineering
by: Terragni, Valerio, et al.
Published: (2024)
by: Terragni, Valerio, et al.
Published: (2024)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
by: Downing, Mara, et al.
Published: (2025)
by: Downing, Mara, et al.
Published: (2025)
Causal Models in Requirement Specifications for Machine Learning: A vision
by: Heyn, Hans-Martin, et al.
Published: (2025)
by: Heyn, Hans-Martin, et al.
Published: (2025)
A Survey of the Metrics, Uses, and Subjects of Diversity-Based Techniques in Software Testing
by: Elgendy, Islam T., et al.
Published: (2023)
by: Elgendy, Islam T., et al.
Published: (2023)
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
by: Yu, Boxi, et al.
Published: (2026)
by: Yu, Boxi, et al.
Published: (2026)
Leveraging LLMs for Formal Software Requirements -- Challenges and Prospects
by: Beg, Arshad, et al.
Published: (2025)
by: Beg, Arshad, et al.
Published: (2025)
Tests4Py: A Benchmark for System Testing
by: Smytzek, Marius, et al.
Published: (2023)
by: Smytzek, Marius, et al.
Published: (2023)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
Early-Stage Requirements Transformation Approaches: A Systematic Review
by: Letsholo, Keletso J.
Published: (2024)
by: Letsholo, Keletso J.
Published: (2024)
Towards a Value-Complemented Framework for Enabling Human Monitoring in Cyber-Physical Systems
by: Pfister, Zoe, et al.
Published: (2025)
by: Pfister, Zoe, et al.
Published: (2025)
Towards an Approach to Pattern-based Domain-Specific Requirements Engineering
by: Chuprina, T., et al.
Published: (2024)
by: Chuprina, T., et al.
Published: (2024)
Prompts Blend Requirements and Solutions: From Intent to Implementation
by: Chakraborty, Shalini, et al.
Published: (2026)
by: Chakraborty, Shalini, et al.
Published: (2026)
From Bugs to Benefits: Improving User Stories by Leveraging Crowd Knowledge with CrUISE-AC
by: Schwedt, Stefan, et al.
Published: (2025)
by: Schwedt, Stefan, et al.
Published: (2025)
Generative Goal Modeling
by: Sharfuddin, Ateeq, et al.
Published: (2025)
by: Sharfuddin, Ateeq, et al.
Published: (2025)
Think Like an Engineer: A Neuro-Symbolic Collaboration Agent for Generative Software Requirements Elicitation and Self-Review
by: Zhang, Sai, et al.
Published: (2025)
by: Zhang, Sai, et al.
Published: (2025)
Validating API Design Requirements for Interoperability: A Static Analysis Approach Using OpenAPI
by: Sundberg, Edwin, et al.
Published: (2025)
by: Sundberg, Edwin, et al.
Published: (2025)
Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based Testing
by: Kitsios, Konstantinos, et al.
Published: (2025)
by: Kitsios, Konstantinos, et al.
Published: (2025)
Large Language Models in Software Documentation and Modeling: A Literature Review and Findings
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
It's Alive! What a Live Object Environment Changes in Software Engineering Practice
by: Grigera, Julián, et al.
Published: (2026)
by: Grigera, Julián, et al.
Published: (2026)
Verifying a Sparse Matrix Algorithm Using Symbolic Execution
by: Wilton, Alexander C.
Published: (2025)
by: Wilton, Alexander C.
Published: (2025)
Auto-repair without test cases: How LLMs fix compilation errors in large industrial embedded code
by: Fu, Han, et al.
Published: (2025)
by: Fu, Han, et al.
Published: (2025)
Model Generation with LLMs: From Requirements to UML Sequence Diagrams
by: Ferrari, Alessio, et al.
Published: (2024)
by: Ferrari, Alessio, et al.
Published: (2024)
Large Language Models (LLMs) for Requirements Engineering (RE): A Systematic Literature Review
by: Zadenoori, Mohammad Amin, et al.
Published: (2025)
by: Zadenoori, Mohammad Amin, et al.
Published: (2025)
Inferring Input Grammars from Code with Symbolic Parsing
by: Bettscheider, Leon, et al.
Published: (2025)
by: Bettscheider, Leon, et al.
Published: (2025)
Testing SSD Firmware with State Data-Aware Fuzzing: Accelerating Coverage in Nondeterministic I/O Environments
by: Yoon, Gangho, et al.
Published: (2025)
by: Yoon, Gangho, et al.
Published: (2025)
Combined Program Analysis Techniques: A Systematic Mapping Study
by: Braione, Pietro, et al.
Published: (2026)
by: Braione, Pietro, et al.
Published: (2026)
Automatically Detecting Numerical Instability in Machine Learning Applications via Soft Assertions
by: Sharmin, Shaila, et al.
Published: (2025)
by: Sharmin, Shaila, et al.
Published: (2025)
Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces
by: Vispute, Neelmani, et al.
Published: (2026)
by: Vispute, Neelmani, et al.
Published: (2026)
Proof of Concept as a First-Class Architectural Decision Instrument
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
by: Antognolli, Bruno Fernando, et al.
Published: (2026)
PyPackIT: Automated Research Software Engineering for Scientific Python Applications on GitHub
by: Ariamajd, Armin, et al.
Published: (2025)
by: Ariamajd, Armin, et al.
Published: (2025)
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
by: Li, Shiyang, et al.
Published: (2026)
by: Li, Shiyang, et al.
Published: (2026)
Similar Items
-
Validating Formal Specifications with LLM-generated Test Cases
by: Cunha, Alcino, et al.
Published: (2025) -
PICKLES: a Natural Language Framework for Requirement Specification and Model-Based Testing
by: Rodríguez, María Belén, et al.
Published: (2026) -
Synthesizing Test Cases for Narrowing Specification Candidates
by: Cunha, Alcino, et al.
Published: (2025) -
On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
by: Salva, Sébastien, et al.
Published: (2025) -
Utilizing Precise and Complete Code Context to Guide LLM in Automatic False Positive Mitigation
by: Chen, Jinbao, et al.
Published: (2024)