Testora: Using Natural Language Intent to Detect Behavioral Regressions
Fuente:
arXiv
Saved in:
| Main Author: | Pradel, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Names Are All You Need: Effective and Safe Regression Test Selection for Python
by: Wang, You, et al.
Published: (2026)
by: Wang, You, et al.
Published: (2026)
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
by: Hu, Huimin, et al.
Published: (2025)
by: Hu, Huimin, et al.
Published: (2025)
Change And Cover: Last-Mile, Pull Request-Based Regression Test Augmentation
by: Zhou, Zitong, et al.
Published: (2026)
by: Zhou, Zitong, et al.
Published: (2026)
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
by: Le-Cong, Thanh, et al.
Published: (2026)
by: Le-Cong, Thanh, et al.
Published: (2026)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
by: Deng, Le, et al.
Published: (2025)
by: Deng, Le, et al.
Published: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
by: Eghbali, Aryaz, et al.
Published: (2024)
by: Eghbali, Aryaz, et al.
Published: (2024)
Artisan: Agentic Artifact Evaluation
by: Baek, Doehyun, et al.
Published: (2026)
by: Baek, Doehyun, et al.
Published: (2026)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
by: Wang, You, et al.
Published: (2025)
by: Wang, You, et al.
Published: (2025)
Evaluating LLM Agents on Automated Software Analysis Tasks
by: Bouzenia, Islem, et al.
Published: (2026)
by: Bouzenia, Islem, et al.
Published: (2026)
RippleGUItester: Change-Aware Exploratory Testing
by: Su, Yanqi, et al.
Published: (2026)
by: Su, Yanqi, et al.
Published: (2026)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
by: Gröninger, Lars, et al.
Published: (2024)
by: Gröninger, Lars, et al.
Published: (2024)
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
by: Bouzenia, Islem, et al.
Published: (2025)
by: Bouzenia, Islem, et al.
Published: (2025)
Treefix: Enabling Execution with a Tree of Prefixes
by: Souza, Beatriz, et al.
Published: (2025)
by: Souza, Beatriz, et al.
Published: (2025)
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
AgentStepper: Interactive Debugging of Software Development Agents
by: Hutter, Robert, et al.
Published: (2026)
by: Hutter, Robert, et al.
Published: (2026)
DyPyBench: A Benchmark of Executable Python Software
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
by: Paltenghi, Matteo, et al.
Published: (2023)
by: Paltenghi, Matteo, et al.
Published: (2023)
Issue2Test: Generating Reproducing Test Cases from Issue Reports
by: Nashid, Noor, et al.
Published: (2025)
by: Nashid, Noor, et al.
Published: (2025)
PyTy: Repairing Static Type Errors in Python
by: Chow, Yiu Wai, et al.
Published: (2024)
by: Chow, Yiu Wai, et al.
Published: (2024)
FlyCatcher: Neural Inference of Runtime Checkers from Tests
by: Souza, Beatriz, et al.
Published: (2026)
by: Souza, Beatriz, et al.
Published: (2026)
PATCH: Empowering Large Language Model with Programmer-Intent Guidance and Collaborative-Behavior Simulation for Automatic Bug Fixing
by: Zhang, Yuwei, et al.
Published: (2025)
by: Zhang, Yuwei, et al.
Published: (2025)
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
by: Joos, Pascal, et al.
Published: (2025)
by: Joos, Pascal, et al.
Published: (2025)
PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages
by: Simsek, Deniz, et al.
Published: (2025)
by: Simsek, Deniz, et al.
Published: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
SmartIntentNN: Towards Smart Contract Intent Detection
by: Huang, Youwei, et al.
Published: (2022)
by: Huang, Youwei, et al.
Published: (2022)
From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets
by: Zhu, Hao-Nan, et al.
Published: (2025)
by: Zhu, Hao-Nan, et al.
Published: (2025)
QITE: Assembly-Level, Cross-Platform Testing of Quantum Computing Platforms
by: Paltenghi, Matteo, et al.
Published: (2025)
by: Paltenghi, Matteo, et al.
Published: (2025)
Fuzz4All: Universal Fuzzing with Large Language Models
by: Xia, Chunqiu Steven, et al.
Published: (2023)
by: Xia, Chunqiu Steven, et al.
Published: (2023)
Agentic AI Software Engineers: Programming with Trust
by: Roychoudhury, Abhik, et al.
Published: (2025)
by: Roychoudhury, Abhik, et al.
Published: (2025)
Semantic Caching and Intent-Driven Context Optimization for Multi-Agent Natural Language to Code Systems
by: Singh, Harmohit
Published: (2026)
by: Singh, Harmohit
Published: (2026)
Detecting Malicious Intents in Smart Contracts with Pre-trained Programming Language Models
by: Huang, Youwei, et al.
Published: (2025)
by: Huang, Youwei, et al.
Published: (2025)
Deep Smart Contract Intent Detection
by: Huang, Youwei, et al.
Published: (2022)
by: Huang, Youwei, et al.
Published: (2022)
Optimizing Knowledge Utilization for Multi-Intent Comment Generation with Large Language Models
by: Li, Shuochuan, et al.
Published: (2025)
by: Li, Shuochuan, et al.
Published: (2025)
Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?
by: Endres, Madeline, et al.
Published: (2023)
by: Endres, Madeline, et al.
Published: (2023)
LLM-Based Repair of Static Nullability Errors
by: Karimipour, Nima, et al.
Published: (2025)
by: Karimipour, Nima, et al.
Published: (2025)
Using Large Language Models for Natural Language Processing Tasks in Requirements Engineering: A Systematic Guideline
by: Vogelsang, Andreas, et al.
Published: (2024)
by: Vogelsang, Andreas, et al.
Published: (2024)
A Vulnerability Code Intent Summary Dataset
by: Huang, Yifan, et al.
Published: (2025)
by: Huang, Yifan, et al.
Published: (2025)
Execution-Aware Program Reduction for WebAssembly via Record and Replay
by: Baek, Doehyun, et al.
Published: (2025)
by: Baek, Doehyun, et al.
Published: (2025)
Exploring the Effect of Multiple Natural Languages on Code Suggestion Using GitHub Copilot
by: Koyanagi, Kei, et al.
Published: (2024)
by: Koyanagi, Kei, et al.
Published: (2024)
Intent Preserving Generation of Diverse and Idiomatic (Code-)Artifacts
by: Westphal, Oliver
Published: (2025)
by: Westphal, Oliver
Published: (2025)
Similar Items
-
Names Are All You Need: Effective and Safe Regression Test Selection for Python
by: Wang, You, et al.
Published: (2026) -
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
by: Hu, Huimin, et al.
Published: (2025) -
Change And Cover: Last-Mile, Pull Request-Based Regression Test Augmentation
by: Zhou, Zitong, et al.
Published: (2026) -
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
by: Le-Cong, Thanh, et al.
Published: (2026) -
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
by: Deng, Le, et al.
Published: (2025)