Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Dong, Du, Mingzhe, Zhang, Jie M., Lin, Zheng, Luo, Meng, Zhang, Qianru, Ng, See-Kiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking LLMs for Unit Test Generation from Real-World Functions
von: Huang, Dong, et al.
Veröffentlicht: (2025)
von: Huang, Dong, et al.
Veröffentlicht: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
von: Du, Mingzhe, et al.
Veröffentlicht: (2025)
von: Du, Mingzhe, et al.
Veröffentlicht: (2025)
Assessing Evaluation Metrics for Neural Test Oracle Generation
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
Automated Discovery of Test Oracles for Database Management Systems Using LLMs
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
KAIJU: An Executive Kernel for Intent-Gated Execution of LLM Agents
von: Guerin, Cormac, et al.
Veröffentlicht: (2026)
von: Guerin, Cormac, et al.
Veröffentlicht: (2026)
Dynamic Scaling of Unit Tests for Code Reward Modeling
von: Ma, Zeyao, et al.
Veröffentlicht: (2025)
von: Ma, Zeyao, et al.
Veröffentlicht: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
von: Liu, Steven, et al.
Veröffentlicht: (2026)
von: Liu, Steven, et al.
Veröffentlicht: (2026)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
GenX: Mastering Code and Test Generation with Execution Feedback
von: Wang, Nan, et al.
Veröffentlicht: (2024)
von: Wang, Nan, et al.
Veröffentlicht: (2024)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
GUI Test Migration via Abstraction and Concretization
von: Zhang, Yakun, et al.
Veröffentlicht: (2024)
von: Zhang, Yakun, et al.
Veröffentlicht: (2024)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
von: Ludwig, Nikolai, et al.
Veröffentlicht: (2026)
von: Ludwig, Nikolai, et al.
Veröffentlicht: (2026)
Multi-Pass Targeted Dynamic Symbolic Execution
von: Yavuz, Tuba
Veröffentlicht: (2024)
von: Yavuz, Tuba
Veröffentlicht: (2024)
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
von: Han, Hojae, et al.
Veröffentlicht: (2026)
von: Han, Hojae, et al.
Veröffentlicht: (2026)
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
Mica: Automated Differential Testing for OCaml Modules
von: Ng, Ernest, et al.
Veröffentlicht: (2024)
von: Ng, Ernest, et al.
Veröffentlicht: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents
von: Wang, Qihao, et al.
Veröffentlicht: (2026)
von: Wang, Qihao, et al.
Veröffentlicht: (2026)
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
von: Pysklo, Hubert M., et al.
Veröffentlicht: (2026)
von: Pysklo, Hubert M., et al.
Veröffentlicht: (2026)
WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements
von: Teoh, Xiwen, et al.
Veröffentlicht: (2026)
von: Teoh, Xiwen, et al.
Veröffentlicht: (2026)
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
von: Tung, Lam Nguyen, et al.
Veröffentlicht: (2024)
von: Tung, Lam Nguyen, et al.
Veröffentlicht: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
Python Symbolic Execution with LLM-powered Code Generation
von: Wang, Wenhan, et al.
Veröffentlicht: (2024)
von: Wang, Wenhan, et al.
Veröffentlicht: (2024)
SEVerA: Verified Synthesis of Self-Evolving Agents
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2026)
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
von: Li, Kenan, et al.
Veröffentlicht: (2026)
von: Li, Kenan, et al.
Veröffentlicht: (2026)
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2026)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
Towards Exception Safety Code Generation with Intermediate Representation Agents Framework
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
CUTECat: Concolic Execution for Computational Law
von: Goutagny, Pierre, et al.
Veröffentlicht: (2024)
von: Goutagny, Pierre, et al.
Veröffentlicht: (2024)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
Test Oracle Automation in the era of LLMs
von: Molina, Facundo, et al.
Veröffentlicht: (2024)
von: Molina, Facundo, et al.
Veröffentlicht: (2024)
From Code Foundation Models to Agents and Applications: A Comprehensive Survey and Practical Guide to Code Intelligence
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
AugmenTest: Enhancing Tests with LLM-Driven Oracles
von: Khandaker, Shaker Mahmud, et al.
Veröffentlicht: (2025)
von: Khandaker, Shaker Mahmud, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Benchmarking LLMs for Unit Test Generation from Real-World Functions
von: Huang, Dong, et al.
Veröffentlicht: (2025) -
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024) -
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024) -
CodeArena: A Collective Evaluation Platform for LLM Code Generation
von: Du, Mingzhe, et al.
Veröffentlicht: (2025) -
Assessing Evaluation Metrics for Neural Test Oracle Generation
von: Shin, Jiho, et al.
Veröffentlicht: (2023)