Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Linxin, Chen, Jiefeng, Huang, Yue, Mishra, Bhavana Dalvi, Wang, Chi, Zhao, Jieyu, Yoon, Jinsung, Pfister, Tomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Finding Logic Bugs in Spatial Database Engines via Affine Equivalent Inputs
von: Deng, Wenjing, et al.
Veröffentlicht: (2024)
von: Deng, Wenjing, et al.
Veröffentlicht: (2024)
Generating Equivalent Representations of Code By A Self-Reflection Approach
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Is Productivity in Quantum Programming Equivalent to Expressiveness?
von: Corrales-Garro, Francini, et al.
Veröffentlicht: (2025)
von: Corrales-Garro, Francini, et al.
Veröffentlicht: (2025)
FormulaCode: Evaluating Agentic Optimization on Large Codebases
von: Sehgal, Atharva, et al.
Veröffentlicht: (2026)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2026)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
von: Chen, Jialong, et al.
Veröffentlicht: (2026)
von: Chen, Jialong, et al.
Veröffentlicht: (2026)
VERT: Verified Equivalent Rust Transpilation with Large Language Models as Few-Shot Learners
von: Yang, Aidan Z. H., et al.
Veröffentlicht: (2024)
von: Yang, Aidan Z. H., et al.
Veröffentlicht: (2024)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
von: Luo, Jane, et al.
Veröffentlicht: (2025)
von: Luo, Jane, et al.
Veröffentlicht: (2025)
Proving Cypher Query Equivalence
von: Tang, Lei, et al.
Veröffentlicht: (2025)
von: Tang, Lei, et al.
Veröffentlicht: (2025)
The Secrets Must Not Flow: Scaling Security Verification to Large Codebases (extended version)
von: Arquint, Linard, et al.
Veröffentlicht: (2025)
von: Arquint, Linard, et al.
Veröffentlicht: (2025)
QEMI: A Quantum Software Stacks Testing Framework via Equivalence Modulo Inputs
von: Luo, Junjie, et al.
Veröffentlicht: (2026)
von: Luo, Junjie, et al.
Veröffentlicht: (2026)
Verify Implementation Equivalence of Large Models
von: Zhan, Qi, et al.
Veröffentlicht: (2026)
von: Zhan, Qi, et al.
Veröffentlicht: (2026)
SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution
von: He, Kang, et al.
Veröffentlicht: (2026)
von: He, Kang, et al.
Veröffentlicht: (2026)
An Empirical Evaluation of Manually Created Equivalent Mutants
von: Straubinger, Philipp, et al.
Veröffentlicht: (2024)
von: Straubinger, Philipp, et al.
Veröffentlicht: (2024)
What can Large Language Models Capture about Code Functional Equivalence?
von: Maveli, Nickil, et al.
Veröffentlicht: (2024)
von: Maveli, Nickil, et al.
Veröffentlicht: (2024)
Large Language Models for Equivalent Mutant Detection: How Far Are We?
von: Tian, Zhao, et al.
Veröffentlicht: (2024)
von: Tian, Zhao, et al.
Veröffentlicht: (2024)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
Code-Survey: An LLM-Driven Methodology for Analyzing Large-Scale Codebases
von: Zheng, Yusheng, et al.
Veröffentlicht: (2024)
von: Zheng, Yusheng, et al.
Veröffentlicht: (2024)
Disproving Program Equivalence with LLMs
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2025)
Refactoring Codebases through Library Design
von: Kovacic, Ziga, et al.
Veröffentlicht: (2025)
von: Kovacic, Ziga, et al.
Veröffentlicht: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
von: Liu, Steven, et al.
Veröffentlicht: (2026)
von: Liu, Steven, et al.
Veröffentlicht: (2026)
Compiler Optimization Testing Based on Optimization-Guided Equivalence Transformations
von: Wu, Jingwen, et al.
Veröffentlicht: (2025)
von: Wu, Jingwen, et al.
Veröffentlicht: (2025)
Prometheus: Towards Long-Horizon Codebase Navigation for Repository-Level Problem Solving
von: Pan, Yue, et al.
Veröffentlicht: (2025)
von: Pan, Yue, et al.
Veröffentlicht: (2025)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
von: Wu, Jiarong, et al.
Veröffentlicht: (2025)
von: Wu, Jiarong, et al.
Veröffentlicht: (2025)
RacerF: Lightweight Static Data Race Detection for C Code
von: Dacík, Tomáš, et al.
Veröffentlicht: (2025)
von: Dacík, Tomáš, et al.
Veröffentlicht: (2025)
RacerF: Data Race Detection with Frama-C (Competition Contribution)
von: Dacík, Tomáš, et al.
Veröffentlicht: (2025)
von: Dacík, Tomáš, et al.
Veröffentlicht: (2025)
Equivalence, Identity, and Unitarity Checking in Black-Box Testing of Quantum Programs
von: Long, Peixun, et al.
Veröffentlicht: (2023)
von: Long, Peixun, et al.
Veröffentlicht: (2023)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Codified Context: Infrastructure for AI Agents in a Complex Codebase
von: Vasilopoulos, Aristidis
Veröffentlicht: (2026)
von: Vasilopoulos, Aristidis
Veröffentlicht: (2026)
Documentation-Guided Agentic Codebase Migration from C to Rust
von: Le-Anh, Minh, et al.
Veröffentlicht: (2026)
von: Le-Anh, Minh, et al.
Veröffentlicht: (2026)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
von: Song, Huatong, et al.
Veröffentlicht: (2026)
von: Song, Huatong, et al.
Veröffentlicht: (2026)
Unseen-Codebases-Domain Data Synthesis and Training Based on Code Graphs
von: Ou, Guangsheng, et al.
Veröffentlicht: (2026)
von: Ou, Guangsheng, et al.
Veröffentlicht: (2026)
ReqToCode: Embedding Requirements Traceability as a Structural Property of the Codebase
von: Schlathölter, Thorsten
Veröffentlicht: (2026)
von: Schlathölter, Thorsten
Veröffentlicht: (2026)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
von: Han, Hojae, et al.
Veröffentlicht: (2025)
von: Han, Hojae, et al.
Veröffentlicht: (2025)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
von: Du, Junjia, et al.
Veröffentlicht: (2025)
von: Du, Junjia, et al.
Veröffentlicht: (2025)
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
von: Ding, Xianzhong, et al.
Veröffentlicht: (2026)
von: Ding, Xianzhong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Finding Logic Bugs in Spatial Database Engines via Affine Equivalent Inputs
von: Deng, Wenjing, et al.
Veröffentlicht: (2024) -
Generating Equivalent Representations of Code By A Self-Reflection Approach
von: Li, Jia, et al.
Veröffentlicht: (2024) -
Is Productivity in Quantum Programming Equivalent to Expressiveness?
von: Corrales-Garro, Francini, et al.
Veröffentlicht: (2025) -
FormulaCode: Evaluating Agentic Optimization on Large Codebases
von: Sehgal, Atharva, et al.
Veröffentlicht: (2026) -
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
von: Chen, Jialong, et al.
Veröffentlicht: (2026)