Agentic Harness for Real-World Compilers
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Yingwei, Li, Cong, Li, Shaohua, Zhang, Yuqun, Su, Zhendong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
API-guided Dataset Synthesis to Finetune Large Code Models
by: Li, Zongjie, et al.
Published: (2024)
by: Li, Zongjie, et al.
Published: (2024)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)
by: Zhang, Zehua, et al.
Published: (2025)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
by: Xu, Jingxuan, et al.
Published: (2025)
by: Xu, Jingxuan, et al.
Published: (2025)
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
by: Li, Lehui, et al.
Published: (2026)
by: Li, Lehui, et al.
Published: (2026)
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces
by: Xu, Duling, et al.
Published: (2026)
by: Xu, Duling, et al.
Published: (2026)
NeuroStrata: Harnessing Neurosymbolic Paradigms for Improved Design, Testability, and Verifiability of Autonomous CPS
by: Zheng, Xi, et al.
Published: (2025)
by: Zheng, Xi, et al.
Published: (2025)
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
by: Kong, Ziqiao, et al.
Published: (2026)
by: Kong, Ziqiao, et al.
Published: (2026)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
by: Li, Chenxin, et al.
Published: (2026)
by: Li, Chenxin, et al.
Published: (2026)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
by: Chen, Mouxiang, et al.
Published: (2026)
by: Chen, Mouxiang, et al.
Published: (2026)
DeepCode: Open Agentic Coding
by: Li, Zongwei, et al.
Published: (2025)
by: Li, Zongwei, et al.
Published: (2025)
LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues
by: Lin, Yalan, et al.
Published: (2024)
by: Lin, Yalan, et al.
Published: (2024)
The New Compiler Stack: A Survey on the Synergy of LLMs and Compilers
by: Zhang, Shuoming, et al.
Published: (2026)
by: Zhang, Shuoming, et al.
Published: (2026)
LEGO-Compiler: Enhancing Neural Compilation Through Translation Composability
by: Zhang, Shuoming, et al.
Published: (2025)
by: Zhang, Shuoming, et al.
Published: (2025)
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
by: Zhang, Zejun, et al.
Published: (2026)
by: Zhang, Zejun, et al.
Published: (2026)
Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute
by: Ma, Yingwei, et al.
Published: (2025)
by: Ma, Yingwei, et al.
Published: (2025)
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
by: Pan, Zhenyu, et al.
Published: (2024)
by: Pan, Zhenyu, et al.
Published: (2024)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench
by: Zou, Qingyun, et al.
Published: (2026)
by: Zou, Qingyun, et al.
Published: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
by: Ouyang, Yipeng, et al.
Published: (2026)
by: Ouyang, Yipeng, et al.
Published: (2026)
PCodeTrans: Translate Decompiled Pseudocode to Compilable and Executable Equivalent
by: Cui, Yuxin, et al.
Published: (2026)
by: Cui, Yuxin, et al.
Published: (2026)
Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement
by: Ma, Yingwei, et al.
Published: (2024)
by: Ma, Yingwei, et al.
Published: (2024)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
by: Ma, Jeffrey Jian, et al.
Published: (2025)
by: Ma, Jeffrey Jian, et al.
Published: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
SCoGen: Scenario-Centric Graph-Based Synthesis of Real-World Code Problems
by: Yao, Xifeng, et al.
Published: (2025)
by: Yao, Xifeng, et al.
Published: (2025)
Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Stop Comparing LLM Agents Without Disclosing the Harness
by: Zhang, Yunbei, et al.
Published: (2026)
by: Zhang, Yunbei, et al.
Published: (2026)
FVSpec: Real-World Property-Based Tests as Lean Challenges
by: Dougherty, Quinn, et al.
Published: (2026)
by: Dougherty, Quinn, et al.
Published: (2026)
LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation
by: Xu, Dong, et al.
Published: (2026)
by: Xu, Dong, et al.
Published: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
UBfuzz: Finding Bugs in Sanitizer Implementations
by: Li, Shaohua, et al.
Published: (2024)
by: Li, Shaohua, et al.
Published: (2024)
Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
by: Ma, Xuyan, et al.
Published: (2025)
by: Ma, Xuyan, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development
by: Casserini, Matteo, et al.
Published: (2026)
by: Casserini, Matteo, et al.
Published: (2026)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
by: Han, Tingxu, et al.
Published: (2026)
by: Han, Tingxu, et al.
Published: (2026)
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
by: Li, Zenan, et al.
Published: (2026)
by: Li, Zenan, et al.
Published: (2026)
QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks
by: Hu, Yaojie, et al.
Published: (2025)
by: Hu, Yaojie, et al.
Published: (2025)
Similar Items
-
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026) -
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
by: Li, Jia, et al.
Published: (2026) -
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026) -
API-guided Dataset Synthesis to Finetune Large Code Models
by: Li, Zongjie, et al.
Published: (2024) -
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
by: Zhang, Zehua, et al.
Published: (2025)