Agent-Based Software Artifact Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhaonan, Zhao, Yanjie, Chen, Zhenpeng, Wang, Zheng, Wang, Haoyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
CommitShield: Tracking Vulnerability Introduction and Fix in Version Control Systems
by: Wu, Zhaonan, et al.
Published: (2025)
by: Wu, Zhaonan, et al.
Published: (2025)
CommitSuite: A Comprehensive Benchmark for Commit Classification and Message Generation
by: Wan, Zirui, et al.
Published: (2026)
by: Wan, Zirui, et al.
Published: (2026)
Research Artifacts in Software Engineering Publications: Status and Trends
by: Liu, Mugeng, et al.
Published: (2024)
by: Liu, Mugeng, et al.
Published: (2024)
Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks
by: Ke, Qiang, et al.
Published: (2026)
by: Ke, Qiang, et al.
Published: (2026)
"Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts
by: Wang, Shenao, et al.
Published: (2026)
by: Wang, Shenao, et al.
Published: (2026)
Towards Reliable Vector Database Management Systems: A Software Testing Roadmap for 2030
by: Wang, Shenao, et al.
Published: (2025)
by: Wang, Shenao, et al.
Published: (2025)
Unsafe by Flow: Uncovering Bidirectional Data-Flow Risks in MCP Ecosystem
by: Hou, Xinyi, et al.
Published: (2026)
by: Hou, Xinyi, et al.
Published: (2026)
WaDec: Decompiling WebAssembly Using Large Language Model
by: She, Xinyu, et al.
Published: (2024)
by: She, Xinyu, et al.
Published: (2024)
Voices from the Frontier: A Comprehensive Analysis of the OpenAI Developer Forum
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
by: Wu, Zehao, et al.
Published: (2025)
by: Wu, Zehao, et al.
Published: (2025)
Large Language Model-Based Agents for Software Engineering: A Survey
by: Liu, Junwei, et al.
Published: (2024)
by: Liu, Junwei, et al.
Published: (2024)
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
by: Rao, Hongzhou, et al.
Published: (2025)
by: Rao, Hongzhou, et al.
Published: (2025)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
Towards Robust Detection of Open Source Software Supply Chain Poisoning Attacks in Industry Environments
by: Zheng, Xinyi, et al.
Published: (2024)
by: Zheng, Xinyi, et al.
Published: (2024)
LLM App Store Analysis: A Vision and Roadmap
by: Zhao, Yanjie, et al.
Published: (2024)
by: Zhao, Yanjie, et al.
Published: (2024)
GPTZoo: A Large-scale Dataset of GPTs for the Research Community
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
Large Language Model Supply Chain: A Research Agenda
by: Wang, Shenao, et al.
Published: (2024)
by: Wang, Shenao, et al.
Published: (2024)
LLM Applications: Current Paradigms and the Next Frontier
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI
by: Treude, Christoph, et al.
Published: (2026)
by: Treude, Christoph, et al.
Published: (2026)
Understanding Large Language Model Supply Chain: Structure, Domain, and Vulnerabilities
by: Hu, Yanzhe, et al.
Published: (2025)
by: Hu, Yanzhe, et al.
Published: (2025)
EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents
by: Guo, Yaoqi, et al.
Published: (2026)
by: Guo, Yaoqi, et al.
Published: (2026)
Toward Understanding Bugs in Vector Database Management Systems
by: Xie, Yinglin, et al.
Published: (2025)
by: Xie, Yinglin, et al.
Published: (2025)
Promptware Engineering: Software Engineering for Prompt-Enabled Systems
by: Chen, Zhenpeng, et al.
Published: (2025)
by: Chen, Zhenpeng, et al.
Published: (2025)
REprompt: Prompt Generation for Intelligent Software Development Guided by Requirements Engineering
by: Shi, Junjie, et al.
Published: (2026)
by: Shi, Junjie, et al.
Published: (2026)
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Bridging Design and Development with Automated Declarative UI Code Generation
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
SoK: Systematizing Software Artifacts Traceability via Associations, Techniques, and Applications
by: Chen, Zhifei, et al.
Published: (2026)
by: Chen, Zhifei, et al.
Published: (2026)
CodeMorph: Mitigating Data Leakage in Large Language Model Assessment
by: Rao, Hongzhou, et al.
Published: (2025)
by: Rao, Hongzhou, et al.
Published: (2025)
DocFetch - Towards Generating Software Documentation from Multiple Software Artifacts
by: Venigalla, Akhila Sri Manasa, et al.
Published: (2025)
by: Venigalla, Akhila Sri Manasa, et al.
Published: (2025)
Revealing Floating-Point Accumulation Orders in Software/Hardware Implementations
by: Xie, Peichen, et al.
Published: (2024)
by: Xie, Peichen, et al.
Published: (2024)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
GPT Store Mining and Analysis
by: Su, Dongxun, et al.
Published: (2024)
by: Su, Dongxun, et al.
Published: (2024)
Large Language Models for Software Engineering: A Systematic Literature Review
by: Hou, Xinyi, et al.
Published: (2023)
by: Hou, Xinyi, et al.
Published: (2023)
Same App, Different Behaviors: Uncovering Device-specific Behaviors in Android Apps
by: Dong, Zikan, et al.
Published: (2024)
by: Dong, Zikan, et al.
Published: (2024)
Towards Trustworthy LLMs for Code: A Data-Centric Synergistic Auditing Framework
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Evaluating and Improving ChatGPT-Based Expansion of Abbreviations
by: Jiang, Yanjie, et al.
Published: (2024)
by: Jiang, Yanjie, et al.
Published: (2024)
Wired for Reuse: Automating Context-Aware Code Adaptation in IDEs via LLM-Based Agent
by: Wang, Taiming, et al.
Published: (2025)
by: Wang, Taiming, et al.
Published: (2025)
Artisan: Agentic Artifact Evaluation
by: Baek, Doehyun, et al.
Published: (2026)
by: Baek, Doehyun, et al.
Published: (2026)
Similar Items
-
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026) -
CommitShield: Tracking Vulnerability Introduction and Fix in Version Control Systems
by: Wu, Zhaonan, et al.
Published: (2025) -
CommitSuite: A Comprehensive Benchmark for Commit Classification and Message Generation
by: Wan, Zirui, et al.
Published: (2026) -
Research Artifacts in Software Engineering Publications: Status and Trends
by: Liu, Mugeng, et al.
Published: (2024) -
Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks
by: Ke, Qiang, et al.
Published: (2026)