From Code to Courtroom: LLMs as the New Software Judges
Fuente:
arXiv
Saved in:
| Main Authors: | He, Junda, Shi, Jieke, Zhuo, Terry Yue, Treude, Christoph, Sun, Jiamou, Xing, Zhenchang, Du, Xiaoning, Lo, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Identifying and Mitigating API Misuse in Large Language Models
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead
by: He, Junda, et al.
Published: (2024)
by: He, Junda, et al.
Published: (2024)
Greening Large Language Models of Code
by: Shi, Jieke, et al.
Published: (2023)
by: Shi, Jieke, et al.
Published: (2023)
Synthesizing Efficient and Permissive Programmatic Runtime Shields for Neural Policies
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
PTMPicker: Facilitating Efficient Pretrained Model Selection for Application Developers
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Efficient and Green Large Language Models for Software Engineering: Literature Review, Vision, and the Road Ahead
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
LLMAID: Identifying AI Capabilities in Android Apps with LLMs
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Ecosystem of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Qualitative Data Analysis in Software Engineering: Techniques and Teaching Insights
by: Treude, Christoph
Published: (2024)
by: Treude, Christoph
Published: (2024)
ACECode: A Reinforcement Learning Framework for Aligning Code Efficiency and Correctness in Code Language Models
by: Yang, Chengran, et al.
Published: (2024)
by: Yang, Chengran, et al.
Published: (2024)
The Shift from Writing to Pruning Software: A Bonsai-Inspired IDE for Reshaping AI Generated Code
by: Kula, Raula Gaikovina, et al.
Published: (2025)
by: Kula, Raula Gaikovina, et al.
Published: (2025)
Bot-Driven Development: From Simple Automation to Autonomous Software Development Bots
by: Treude, Christoph, et al.
Published: (2024)
by: Treude, Christoph, et al.
Published: (2024)
Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient Shorthand
by: Sun, Zhensu, et al.
Published: (2025)
by: Sun, Zhensu, et al.
Published: (2025)
Think Like Human Developers: Harnessing Community Knowledge for Structured Code Reasoning
by: Yang, Chengran, et al.
Published: (2025)
by: Yang, Chengran, et al.
Published: (2025)
Do Chase Your Tail! Missing Key Aspects Augmentation in Textual Vulnerability Descriptions of Long-tail Software through Feature Inference
by: Han, Linyi, et al.
Published: (2024)
by: Han, Linyi, et al.
Published: (2024)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Finding Safety Violations of AI-Enabled Control Systems through the Lens of Synthesized Proxy Programs
by: Shi, Jieke, et al.
Published: (2024)
by: Shi, Jieke, et al.
Published: (2024)
Curiosity-Driven Testing for Sequential Decision-Making Process
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Gender Influence on Student Teams' Online Communication in Software Engineering Education
by: Garcia, Rita, et al.
Published: (2025)
by: Garcia, Rita, et al.
Published: (2025)
Accountable Agents in Software Engineering: An Analysis of Terms of Service and a Research Roadmap
by: Treude, Christoph
Published: (2026)
by: Treude, Christoph
Published: (2026)
AI Slop and the Software Commons
by: Baltes, Sebastian, et al.
Published: (2026)
by: Baltes, Sebastian, et al.
Published: (2026)
Enhancing Source Code Representations for Deep Learning with Static Analysis
by: Guan, Xueting, et al.
Published: (2024)
by: Guan, Xueting, et al.
Published: (2024)
GenAI Is No Silver Bullet for Qualitative Research in Software Engineering
by: Ernst, Neil A., et al.
Published: (2026)
by: Ernst, Neil A., et al.
Published: (2026)
APIDocBooster: An Extract-Then-Abstract Framework Leveraging Large Language Models for Augmenting API Documentation
by: Yang, Chengran, et al.
Published: (2023)
by: Yang, Chengran, et al.
Published: (2023)
A Functional Software Reference Architecture for LLM-Integrated Systems
by: Bucaioni, Alessio, et al.
Published: (2025)
by: Bucaioni, Alessio, et al.
Published: (2025)
Artificial Intelligence for Software Architecture: Literature Review and the Road Ahead
by: Bucaioni, Alessio, et al.
Published: (2025)
by: Bucaioni, Alessio, et al.
Published: (2025)
Adapting Installation Instructions in Rapidly Evolving Software Ecosystems
by: Gao, Haoyu, et al.
Published: (2023)
by: Gao, Haoyu, et al.
Published: (2023)
Deterministic vs. Probabilistic Summarisation: An Empirical Trade-off Study in Design Pattern Centric Java Code
by: Nazar, Najam, et al.
Published: (2026)
by: Nazar, Najam, et al.
Published: (2026)
Walking the Tightrope of LLMs for Software Development: A Practitioners' Perspective
by: Ferino, Samuel, et al.
Published: (2025)
by: Ferino, Samuel, et al.
Published: (2025)
Can LLMs Deobfuscate Binary Code? A Systematic Analysis of Large Language Models into Pseudocode Deobfuscation
by: Hu, Li, et al.
Published: (2026)
by: Hu, Li, et al.
Published: (2026)
Domain-constrained Synthesis of Inconsistent Key Aspects in Textual Vulnerability Descriptions
by: Han, Linyi, et al.
Published: (2025)
by: Han, Linyi, et al.
Published: (2025)
Unveiling Memorization in Code Models
by: Yang, Zhou, et al.
Published: (2023)
by: Yang, Zhou, et al.
Published: (2023)
"An Endless Stream of AI Slop": The Growing Burden of AI-Assisted Software Development
by: Baltes, Sebastian, et al.
Published: (2026)
by: Baltes, Sebastian, et al.
Published: (2026)
Open Source Software Development Tool Installation: Challenges and Strategies For Novice Developers
by: Salerno, Larissa, et al.
Published: (2024)
by: Salerno, Larissa, et al.
Published: (2024)
Trust in Software Supply Chains: Blockchain-Enabled SBOM and the AIBOM Future
by: Xia, Boming, et al.
Published: (2023)
by: Xia, Boming, et al.
Published: (2023)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
Novice Developers' Perspectives on Adopting LLMs for Software Development: A Systematic Literature Review
by: Ferino, Samuel, et al.
Published: (2025)
by: Ferino, Samuel, et al.
Published: (2025)
Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI
by: Treude, Christoph, et al.
Published: (2026)
by: Treude, Christoph, et al.
Published: (2026)
Similar Items
-
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025) -
Identifying and Mitigating API Misuse in Large Language Models
by: Zhuo, Terry Yue, et al.
Published: (2025) -
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026) -
LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead
by: He, Junda, et al.
Published: (2024) -
Greening Large Language Models of Code
by: Shi, Jieke, et al.
Published: (2023)