Engineering AI Judge Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Jiahuei, Lin, Dayi, Zhang, Sky, Hassan, Ahmed E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024)
by: Hassan, Ahmed E., et al.
Published: (2024)
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025)
by: Hassan, Ahmed E., et al.
Published: (2025)
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
by: Rombaut, Benjamin, et al.
Published: (2024)
by: Rombaut, Benjamin, et al.
Published: (2024)
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
by: Gallaba, Keheliya, et al.
Published: (2025)
by: Gallaba, Keheliya, et al.
Published: (2025)
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
by: Fan, Zhiyu, et al.
Published: (2025)
by: Fan, Zhiyu, et al.
Published: (2025)
Data Quality Antipatterns for Software Analytics
by: Bhatia, Aaditya, et al.
Published: (2024)
by: Bhatia, Aaditya, et al.
Published: (2024)
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
by: Rajbahadur, Gopi Krishnan, et al.
Published: (2024)
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
by: Vasilevski, Kirill, et al.
Published: (2025)
by: Vasilevski, Kirill, et al.
Published: (2025)
AIDev: Studying AI Coding Agents on GitHub
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025)
by: Oliva, Gustavo A., et al.
Published: (2025)
GenAI for Simulation Model in Model-Based Systems Engineering
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Unified Software Engineering Agent as AI Software Engineer
by: Applis, Leonhard, et al.
Published: (2025)
by: Applis, Leonhard, et al.
Published: (2025)
The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Software Engineering and Foundation Models: Insights from Industry Blogs Using a Jury of Foundation Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Software Performance Engineering for Foundation Model-Powered Software
by: Zhang, Haoxiang, et al.
Published: (2024)
by: Zhang, Haoxiang, et al.
Published: (2024)
Keeping Deep Learning Models in Check: A History-Based Approach to Mitigate Overfitting
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
by: Zhao, Zixiao, et al.
Published: (2026)
by: Zhao, Zixiao, et al.
Published: (2026)
Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
by: Sartaj, Hassan, et al.
Published: (2025)
by: Sartaj, Hassan, et al.
Published: (2025)
A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
by: Jewitt, James, et al.
Published: (2025)
by: Jewitt, James, et al.
Published: (2025)
Leveraging LLMs for User Stories in AI Systems: UStAI Dataset
by: Yamani, Asma, et al.
Published: (2025)
by: Yamani, Asma, et al.
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation
by: Tantithamthavorn, Kla, et al.
Published: (2026)
by: Tantithamthavorn, Kla, et al.
Published: (2026)
Greening AI-enabled Systems with Software Engineering: A Research Agenda for Environmentally Sustainable AI Practices
by: Cruz, Luís, et al.
Published: (2025)
by: Cruz, Luís, et al.
Published: (2025)
The Cognitive Circuit Breaker: A Systems Engineering Framework for Intrinsic AI Reliability
by: Pan, Jonathan
Published: (2026)
by: Pan, Jonathan
Published: (2026)
LLMs Integration in Software Engineering Team Projects: Roles, Impact, and a Pedagogical Design Space for AI Tools in Computing Education
by: Kharrufa, Ahmed, et al.
Published: (2024)
by: Kharrufa, Ahmed, et al.
Published: (2024)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
by: Tan, Jingwen, et al.
Published: (2024)
by: Tan, Jingwen, et al.
Published: (2024)
HAFixAgent: History-Aware Program Repair Agent
by: Shi, Yu, et al.
Published: (2025)
by: Shi, Yu, et al.
Published: (2025)
A Case Study on AI Engineering Practices: Developing an Autonomous Stock Trading System
by: Grote, Marcel, et al.
Published: (2023)
by: Grote, Marcel, et al.
Published: (2023)
On the Impact of Black-box Deployment Strategies for Edge AI on Latency and Model Performance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
From Code-Centric to Intent-Centric Software Engineering: A Reflexive Thematic Analysis of Generative AI, Agentic Systems, and Engineering Accountability
by: De La Cruz, Elyson
Published: (2026)
by: De La Cruz, Elyson
Published: (2026)
Evolution without an Oracle: Driving Effective Evolution with LLM Judges
by: Zhao, Zhe, et al.
Published: (2025)
by: Zhao, Zhe, et al.
Published: (2025)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
AI-Tutoring in Software Engineering Education
by: Frankford, Eduard, et al.
Published: (2024)
by: Frankford, Eduard, et al.
Published: (2024)
OmniLLP: Enhancing LLM-based Log Level Prediction with Context-Aware Retrieval
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
by: Ouatiti, Youssef Esseddiq, et al.
Published: (2025)
Agentic AI Software Engineers: Programming with Trust
by: Roychoudhury, Abhik, et al.
Published: (2025)
by: Roychoudhury, Abhik, et al.
Published: (2025)
SimClone: Detecting Tabular Data Clones using Value Similarity
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
An Empirical Study of Challenges in Machine Learning Asset Management
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
Similar Items
-
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
by: Hassan, Ahmed E., et al.
Published: (2024) -
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
by: Hassan, Ahmed E., et al.
Published: (2024) -
Rethinking Software Engineering in the Foundation Model Era: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
by: Hassan, Ahmed E., et al.
Published: (2024) -
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
by: Hassan, Ahmed E., et al.
Published: (2025) -
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
by: Rombaut, Benjamin, et al.
Published: (2024)