Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
Fuente:
arXiv
Saved in:
| Main Authors: | Fandina, Ora Nova, Farchi, Eitan, Froimovich, Shmulik, Gal, Raviv, Ibraheem, Wesam, Katan, Rami, Podolsky, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)
by: Froimovich, Shmulik, et al.
Published: (2025)
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025)
by: Dreyfuss, Itay, et al.
Published: (2025)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
by: Fandina, Ora Nova, et al.
Published: (2024)
by: Fandina, Ora Nova, et al.
Published: (2024)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Black-Box Bug-Amplification for Multithreaded Software
by: Weiss, Yeshayahu, et al.
Published: (2025)
by: Weiss, Yeshayahu, et al.
Published: (2025)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
Generalized Coverage Criteria for Combinatorial Sequence Testing
by: Elyasaf, Achiya, et al.
Published: (2022)
by: Elyasaf, Achiya, et al.
Published: (2022)
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
by: Farchi, Eitan, et al.
Published: (2025)
by: Farchi, Eitan, et al.
Published: (2025)
HintPilot: LLM-based Compiler Hint Synthesis for Code Optimization
by: Jiang, Hanyun, et al.
Published: (2026)
by: Jiang, Hanyun, et al.
Published: (2026)
MioHint: LLM-assisted Mutation for Whitebox API Testing
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
TORAI: Multi-source Root Cause Analysis for Blind Spots in Microservice Service Call Graph
by: Pham, Luan, et al.
Published: (2026)
by: Pham, Luan, et al.
Published: (2026)
Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation
by: Ouyang, Shuyin, et al.
Published: (2026)
by: Ouyang, Shuyin, et al.
Published: (2026)
"Should I Give Up Now?" Investigating LLM Pitfalls in Software Engineering
by: Tie, Jiessie, et al.
Published: (2024)
by: Tie, Jiessie, et al.
Published: (2024)
A Pilot Study on LLM-Based Agentic Translation from Android to iOS: Pitfalls and Insights
by: Zeng, Zhili, et al.
Published: (2025)
by: Zeng, Zhili, et al.
Published: (2025)
Evaluating and Mitigating Errors in LLM-Generated Web API Integrations
by: Maninger, Daniel, et al.
Published: (2025)
by: Maninger, Daniel, et al.
Published: (2025)
Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
by: Milliken, Louis, et al.
Published: (2024)
by: Milliken, Louis, et al.
Published: (2024)
Bridge and Hint: Extending Pre-trained Language Models for Long-Range Code
by: Chen, Yujia, et al.
Published: (2024)
by: Chen, Yujia, et al.
Published: (2024)
The Promise and Pitfalls of WebAssembly: Perspectives from the Industry
by: He, Ningyu, et al.
Published: (2025)
by: He, Ningyu, et al.
Published: (2025)
Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness
by: Ouédraogo, Wendkûuni C., et al.
Published: (2025)
by: Ouédraogo, Wendkûuni C., et al.
Published: (2025)
Hints Help Finding and Fixing Bugs Differently in Python and Text-based Program Representations
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2025)
by: Dristi, Simantika Bhattacharjee, et al.
Published: (2025)
Beyond Rules: LLM-Powered Linting for Quantum Programs
by: Cassieri, Pietro, et al.
Published: (2026)
by: Cassieri, Pietro, et al.
Published: (2026)
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
by: Velasco, Alejandro, et al.
Published: (2025)
by: Velasco, Alejandro, et al.
Published: (2025)
Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation
by: Sarker, Laboni, et al.
Published: (2024)
by: Sarker, Laboni, et al.
Published: (2024)
Online Probabilistic Metric Embedding: A General Framework for Bypassing Inherent Bounds
by: Bartal, Yair, et al.
Published: (2024)
by: Bartal, Yair, et al.
Published: (2024)
Beyond words and actions: Exploring Multimodal Analytics and Collaboration in the Digital Age
by: Miranda, Diego, et al.
Published: (2024)
by: Miranda, Diego, et al.
Published: (2024)
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
by: He, Fusen, et al.
Published: (2024)
by: He, Fusen, et al.
Published: (2024)
Revisiting Out-of-Distribution Detection in Real-time Object Detection: From Benchmark Pitfalls to a New Mitigation Paradigm
by: Wu, Changshun, et al.
Published: (2025)
by: Wu, Changshun, et al.
Published: (2025)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
by: Zhang, Ruixin, et al.
Published: (2026)
by: Zhang, Ruixin, et al.
Published: (2026)
Taxonomy of the Retrieval System Framework: Pitfalls and Paradigms
by: Shah, Deep, et al.
Published: (2026)
by: Shah, Deep, et al.
Published: (2026)
Similar Items
-
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025) -
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025) -
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024) -
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025) -
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025)