Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
Fuente:
arXiv
Saved in:
| Main Authors: | Fandina, Ora Nova, Amram, Gal, Farchi, Eitan, Froimovich, Shmulik, Gal, Raviv, Ibraheem, Wesam, Katan, Rami, Podolsky, Alice, Raz, Orna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)
by: Froimovich, Shmulik, et al.
Published: (2025)
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
by: Fandina, Ora Nova, et al.
Published: (2024)
by: Fandina, Ora Nova, et al.
Published: (2024)
Generating Unseen Code Tests In Infinitum
by: Zalmanovici, Marcel, et al.
Published: (2024)
by: Zalmanovici, Marcel, et al.
Published: (2024)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025)
by: Dreyfuss, Itay, et al.
Published: (2025)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Black-Box Bug-Amplification for Multithreaded Software
by: Weiss, Yeshayahu, et al.
Published: (2025)
by: Weiss, Yeshayahu, et al.
Published: (2025)
Detecting Zariski Pairs by Algorithms and Computational Classification in Conic Line Arrangements
by: Amram, Meirav, et al.
Published: (2026)
by: Amram, Meirav, et al.
Published: (2026)
Cost Function Unrolling in Unsupervised Optical Flow
by: Lifshitz, Gal, et al.
Published: (2020)
by: Lifshitz, Gal, et al.
Published: (2020)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding
by: Nevo, Ziv, et al.
Published: (2025)
by: Nevo, Ziv, et al.
Published: (2025)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Online Probabilistic Metric Embedding: A General Framework for Bypassing Inherent Bounds
by: Bartal, Yair, et al.
Published: (2024)
by: Bartal, Yair, et al.
Published: (2024)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Ecosystem Engineering by Mussels ( Margaritifera margaritifera ) Influences the Behaviour of Their Host Fish Brown Trout ( Salmo trutta ) Under Various Flows
by: Magnus Lovén Wallerius, et al.
Published: (2026)
by: Magnus Lovén Wallerius, et al.
Published: (2026)
On the Robustness of Diffusion-Based Image Compression to Bit-Flip Errors
by: Vaisman, Amit, et al.
Published: (2026)
by: Vaisman, Amit, et al.
Published: (2026)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
Automata Models for Effective Bug Pattern Description
by: Yaacov, Tom, et al.
Published: (2025)
by: Yaacov, Tom, et al.
Published: (2025)
L-SR1: Learned Symmetric-Rank-One Preconditioning
by: Lifshitz, Gal, et al.
Published: (2025)
by: Lifshitz, Gal, et al.
Published: (2025)
Systole-Conditioned Generative Cardiac Motion
by: Zuler, Shahar, et al.
Published: (2025)
by: Zuler, Shahar, et al.
Published: (2025)
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
by: HIdekel, Ido Nitzan, et al.
Published: (2025)
by: HIdekel, Ido Nitzan, et al.
Published: (2025)
Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds
by: Shaar, Eitan, et al.
Published: (2025)
by: Shaar, Eitan, et al.
Published: (2025)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
Adoption of Generative AI Technologies: Insights From the UTAUT2 Model, Personality Characteristics, and Behavioural Factors
by: Tali Gazit, et al.
Published: (2025)
by: Tali Gazit, et al.
Published: (2025)
Implementing Reinforcement Learning Datacenter Congestion Control in NVIDIA NICs
by: Fuhrer, Benjamin, et al.
Published: (2022)
by: Fuhrer, Benjamin, et al.
Published: (2022)
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Bone Diseases as Indicators of Animal Health in the Early Modern Age Assemblage From the Castle of Dombóvár‐Gólyavár in Context With Other Coeval Cases From Hungary
by: Erika Gál, et al.
Published: (2025)
by: Erika Gál, et al.
Published: (2025)
Linear Codes for Hyperdimensional Computing
by: Raviv, Netanel
Published: (2024)
by: Raviv, Netanel
Published: (2024)
Reply to commentary by Offer Rozenstein on ‘Is the crop evapotranspiration rate a good surrogate for the recommended irrigation rate?’
by: Shmulik P. Friedman
Published: (2024)
by: Shmulik P. Friedman
Published: (2024)
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
by: Farchi, Eitan, et al.
Published: (2025)
by: Farchi, Eitan, et al.
Published: (2025)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Similar Items
-
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025) -
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025) -
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025) -
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024) -
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)