LaajMeter: A Framework for LaaJ Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Ackerman, Samuel, Amram, Gal, Fandina, Ora Nova, Farchi, Eitan, Froimovich, Shmulik, Gal, Raviv, Ibraheem, Wesam, Ziv, Avi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025)
by: Froimovich, Shmulik, et al.
Published: (2025)
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
by: Fandina, Ora Nova, et al.
Published: (2024)
by: Fandina, Ora Nova, et al.
Published: (2024)
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Black-Box Bug-Amplification for Multithreaded Software
by: Weiss, Yeshayahu, et al.
Published: (2025)
by: Weiss, Yeshayahu, et al.
Published: (2025)
Automata Models for Effective Bug Pattern Description
by: Yaacov, Tom, et al.
Published: (2025)
by: Yaacov, Tom, et al.
Published: (2025)
CARE: Counselor-Aligned Response Engine for Online Mental-Health Support
by: Astrin, Hagai, et al.
Published: (2026)
by: Astrin, Hagai, et al.
Published: (2026)
Detecting Zariski Pairs by Algorithms and Computational Classification in Conic Line Arrangements
by: Amram, Meirav, et al.
Published: (2026)
by: Amram, Meirav, et al.
Published: (2026)
Cost Function Unrolling in Unsupervised Optical Flow
by: Lifshitz, Gal, et al.
Published: (2020)
by: Lifshitz, Gal, et al.
Published: (2020)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
Online Probabilistic Metric Embedding: A General Framework for Bypassing Inherent Bounds
by: Bartal, Yair, et al.
Published: (2024)
by: Bartal, Yair, et al.
Published: (2024)
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Classifier-Guided Captioning Across Modalities
by: Shaulov, Ariel, et al.
Published: (2025)
by: Shaulov, Ariel, et al.
Published: (2025)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Probing Subphonemes in Morphology Models
by: Astrach, Gal, et al.
Published: (2025)
by: Astrach, Gal, et al.
Published: (2025)
Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models
by: Steinberg, Jonathan, et al.
Published: (2026)
by: Steinberg, Jonathan, et al.
Published: (2026)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
by: Aperstein, Yehudit, et al.
Published: (2025)
by: Aperstein, Yehudit, et al.
Published: (2025)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
by: Shaar, Eitan, et al.
Published: (2026)
by: Shaar, Eitan, et al.
Published: (2026)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
by: Feng, Zhili, et al.
Published: (2025)
by: Feng, Zhili, et al.
Published: (2025)
Resistant or Favorable? Chinese Learners’ Beliefs towards Task-Based Language Learning and Teaching
by: Damien Le Gal
Published: (2015)
by: Damien Le Gal
Published: (2015)
LR-DWM: Efficient Watermarking for Diffusion Language Models
by: Raban, Ofek, et al.
Published: (2026)
by: Raban, Ofek, et al.
Published: (2026)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
by: Yona, Gal, et al.
Published: (2024)
by: Yona, Gal, et al.
Published: (2024)
Hallucinations Undermine Trust; Metacognition is a Way Forward
by: Yona, Gal, et al.
Published: (2026)
by: Yona, Gal, et al.
Published: (2026)
Ecosystem Engineering by Mussels ( Margaritifera margaritifera ) Influences the Behaviour of Their Host Fish Brown Trout ( Salmo trutta ) Under Various Flows
by: Magnus Lovén Wallerius, et al.
Published: (2026)
by: Magnus Lovén Wallerius, et al.
Published: (2026)
Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework
by: Buskila, Avi-ad Avraam
Published: (2026)
by: Buskila, Avi-ad Avraam
Published: (2026)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
by: Aharon, Eliya Naomi, et al.
Published: (2026)
by: Aharon, Eliya Naomi, et al.
Published: (2026)
Similar Items
-
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025) -
Quality Evaluation of COBOL to Java Code Transformation
by: Froimovich, Shmulik, et al.
Published: (2025) -
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025) -
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025) -
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)