Jury: A Comprehensive Evaluation Toolkit
Fuente:
arXiv
Saved in:
| Main Authors: | Cavusoglu, Devrim, Sen, Secil, Sert, Ulas, Altinuc, Sinan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
by: Guo, Dongxin
Published: (2026)
by: Guo, Dongxin
Published: (2026)
When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
by: Coen, Michael H.
Published: (2025)
by: Coen, Michael H.
Published: (2025)
Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
by: Chelombitko, Iaroslav, et al.
Published: (2024)
by: Chelombitko, Iaroslav, et al.
Published: (2024)
How to Evaluate Medical AI
by: Kopanichuk, Ilia, et al.
Published: (2025)
by: Kopanichuk, Ilia, et al.
Published: (2025)
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
by: Yan, Lingyong, et al.
Published: (2026)
by: Yan, Lingyong, et al.
Published: (2026)
Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation
by: Wu, Jikun, et al.
Published: (2026)
by: Wu, Jikun, et al.
Published: (2026)
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
by: Pan, Xinghan
Published: (2025)
by: Pan, Xinghan
Published: (2025)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
by: Tao, Xinli, et al.
Published: (2025)
by: Tao, Xinli, et al.
Published: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
by: Rodriguez, David, et al.
Published: (2025)
by: Rodriguez, David, et al.
Published: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
by: Hasan, Mohammed Rakibul
Published: (2026)
by: Hasan, Mohammed Rakibul
Published: (2026)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
by: Loi, Dario, et al.
Published: (2025)
by: Loi, Dario, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
NL2LOGIC: AST-Guided Translation of Natural Language into First-Order Logic with Large Language Models
by: Putra, Rizky Ramadhana, et al.
Published: (2026)
by: Putra, Rizky Ramadhana, et al.
Published: (2026)
Structured Prompting and Feedback-Guided Reasoning with LLMs for Data Interpretation
by: Rath, Amit
Published: (2025)
by: Rath, Amit
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
by: Linzmayer, Robin, et al.
Published: (2026)
by: Linzmayer, Robin, et al.
Published: (2026)
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
by: Adjovi, Mahounan Pericles, et al.
Published: (2026)
OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
by: Bommarito II, Michael J.
Published: (2025)
by: Bommarito II, Michael J.
Published: (2025)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
by: Poddar, Aheli, et al.
Published: (2025)
by: Poddar, Aheli, et al.
Published: (2025)
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
by: V., Diego Fajardo, et al.
Published: (2025)
by: V., Diego Fajardo, et al.
Published: (2025)
TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments
by: Sakizli, Furkan
Published: (2026)
by: Sakizli, Furkan
Published: (2026)
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
by: Fostiropoulos, Iordanis, et al.
Published: (2026)
by: Fostiropoulos, Iordanis, et al.
Published: (2026)
Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework
by: Liu, Jinyan, et al.
Published: (2025)
by: Liu, Jinyan, et al.
Published: (2025)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
by: Shah, Aaryan, et al.
Published: (2026)
by: Shah, Aaryan, et al.
Published: (2026)
Subjective Question Generation and Answer Evaluation using NLP
by: Islam, G. M. Refatul, et al.
Published: (2025)
by: Islam, G. M. Refatul, et al.
Published: (2025)
DisGeM: Distractor Generation for Multiple Choice Questions with Span Masking
by: Cavusoglu, Devrim, et al.
Published: (2024)
by: Cavusoglu, Devrim, et al.
Published: (2024)
AskSport: Web Application for Sports Question-Answering
by: Onofre, Enzo B, et al.
Published: (2025)
by: Onofre, Enzo B, et al.
Published: (2025)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
by: Dayarathne, Ranul, et al.
Published: (2025)
by: Dayarathne, Ranul, et al.
Published: (2025)
ReTreVal: Reasoning Tree with Validation -- A Hybrid Framework for Enhanced LLM Multi-Step Reasoning
by: HS, Abhishek, et al.
Published: (2026)
by: HS, Abhishek, et al.
Published: (2026)
Challenges and Opportunities of NLP for HR Applications: A Discussion Paper
by: Leidner, Jochen L., et al.
Published: (2024)
by: Leidner, Jochen L., et al.
Published: (2024)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
by: Berijanian, Maryam, et al.
Published: (2025)
by: Berijanian, Maryam, et al.
Published: (2025)
Exploring LLMs for User Story Extraction from Mockups
by: Firmenich, Diego, et al.
Published: (2026)
by: Firmenich, Diego, et al.
Published: (2026)
Evaluating Large Language Models for IUCN Red List Species Information
by: Uryu, Shinya
Published: (2025)
by: Uryu, Shinya
Published: (2025)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
by: Manczak, Blazej, et al.
Published: (2025)
by: Manczak, Blazej, et al.
Published: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
An Industrial-Scale Retrieval-Augmented Generation Framework for Requirements Engineering: Empirical Evaluation with Automotive Manufacturing Data
by: Khalid, Muhammad, et al.
Published: (2026)
by: Khalid, Muhammad, et al.
Published: (2026)
Similar Items
-
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
by: Guo, Dongxin
Published: (2026) -
When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
by: Coen, Michael H.
Published: (2025) -
Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
by: Chelombitko, Iaroslav, et al.
Published: (2024) -
How to Evaluate Medical AI
by: Kopanichuk, Ilia, et al.
Published: (2025) -
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
by: Yan, Lingyong, et al.
Published: (2026)