Jury: A Comprehensive Evaluation Toolkit
Fuente:
arXiv
Salvato in:
| Autori principali: | Cavusoglu, Devrim, Sen, Secil, Sert, Ulas, Altinuc, Sinan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
di: Guo, Dongxin
Pubblicazione: (2026)
di: Guo, Dongxin
Pubblicazione: (2026)
When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
di: Coen, Michael H.
Pubblicazione: (2025)
di: Coen, Michael H.
Pubblicazione: (2025)
Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
di: Chelombitko, Iaroslav, et al.
Pubblicazione: (2024)
di: Chelombitko, Iaroslav, et al.
Pubblicazione: (2024)
How to Evaluate Medical AI
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025)
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
di: Yan, Lingyong, et al.
Pubblicazione: (2026)
di: Yan, Lingyong, et al.
Pubblicazione: (2026)
Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation
di: Wu, Jikun, et al.
Pubblicazione: (2026)
di: Wu, Jikun, et al.
Pubblicazione: (2026)
When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
di: Pan, Xinghan
Pubblicazione: (2025)
di: Pan, Xinghan
Pubblicazione: (2025)
OEMA: Ontology-Enhanced Multi-Agent Collaboration Framework for Zero-Shot Clinical Named Entity Recognition
di: Tao, Xinli, et al.
Pubblicazione: (2025)
di: Tao, Xinli, et al.
Pubblicazione: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
di: Rodriguez, David, et al.
Pubblicazione: (2025)
di: Rodriguez, David, et al.
Pubblicazione: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
di: Michail, Andrianos, et al.
Pubblicazione: (2024)
di: Michail, Andrianos, et al.
Pubblicazione: (2024)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
di: Hasan, Mohammed Rakibul
Pubblicazione: (2026)
di: Hasan, Mohammed Rakibul
Pubblicazione: (2026)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
di: Weng, Haojun, et al.
Pubblicazione: (2026)
di: Weng, Haojun, et al.
Pubblicazione: (2026)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
di: Loi, Dario, et al.
Pubblicazione: (2025)
di: Loi, Dario, et al.
Pubblicazione: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
NL2LOGIC: AST-Guided Translation of Natural Language into First-Order Logic with Large Language Models
di: Putra, Rizky Ramadhana, et al.
Pubblicazione: (2026)
di: Putra, Rizky Ramadhana, et al.
Pubblicazione: (2026)
Structured Prompting and Feedback-Guided Reasoning with LLMs for Data Interpretation
di: Rath, Amit
Pubblicazione: (2025)
di: Rath, Amit
Pubblicazione: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
di: Linzmayer, Robin, et al.
Pubblicazione: (2026)
di: Linzmayer, Robin, et al.
Pubblicazione: (2026)
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph
di: Bommarito II, Michael J.
Pubblicazione: (2025)
di: Bommarito II, Michael J.
Pubblicazione: (2025)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
di: Poddar, Aheli, et al.
Pubblicazione: (2025)
di: Poddar, Aheli, et al.
Pubblicazione: (2025)
MedPI: Evaluating AI Systems in Medical Patient-facing Interactions
di: V., Diego Fajardo, et al.
Pubblicazione: (2025)
di: V., Diego Fajardo, et al.
Pubblicazione: (2025)
TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments
di: Sakizli, Furkan
Pubblicazione: (2026)
di: Sakizli, Furkan
Pubblicazione: (2026)
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
di: Fostiropoulos, Iordanis, et al.
Pubblicazione: (2026)
di: Fostiropoulos, Iordanis, et al.
Pubblicazione: (2026)
Evaluating the Challenges of LLMs in Real-world Medical Follow-up: A Comparative Study and An Optimized Framework
di: Liu, Jinyan, et al.
Pubblicazione: (2025)
di: Liu, Jinyan, et al.
Pubblicazione: (2025)
Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters
di: Shah, Aaryan, et al.
Pubblicazione: (2026)
di: Shah, Aaryan, et al.
Pubblicazione: (2026)
Subjective Question Generation and Answer Evaluation using NLP
di: Islam, G. M. Refatul, et al.
Pubblicazione: (2025)
di: Islam, G. M. Refatul, et al.
Pubblicazione: (2025)
DisGeM: Distractor Generation for Multiple Choice Questions with Span Masking
di: Cavusoglu, Devrim, et al.
Pubblicazione: (2024)
di: Cavusoglu, Devrim, et al.
Pubblicazione: (2024)
AskSport: Web Application for Sports Question-Answering
di: Onofre, Enzo B, et al.
Pubblicazione: (2025)
di: Onofre, Enzo B, et al.
Pubblicazione: (2025)
Comparing the Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature
di: Dayarathne, Ranul, et al.
Pubblicazione: (2025)
di: Dayarathne, Ranul, et al.
Pubblicazione: (2025)
ReTreVal: Reasoning Tree with Validation -- A Hybrid Framework for Enhanced LLM Multi-Step Reasoning
di: HS, Abhishek, et al.
Pubblicazione: (2026)
di: HS, Abhishek, et al.
Pubblicazione: (2026)
Challenges and Opportunities of NLP for HR Applications: A Discussion Paper
di: Leidner, Jochen L., et al.
Pubblicazione: (2024)
di: Leidner, Jochen L., et al.
Pubblicazione: (2024)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
di: Berijanian, Maryam, et al.
Pubblicazione: (2025)
di: Berijanian, Maryam, et al.
Pubblicazione: (2025)
Exploring LLMs for User Story Extraction from Mockups
di: Firmenich, Diego, et al.
Pubblicazione: (2026)
di: Firmenich, Diego, et al.
Pubblicazione: (2026)
Evaluating Large Language Models for IUCN Red List Species Information
di: Uryu, Shinya
Pubblicazione: (2025)
di: Uryu, Shinya
Pubblicazione: (2025)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
di: Manczak, Blazej, et al.
Pubblicazione: (2025)
di: Manczak, Blazej, et al.
Pubblicazione: (2025)
BLT: Can Large Language Models Handle Basic Legal Text?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
An Industrial-Scale Retrieval-Augmented Generation Framework for Requirements Engineering: Empirical Evaluation with Automotive Manufacturing Data
di: Khalid, Muhammad, et al.
Pubblicazione: (2026)
di: Khalid, Muhammad, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
di: Guo, Dongxin
Pubblicazione: (2026) -
When F1 Fails: Granularity-Aware Evaluation for Dialogue Topic Segmentation
di: Coen, Michael H.
Pubblicazione: (2025) -
Qtok: A Comprehensive Framework for Evaluating Multilingual Tokenizer Quality in Large Language Models
di: Chelombitko, Iaroslav, et al.
Pubblicazione: (2024) -
How to Evaluate Medical AI
di: Kopanichuk, Ilia, et al.
Pubblicazione: (2025) -
Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation
di: Yan, Lingyong, et al.
Pubblicazione: (2026)