YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913865770467328 |
|---|---|
| author | D'Souza, Jennifer Giglou, Hamed Babaei Münch, Quentin |
| author_facet | D'Souza, Jennifer Giglou, Hamed Babaei Münch, Quentin |
| contents | Large Language Models (LLMs) drive scientific question-answering on modern search engines, yet their evaluation robustness remains underexplored. We introduce YESciEval, an open-source framework that combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators. We release multidisciplinary scienceQ&A datasets, including adversarial variants, with evaluation scores from multiple LLMs. Independent of proprietary models and human feedback, our approach enables scalable, cost-free evaluation. By advancing reliable LLM-as-a-judge models, this work supports AI alignment and fosters robust, transparent evaluation essential for scientific inquiry. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14279 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering D'Souza, Jennifer Giglou, Hamed Babaei Münch, Quentin Computation and Language Artificial Intelligence Large Language Models (LLMs) drive scientific question-answering on modern search engines, yet their evaluation robustness remains underexplored. We introduce YESciEval, an open-source framework that combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators. We release multidisciplinary scienceQ&A datasets, including adversarial variants, with evaluation scores from multiple LLMs. Independent of proprietary models and human feedback, our approach enables scalable, cost-free evaluation. By advancing reliable LLM-as-a-judge models, this work supports AI alignment and fosters robust, transparent evaluation essential for scientific inquiry. |
| title | YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.14279 |