LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915833376145408 |
|---|---|
| author | Saad-Falcon, Jon Vivek, Rajan Berrios, William Naik, Nandita Shankar Franklin, Matija Vidgen, Bertie Singh, Amanpreet Kiela, Douwe Mehri, Shikib |
| author_facet | Saad-Falcon, Jon Vivek, Rajan Berrios, William Naik, Nandita Shankar Franklin, Matija Vidgen, Bertie Singh, Amanpreet Kiela, Douwe Mehri, Shikib |
| contents | As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics provide only coarse, difficult-to-interpret signals. We introduce natural language unit tests, a paradigm that decomposes response quality into explicit, testable criteria, along with a unified scoring model, LMUnit, which combines multi-objective training across preferences, direct ratings, and natural language rationales. Through controlled human studies, we show this paradigm significantly improves inter-annotator agreement and enables more effective LLM development workflows. LMUnit achieves state-of-the-art performance on evaluation benchmarks (FLASK, BigGenBench) and competitive results on RewardBench. These results validate both our proposed paradigm and scoring model, suggesting a promising path forward for language model evaluation and development. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_13091 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LMUnit: Fine-grained Evaluation with Natural Language Unit Tests Saad-Falcon, Jon Vivek, Rajan Berrios, William Naik, Nandita Shankar Franklin, Matija Vidgen, Bertie Singh, Amanpreet Kiela, Douwe Mehri, Shikib Computation and Language Artificial Intelligence As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics provide only coarse, difficult-to-interpret signals. We introduce natural language unit tests, a paradigm that decomposes response quality into explicit, testable criteria, along with a unified scoring model, LMUnit, which combines multi-objective training across preferences, direct ratings, and natural language rationales. Through controlled human studies, we show this paradigm significantly improves inter-annotator agreement and enables more effective LLM development workflows. LMUnit achieves state-of-the-art performance on evaluation benchmarks (FLASK, BigGenBench) and competitive results on RewardBench. These results validate both our proposed paradigm and scoring model, suggesting a promising path forward for language model evaluation and development. |
| title | LMUnit: Fine-grained Evaluation with Natural Language Unit Tests |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2412.13091 |