Evalverse: Unified and Accessible Library for Large Language Model Evaluation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Jihoo, Song, Wonho, Kim, Dahyun, Kim, Yunsu, Kim, Yungi, Park, Chanjun
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912061583261696
author Kim, Jihoo
Song, Wonho
Kim, Dahyun
Kim, Yunsu
Kim, Yungi
Park, Chanjun
author_facet Kim, Jihoo
Song, Wonho
Kim, Dahyun
Kim, Yunsu
Kim, Yungi
Park, Chanjun
contents This paper introduces Evalverse, a novel library that streamlines the evaluation of Large Language Models (LLMs) by unifying disparate evaluation tools into a single, user-friendly framework. Evalverse enables individuals with limited knowledge of artificial intelligence to easily request LLM evaluations and receive detailed reports, facilitated by an integration with communication platforms like Slack. Thus, Evalverse serves as a powerful tool for the comprehensive assessment of LLMs, offering both researchers and practitioners a centralized and easily accessible evaluation framework. Finally, we also provide a demo video for Evalverse, showcasing its capabilities and implementation in a two-minute format.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00943
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evalverse: Unified and Accessible Library for Large Language Model Evaluation
Kim, Jihoo
Song, Wonho
Kim, Dahyun
Kim, Yunsu
Kim, Yungi
Park, Chanjun
Computation and Language
Artificial Intelligence
This paper introduces Evalverse, a novel library that streamlines the evaluation of Large Language Models (LLMs) by unifying disparate evaluation tools into a single, user-friendly framework. Evalverse enables individuals with limited knowledge of artificial intelligence to easily request LLM evaluations and receive detailed reports, facilitated by an integration with communication platforms like Slack. Thus, Evalverse serves as a powerful tool for the comprehensive assessment of LLMs, offering both researchers and practitioners a centralized and easily accessible evaluation framework. Finally, we also provide a demo video for Evalverse, showcasing its capabilities and implementation in a two-minute format.
title Evalverse: Unified and Accessible Library for Large Language Model Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.00943