Can we trust the evaluation on ChatGPT?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aiyappa, Rachith, An, Jisun, Kwak, Haewoon, Ahn, Yong-Yeol
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910572356829184
author Aiyappa, Rachith
An, Jisun
Kwak, Haewoon
Ahn, Yong-Yeol
author_facet Aiyappa, Rachith
An, Jisun
Kwak, Haewoon
Ahn, Yong-Yeol
contents ChatGPT, the first large language model (LLM) with mass adoption, has demonstrated remarkable performance in numerous natural language tasks. Despite its evident usefulness, evaluating ChatGPT's performance in diverse problem domains remains challenging due to the closed nature of the model and its continuous updates via Reinforcement Learning from Human Feedback (RLHF). We highlight the issue of data contamination in ChatGPT evaluations, with a case study of the task of stance detection. We discuss the challenge of preventing data contamination and ensuring fair model evaluation in the age of closed and continuously trained models.
format Preprint
id arxiv_https___arxiv_org_abs_2303_12767
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Can we trust the evaluation on ChatGPT?
Aiyappa, Rachith
An, Jisun
Kwak, Haewoon
Ahn, Yong-Yeol
Computation and Language
Artificial Intelligence
Machine Learning
ChatGPT, the first large language model (LLM) with mass adoption, has demonstrated remarkable performance in numerous natural language tasks. Despite its evident usefulness, evaluating ChatGPT's performance in diverse problem domains remains challenging due to the closed nature of the model and its continuous updates via Reinforcement Learning from Human Feedback (RLHF). We highlight the issue of data contamination in ChatGPT evaluations, with a case study of the task of stance detection. We discuss the challenge of preventing data contamination and ensuring fair model evaluation in the age of closed and continuously trained models.
title Can we trust the evaluation on ChatGPT?
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2303.12767