Can we trust the evaluation on ChatGPT?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910572356829184 |
|---|---|
| author | Aiyappa, Rachith An, Jisun Kwak, Haewoon Ahn, Yong-Yeol |
| author_facet | Aiyappa, Rachith An, Jisun Kwak, Haewoon Ahn, Yong-Yeol |
| contents | ChatGPT, the first large language model (LLM) with mass adoption, has demonstrated remarkable performance in numerous natural language tasks. Despite its evident usefulness, evaluating ChatGPT's performance in diverse problem domains remains challenging due to the closed nature of the model and its continuous updates via Reinforcement Learning from Human Feedback (RLHF). We highlight the issue of data contamination in ChatGPT evaluations, with a case study of the task of stance detection. We discuss the challenge of preventing data contamination and ensuring fair model evaluation in the age of closed and continuously trained models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2303_12767 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Can we trust the evaluation on ChatGPT? Aiyappa, Rachith An, Jisun Kwak, Haewoon Ahn, Yong-Yeol Computation and Language Artificial Intelligence Machine Learning ChatGPT, the first large language model (LLM) with mass adoption, has demonstrated remarkable performance in numerous natural language tasks. Despite its evident usefulness, evaluating ChatGPT's performance in diverse problem domains remains challenging due to the closed nature of the model and its continuous updates via Reinforcement Learning from Human Feedback (RLHF). We highlight the issue of data contamination in ChatGPT evaluations, with a case study of the task of stance detection. We discuss the challenge of preventing data contamination and ensuring fair model evaluation in the age of closed and continuously trained models. |
| title | Can we trust the evaluation on ChatGPT? |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2303.12767 |