ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tu, Shangqing, Li, Chunyang, Yu, Jifan, Wang, Xiaozhi, Hou, Lei, Li, Juanzi
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913393761320960
author Tu, Shangqing
Li, Chunyang
Yu, Jifan
Wang, Xiaozhi
Hou, Lei
Li, Juanzi
author_facet Tu, Shangqing
Li, Chunyang
Yu, Jifan
Wang, Xiaozhi
Hou, Lei
Li, Juanzi
contents ChatGPT has achieved great success and can be considered to have acquired an infrastructural status. There are abundant works for evaluating ChatGPT on benchmarks. However, existing benchmarks encounter two challenges: (1) Disregard for periodical evaluation and (2) Lack of fine-grained features. In this paper, we construct ChatLog, an ever-updating dataset with large-scale records of diverse long-form ChatGPT responses for 21 NLP benchmarks from March, 2023 to now. We conduct a comprehensive performance evaluation to find that most capabilities of ChatGPT improve over time except for some abilities, and there exists a step-wise evolving pattern of ChatGPT. We further analyze the inherent characteristics of ChatGPT by extracting the knowledge and linguistic features. We find some stable features that stay unchanged and apply them on the detection of ChatGPT-generated texts to improve the robustness of cross-version detection. We will continuously maintain our project at \url{https://github.com/THU-KEG/ChatLog/}.
format Preprint
id arxiv_https___arxiv_org_abs_2304_14106
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
Tu, Shangqing
Li, Chunyang
Yu, Jifan
Wang, Xiaozhi
Hou, Lei
Li, Juanzi
Computation and Language
Artificial Intelligence
ChatGPT has achieved great success and can be considered to have acquired an infrastructural status. There are abundant works for evaluating ChatGPT on benchmarks. However, existing benchmarks encounter two challenges: (1) Disregard for periodical evaluation and (2) Lack of fine-grained features. In this paper, we construct ChatLog, an ever-updating dataset with large-scale records of diverse long-form ChatGPT responses for 21 NLP benchmarks from March, 2023 to now. We conduct a comprehensive performance evaluation to find that most capabilities of ChatGPT improve over time except for some abilities, and there exists a step-wise evolving pattern of ChatGPT. We further analyze the inherent characteristics of ChatGPT by extracting the knowledge and linguistic features. We find some stable features that stay unchanged and apply them on the detection of ChatGPT-generated texts to improve the robustness of cross-version detection. We will continuously maintain our project at \url{https://github.com/THU-KEG/ChatLog/}.
title ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2304.14106