LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zheng, Lianmin, Chiang, Wei-Lin, Sheng, Ying, Li, Tianle, Zhuang, Siyuan, Wu, Zhanghao, Zhuang, Yonghao, Li, Zhuohan, Lin, Zi, Xing, Eric P., Gonzalez, Joseph E., Stoica, Ion, Zhang, Hao
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911792130686976
author Zheng, Lianmin
Chiang, Wei-Lin
Sheng, Ying
Li, Tianle
Zhuang, Siyuan
Wu, Zhanghao
Zhuang, Yonghao
Li, Zhuohan
Lin, Zi
Xing, Eric P.
Gonzalez, Joseph E.
Stoica, Ion
Zhang, Hao
author_facet Zheng, Lianmin
Chiang, Wei-Lin
Sheng, Ying
Li, Tianle
Zhuang, Siyuan
Wu, Zhanghao
Zhuang, Yonghao
Li, Zhuohan
Lin, Zi
Xing, Eric P.
Gonzalez, Joseph E.
Stoica, Ion
Zhang, Hao
contents Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11998
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Zheng, Lianmin
Chiang, Wei-Lin
Sheng, Ying
Li, Tianle
Zhuang, Siyuan
Wu, Zhanghao
Zhuang, Yonghao
Li, Zhuohan
Lin, Zi
Xing, Eric P.
Gonzalez, Joseph E.
Stoica, Ion
Zhang, Hao
Computation and Language
Artificial Intelligence
Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.
title LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2309.11998