ShareChat: A Dataset of Chatbot Conversations in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Yueru, Nguyen, Tuc, Su, Bo, Lieffers, Melissa, Le, Thai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914573401980928
author Yan, Yueru
Nguyen, Tuc
Su, Bo
Lieffers, Melissa
Le, Thai
author_facet Yan, Yueru
Nguyen, Tuc
Su, Bo
Lieffers, Melissa
Le, Thai
contents By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system performance. To bridge this gap, we present ShareChat, the first large-scale corpus of 142,808 conversations (660,293 turns) collected from publicly shared URLs on ChatGPT, Perplexity, Grok, Gemini, and Claude. ShareChat preserves native platform affordances, including citations, thinking traces, and code artifacts, across 95 languages and the period from April 2023 to October 2025, complementing existing corpora that homogenize these interactions. To demonstrate the dataset's evaluative utility, we present three case studies: a conversation completeness analysis assessing cross-platform differences in intent satisfaction, a source grounding analysis comparing citation strategies between search-augmented systems, and a temporal analysis revealing divergent response latency dynamics. Together, these analyses demonstrate research questions that are inaccessible to single-platform or stripped-affordance corpora. The dataset is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShareChat: A Dataset of Chatbot Conversations in the Wild
Yan, Yueru
Nguyen, Tuc
Su, Bo
Lieffers, Melissa
Le, Thai
Computation and Language
Artificial Intelligence
Human-Computer Interaction
By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system performance. To bridge this gap, we present ShareChat, the first large-scale corpus of 142,808 conversations (660,293 turns) collected from publicly shared URLs on ChatGPT, Perplexity, Grok, Gemini, and Claude. ShareChat preserves native platform affordances, including citations, thinking traces, and code artifacts, across 95 languages and the period from April 2023 to October 2025, complementing existing corpora that homogenize these interactions. To demonstrate the dataset's evaluative utility, we present three case studies: a conversation completeness analysis assessing cross-platform differences in intent satisfaction, a source grounding analysis comparing citation strategies between search-augmented systems, and a temporal analysis revealing divergent response latency dynamics. Together, these analyses demonstrate research questions that are inaccessible to single-platform or stripped-affordance corpora. The dataset is publicly available.
title ShareChat: A Dataset of Chatbot Conversations in the Wild
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2512.17843