User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: King, Jennifer, Klyman, Kevin, Capstick, Emily, Saade, Tiffany, Hsieh, Victoria
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912573943709696
author King, Jennifer
Klyman, Kevin
Capstick, Emily
Saade, Tiffany
Hsieh, Victoria
author_facet King, Jennifer
Klyman, Kevin
Capstick, Emily
Saade, Tiffany
Hsieh, Victoria
contents Hundreds of millions of people now regularly interact with large language models via chatbots. Model developers are eager to acquire new sources of high-quality training data as they race to improve model capabilities and win market share. This paper analyzes the privacy policies of six U.S. frontier AI developers to understand how they use their users' chats to train models. Drawing primarily on the California Consumer Privacy Act, we develop a novel qualitative coding schema that we apply to each developer's relevant privacy policies to compare data collection and use practices across the six companies. We find that all six developers appear to employ their users' chat data to train and improve their models by default, and that some retain this data indefinitely. Developers may collect and train on personal information disclosed in chats, including sensitive information such as biometric and health data, as well as files uploaded by users. Four of the six companies we examined appear to include children's chat data for model training, as well as customer data from other products. On the whole, developers' privacy policies often lack essential information about their practices, highlighting the need for greater transparency and accountability. We address the implications of users' lack of consent for the use of their chat data for model training, data security issues arising from indefinite chat data retention, and training on children's chat data. We conclude by providing recommendations to policymakers and developers to address the data privacy challenges posed by LLM-powered chatbots.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05382
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
King, Jennifer
Klyman, Kevin
Capstick, Emily
Saade, Tiffany
Hsieh, Victoria
Computers and Society
Artificial Intelligence
Cryptography and Security
Hundreds of millions of people now regularly interact with large language models via chatbots. Model developers are eager to acquire new sources of high-quality training data as they race to improve model capabilities and win market share. This paper analyzes the privacy policies of six U.S. frontier AI developers to understand how they use their users' chats to train models. Drawing primarily on the California Consumer Privacy Act, we develop a novel qualitative coding schema that we apply to each developer's relevant privacy policies to compare data collection and use practices across the six companies. We find that all six developers appear to employ their users' chat data to train and improve their models by default, and that some retain this data indefinitely. Developers may collect and train on personal information disclosed in chats, including sensitive information such as biometric and health data, as well as files uploaded by users. Four of the six companies we examined appear to include children's chat data for model training, as well as customer data from other products. On the whole, developers' privacy policies often lack essential information about their practices, highlighting the need for greater transparency and accountability. We address the implications of users' lack of consent for the use of their chat data for model training, data security issues arising from indefinite chat data retention, and training on children's chat data. We conclude by providing recommendations to policymakers and developers to address the data privacy challenges posed by LLM-powered chatbots.
title User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
topic Computers and Society
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2509.05382