Urania: Differentially Private Insights into AI Use

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Daogao, Cohen, Edith, Ghazi, Badih, Kairouz, Peter, Kamath, Pritish, Knop, Alexander, Kumar, Ravi, Manurangsi, Pasin, Sealfon, Adam, Yu, Da, Zhang, Chiyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912602112655360
author Liu, Daogao
Cohen, Edith
Ghazi, Badih
Kairouz, Peter
Kamath, Pritish
Knop, Alexander
Kumar, Ravi
Manurangsi, Pasin
Sealfon, Adam
Yu, Da
Zhang, Chiyuan
author_facet Liu, Daogao
Cohen, Edith
Ghazi, Badih
Kairouz, Peter
Kamath, Pritish
Knop, Alexander
Kumar, Ravi
Manurangsi, Pasin
Sealfon, Adam
Yu, Da
Zhang, Chiyuan
contents We introduce $Urania$, a novel framework for generating insights about LLM chatbot interactions with rigorous differential privacy (DP) guarantees. The framework employs a private clustering mechanism and innovative keyword extraction methods, including frequency-based, TF-IDF-based, and LLM-guided approaches. By leveraging DP tools such as clustering, partition selection, and histogram-based summarization, $Urania$ provides end-to-end privacy protection. Our evaluation assesses lexical and semantic content preservation, pair similarity, and LLM-based metrics, benchmarking against a non-private Clio-inspired pipeline (Tamkin et al., 2024). Moreover, we develop a simple empirical privacy evaluation that demonstrates the enhanced robustness of our DP pipeline. The results show the framework's ability to extract meaningful conversational insights while maintaining stringent user privacy, effectively balancing data utility with privacy preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Urania: Differentially Private Insights into AI Use
Liu, Daogao
Cohen, Edith
Ghazi, Badih
Kairouz, Peter
Kamath, Pritish
Knop, Alexander
Kumar, Ravi
Manurangsi, Pasin
Sealfon, Adam
Yu, Da
Zhang, Chiyuan
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Computers and Society
We introduce $Urania$, a novel framework for generating insights about LLM chatbot interactions with rigorous differential privacy (DP) guarantees. The framework employs a private clustering mechanism and innovative keyword extraction methods, including frequency-based, TF-IDF-based, and LLM-guided approaches. By leveraging DP tools such as clustering, partition selection, and histogram-based summarization, $Urania$ provides end-to-end privacy protection. Our evaluation assesses lexical and semantic content preservation, pair similarity, and LLM-based metrics, benchmarking against a non-private Clio-inspired pipeline (Tamkin et al., 2024). Moreover, we develop a simple empirical privacy evaluation that demonstrates the enhanced robustness of our DP pipeline. The results show the framework's ability to extract meaningful conversational insights while maintaining stringent user privacy, effectively balancing data utility with privacy preservation.
title Urania: Differentially Private Insights into AI Use
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Computers and Society
url https://arxiv.org/abs/2506.04681