Urania: Differentially Private Insights into AI Use
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912602112655360 |
|---|---|
| author | Liu, Daogao Cohen, Edith Ghazi, Badih Kairouz, Peter Kamath, Pritish Knop, Alexander Kumar, Ravi Manurangsi, Pasin Sealfon, Adam Yu, Da Zhang, Chiyuan |
| author_facet | Liu, Daogao Cohen, Edith Ghazi, Badih Kairouz, Peter Kamath, Pritish Knop, Alexander Kumar, Ravi Manurangsi, Pasin Sealfon, Adam Yu, Da Zhang, Chiyuan |
| contents | We introduce $Urania$, a novel framework for generating insights about LLM chatbot interactions with rigorous differential privacy (DP) guarantees. The framework employs a private clustering mechanism and innovative keyword extraction methods, including frequency-based, TF-IDF-based, and LLM-guided approaches. By leveraging DP tools such as clustering, partition selection, and histogram-based summarization, $Urania$ provides end-to-end privacy protection. Our evaluation assesses lexical and semantic content preservation, pair similarity, and LLM-based metrics, benchmarking against a non-private Clio-inspired pipeline (Tamkin et al., 2024). Moreover, we develop a simple empirical privacy evaluation that demonstrates the enhanced robustness of our DP pipeline. The results show the framework's ability to extract meaningful conversational insights while maintaining stringent user privacy, effectively balancing data utility with privacy preservation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_04681 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Urania: Differentially Private Insights into AI Use Liu, Daogao Cohen, Edith Ghazi, Badih Kairouz, Peter Kamath, Pritish Knop, Alexander Kumar, Ravi Manurangsi, Pasin Sealfon, Adam Yu, Da Zhang, Chiyuan Machine Learning Artificial Intelligence Computation and Language Cryptography and Security Computers and Society We introduce $Urania$, a novel framework for generating insights about LLM chatbot interactions with rigorous differential privacy (DP) guarantees. The framework employs a private clustering mechanism and innovative keyword extraction methods, including frequency-based, TF-IDF-based, and LLM-guided approaches. By leveraging DP tools such as clustering, partition selection, and histogram-based summarization, $Urania$ provides end-to-end privacy protection. Our evaluation assesses lexical and semantic content preservation, pair similarity, and LLM-based metrics, benchmarking against a non-private Clio-inspired pipeline (Tamkin et al., 2024). Moreover, we develop a simple empirical privacy evaluation that demonstrates the enhanced robustness of our DP pipeline. The results show the framework's ability to extract meaningful conversational insights while maintaining stringent user privacy, effectively balancing data utility with privacy preservation. |
| title | Urania: Differentially Private Insights into AI Use |
| topic | Machine Learning Artificial Intelligence Computation and Language Cryptography and Security Computers and Society |
| url | https://arxiv.org/abs/2506.04681 |