A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Rachel, Hutson, Jevan, Agnew, William, Huda, Imaad, Kohno, Tadayoshi, Morgenstern, Jamie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
by: Lee, Chung Peng, et al.
Published: (2025)
by: Lee, Chung Peng, et al.
Published: (2025)
Unencrypted Flying Objects: Security Lessons from University Small Satellite Developers and Their Code
by: McAmis, Rachel, et al.
Published: (2025)
by: McAmis, Rachel, et al.
Published: (2025)
Attacking the Diebold Signature Variant -- RSA Signatures with Unverified High-order Padding
by: Gardner, Ryan W., et al.
Published: (2024)
by: Gardner, Ryan W., et al.
Published: (2024)
[Extended] Ethics in Computer Security Research: A Data-Driven Assessment of the Past, the Present, and the Possible Future
by: Ramulu, Harshini Sri, et al.
Published: (2025)
by: Ramulu, Harshini Sri, et al.
Published: (2025)
Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
by: Hong, Rachel, et al.
Published: (2024)
by: Hong, Rachel, et al.
Published: (2024)
Extending the Formalism and Theoretical Foundations of Cryptography to AI
by: Villa, Federico, et al.
Published: (2026)
by: Villa, Federico, et al.
Published: (2026)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
by: Iqbal, Umar, et al.
Published: (2023)
by: Iqbal, Umar, et al.
Published: (2023)
IDCloak: A Practical Secure Multi-party Dataset Join Framework for Vertical Privacy-preserving Machine Learning
by: Chen, Shuyu, et al.
Published: (2025)
by: Chen, Shuyu, et al.
Published: (2025)
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies
by: Srinath, Mukund, et al.
Published: (2020)
by: Srinath, Mukund, et al.
Published: (2020)
Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users
by: Bhardwaj, Arth, et al.
Published: (2026)
by: Bhardwaj, Arth, et al.
Published: (2026)
Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable
by: Bertran, Martin, et al.
Published: (2024)
by: Bertran, Martin, et al.
Published: (2024)
BinPool: A Dataset of Vulnerabilities for Binary Security Analysis
by: Arasteh, Sima, et al.
Published: (2025)
by: Arasteh, Sima, et al.
Published: (2025)
algoXSSF: Detection and analysis of cross-site request forgery (XSRF) and cross-site scripting (XSS) attacks via Machine learning algorithms
by: Kshetri, Naresh, et al.
Published: (2024)
by: Kshetri, Naresh, et al.
Published: (2024)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Indifferential Privacy: A New Paradigm and Its Applications to Optimal Matching in Dark Pool Auctions
by: Polychroniadou, Antigoni, et al.
Published: (2025)
by: Polychroniadou, Antigoni, et al.
Published: (2025)
Unveiling Usability Challenges in Web Privacy Controls
by: Masood, Rahat, et al.
Published: (2025)
by: Masood, Rahat, et al.
Published: (2025)
Understanding Privacy Norms through Web Forms
by: Cui, Hao, et al.
Published: (2024)
by: Cui, Hao, et al.
Published: (2024)
"Sign in with ... Privacy'': Timely Disclosure of Privacy Differences among Web SSO Login Options
by: Morkonda, Srivathsan G., et al.
Published: (2022)
by: Morkonda, Srivathsan G., et al.
Published: (2022)
Distortion Search, A Web Search Privacy Heuristic
by: Mivule, Kato, et al.
Published: (2025)
by: Mivule, Kato, et al.
Published: (2025)
SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors
by: Wei, Miranda, et al.
Published: (2024)
by: Wei, Miranda, et al.
Published: (2024)
FiberPool: Leveraging Multiple Blockchains for Decentralized Pooled Mining
by: Sakurai, Akira, et al.
Published: (2025)
by: Sakurai, Akira, et al.
Published: (2025)
Technical Report: The Need for a (Research) Sandstorm through the Privacy Sandbox
by: Beugin, Yohan, et al.
Published: (2025)
by: Beugin, Yohan, et al.
Published: (2025)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Revisiting Privacy Leakage in Machine Unlearning: Membership Inference Beyond the Forgotten Set
by: Fu, Jie, et al.
Published: (2026)
by: Fu, Jie, et al.
Published: (2026)
Big Bird: Resilient Privacy Budgeting Across Untrusted Web Domains
by: Tholoniat, Pierre, et al.
Published: (2025)
by: Tholoniat, Pierre, et al.
Published: (2025)
Privacy-Enhanced Database Synthesis for Benchmark Publishing (Technical Report)
by: Ge, Yunqing, et al.
Published: (2024)
by: Ge, Yunqing, et al.
Published: (2024)
Privacy-Preserving Logistic Regression Training on Large Datasets
by: Chiang, John
Published: (2024)
by: Chiang, John
Published: (2024)
Towards Automating Data Access Permissions in AI Agents
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
Global Web, Local Privacy? An International Review of Web Tracking
by: Yu, Harry, et al.
Published: (2026)
by: Yu, Harry, et al.
Published: (2026)
When Machine Learning Meets Vulnerability Discovery: Challenges and Lessons Learned
by: Arasteh, Sima, et al.
Published: (2025)
by: Arasteh, Sima, et al.
Published: (2025)
Efficient Privacy-Preserving Approximation of the Kidney Exchange Problem
by: Breuer, Malte, et al.
Published: (2023)
by: Breuer, Malte, et al.
Published: (2023)
Towards Privacy-Preserving LLM Inference via Covariant Obfuscation (Technical Report)
by: Lin, Yu, et al.
Published: (2026)
by: Lin, Yu, et al.
Published: (2026)
Privacy-Aware White and Black List Searching for Fraud Analysis
by: Buchanan, William J, et al.
Published: (2025)
by: Buchanan, William J, et al.
Published: (2025)
A Unified Framework for Adversary-Aware Differential Privacy Bounds
by: Swanberg, Marika, et al.
Published: (2025)
by: Swanberg, Marika, et al.
Published: (2025)
Measuring Privacy vs. Fidelity in Synthetic Social Media Datasets
by: Tari, Henry, et al.
Published: (2026)
by: Tari, Henry, et al.
Published: (2026)
From Blocking to Breaking: Evaluating the Impact of Adblockers on Web Usability
by: Roongta, Ritik, et al.
Published: (2024)
by: Roongta, Ritik, et al.
Published: (2024)
Memory Scraping Attack on Xilinx FPGAs: Private Data Extraction from Terminated Processes
by: Madabhushi, Bharadwaj, et al.
Published: (2024)
by: Madabhushi, Bharadwaj, et al.
Published: (2024)
The Evolution Of The Digital Inheritance: Legal, Technical, And Practical Dimensions Of Cryptocurrency Transfer Through Succession In French-Inspired Legal Systems
by: Carata, Cristina, et al.
Published: (2024)
by: Carata, Cristina, et al.
Published: (2024)
The Illusion of Anonymity: Uncovering the Impact of User Actions on Privacy in Web3 Social Ecosystems
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Advances in Differential Privacy and Differentially Private Machine Learning
by: Das, Saswat, et al.
Published: (2024)
by: Das, Saswat, et al.
Published: (2024)
Similar Items
-
How Do Data Owners Say No? A Case Study of Data Consent Mechanisms in Web-Scraped Vision-Language AI Training Datasets
by: Lee, Chung Peng, et al.
Published: (2025) -
Unencrypted Flying Objects: Security Lessons from University Small Satellite Developers and Their Code
by: McAmis, Rachel, et al.
Published: (2025) -
Attacking the Diebold Signature Variant -- RSA Signatures with Unverified High-order Padding
by: Gardner, Ryan W., et al.
Published: (2024) -
[Extended] Ethics in Computer Security Research: A Data-Driven Assessment of the Past, the Present, and the Possible Future
by: Ramulu, Harshini Sri, et al.
Published: (2025) -
Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
by: Hong, Rachel, et al.
Published: (2024)