Consent in Crisis: The Rapid Decline of the AI Data Commons
Fuente:
arXiv
Saved in:
| Main Authors: | Longpre, Shayne, Mahari, Robert, Lee, Ariel, Lund, Campbell, Oderinwale, Hamidah, Brannon, William, Saxena, Nayan, Obeng-Marnu, Naana, South, Tobin, Hunter, Cole, Klyman, Kevin, Klamm, Christopher, Schoelkopf, Hailey, Singh, Nikhil, Cherep, Manuel, Anis, Ahmad, Dinh, An, Chitongo, Caroline, Yin, Da, Sileo, Damien, Mataciunas, Deividas, Misra, Diganta, Alghamdi, Emad, Shippole, Enrico, Zhang, Jianguo, Materzynska, Joanna, Qian, Kun, Tiwary, Kush, Miranda, Lester, Dey, Manan, Liang, Minnie, Hamdy, Mohammed, Muennighoff, Niklas, Ye, Seonghyeon, Kim, Seungone, Mohanty, Shrestha, Gupta, Vipul, Sharma, Vivek, Chien, Vu Minh, Zhou, Xuhui, Li, Yizhi, Xiong, Caiming, Villa, Luis, Biderman, Stella, Li, Hanlin, Ippolito, Daphne, Hooker, Sara, Kabbara, Jad, Pentland, Sandy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?
by: Longpre, Shayne, et al.
Published: (2024)
by: Longpre, Shayne, et al.
Published: (2024)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
by: Oderinwale, Hamidah, et al.
Published: (2024)
by: Oderinwale, Hamidah, et al.
Published: (2024)
Bridging the Data Provenance Gap Across Text, Speech and Video
by: Longpre, Shayne, et al.
Published: (2024)
by: Longpre, Shayne, et al.
Published: (2024)
Procedural Knowledge Libraries: Towards Executable (Research) Memory
by: Oderinwale, Hamidah
Published: (2025)
by: Oderinwale, Hamidah
Published: (2025)
The Economics of AI Training Data: A Research Agenda
by: Oderinwale, Hamidah, et al.
Published: (2025)
by: Oderinwale, Hamidah, et al.
Published: (2025)
Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
by: Laufer, Benjamin, et al.
Published: (2025)
by: Laufer, Benjamin, et al.
Published: (2025)
LePaRD: A Large-Scale Dataset of Judges Citing Precedents
by: Mahari, Robert, et al.
Published: (2023)
by: Mahari, Robert, et al.
Published: (2023)
Insights from an experiment crowdsourcing data from thousands of US Amazon users: The importance of transparency, money, and data use
by: Berke, Alex, et al.
Published: (2024)
by: Berke, Alex, et al.
Published: (2024)
zkTax: A pragmatic way to support zero-knowledge tax disclosures
by: Berke, Alex, et al.
Published: (2023)
by: Berke, Alex, et al.
Published: (2023)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Building trust in luxury brands through behavioral analytics of customer experience
by: Cherep, Nataliia
Published: (2025)
by: Cherep, Nataliia
Published: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling
by: Jiang, Hang, et al.
Published: (2024)
by: Jiang, Hang, et al.
Published: (2024)
Chapter 2 Supplementary File
by: Ghani, Hamidah
Published: (2025)
by: Ghani, Hamidah
Published: (2025)
The Other, the New, and the Old. Three Images of Leith in Irvine Welsh’s Porno
by: Deividas Zibalas
Published: (2021)
by: Deividas Zibalas
Published: (2021)
Los impulsos en la concepción materialista de la razón de Max Horkheimer
by: Paula García Cherep
Published: (2021)
by: Paula García Cherep
Published: (2021)
Suppressing Pink Elephants with Direct Principle Feedback
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
Community-centric modeling of citation dynamics explains collective citation patterns in science, law, and patents
by: Kojaku, Sadamori, et al.
Published: (2025)
by: Kojaku, Sadamori, et al.
Published: (2025)
ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings
by: Brannon, William, et al.
Published: (2023)
by: Brannon, William, et al.
Published: (2023)
Contrastive Learning from Synthetic Audio Doppelgängers
by: Cherep, Manuel, et al.
Published: (2024)
by: Cherep, Manuel, et al.
Published: (2024)
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Bibliographic Instruction for Freshman Students at Florida International University.
by: Dunbar, H. Minnie
Published: (1986)
by: Dunbar, H. Minnie
Published: (1986)
Acceptable Use Policies for Foundation Models
by: Klyman, Kevin
Published: (2024)
by: Klyman, Kevin
Published: (2024)
On the Relationship between Truth and Political Bias in Language Models
by: Fulay, Suyash, et al.
Published: (2024)
by: Fulay, Suyash, et al.
Published: (2024)
Brunei Darussalam: Problems encountered in data compilation
by: Haji Ladis, Hajah Hamidah
Published: (1997)
by: Haji Ladis, Hajah Hamidah
Published: (1997)
The Service Implications of a Rhetorical Approach to Information Literacy
by: Brannon, Brittany
Published: (2017)
by: Brannon, Brittany
Published: (2017)
Examining the Fieldwork Experience from the Site Supervisor Perspective: A Mixed-Methods Study Using Vygotsky's Zone of Proximal Development Theory
by: Brannon, Sian
Published: (2013)
by: Brannon, Sian
Published: (2013)
Assessment in Fieldwork Courses: What Are We Rating?
by: Brannon, Sian
Published: (2014)
by: Brannon, Sian
Published: (2014)
Say No to Speed Bumps!
by: Brannon, Sian
Published: (2010)
by: Brannon, Sian
Published: (2010)
Intentionality is a Design Decision: Measuring Functional Intentionality for Accountable AI Systems
by: Chiappetta, Allessia, et al.
Published: (2026)
by: Chiappetta, Allessia, et al.
Published: (2026)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
by: Sileo, Damien
Published: (2025)
by: Sileo, Damien
Published: (2025)
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
by: Sileo, Damien
Published: (2024)
by: Sileo, Damien
Published: (2024)
Discurso de aceptación de la orden Gustavo Machado
by: Enriqueta Sileo
Published: (2015)
by: Enriqueta Sileo
Published: (2015)
Handbook for Louisiana Library Trustees.
by: Lynch, Minnie-Lou, Ed.
Published: (1980)
by: Lynch, Minnie-Lou, Ed.
Published: (1980)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
The 2025 Foundation Model Transparency Index
by: Wan, Alexander, et al.
Published: (2025)
by: Wan, Alexander, et al.
Published: (2025)
Na ante-sala da discriminação: o preço dos atributos de sexo ecor no Brasil (19891999)
by: Ciro Biderman
Published: (2004)
by: Ciro Biderman
Published: (2004)
Verifiable evaluations of machine learning models using zkSNARKs
by: South, Tobin, et al.
Published: (2024)
by: South, Tobin, et al.
Published: (2024)
Similar Items
-
Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?
by: Longpre, Shayne, et al.
Published: (2024) -
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
by: Oderinwale, Hamidah, et al.
Published: (2024) -
Bridging the Data Provenance Gap Across Text, Speech and Video
by: Longpre, Shayne, et al.
Published: (2024) -
Procedural Knowledge Libraries: Towards Executable (Research) Memory
by: Oderinwale, Hamidah
Published: (2025) -
The Economics of AI Training Data: A Research Agenda
by: Oderinwale, Hamidah, et al.
Published: (2025)