Extracting memorized pieces of (copyrighted) books from open-weight language models
Fuente:
arXiv
Saved in:
| Main Authors: | Cooper, A. Feder, Lemley, Mark A., Casasola, Allison, Ahmed, Ahmed, Gokaslan, Aaron, Cyphert, Amy B., De Sa, Christopher, Ho, Daniel E., Liang, Percy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
Estimating near-verbatim extraction risk in language models with decoding-constrained beam search
by: Cooper, A. Feder, et al.
Published: (2026)
by: Cooper, A. Feder, et al.
Published: (2026)
Non-Determinism and the Lawlessness of Machine Learning Code
by: Cooper, A. Feder, et al.
Published: (2022)
by: Cooper, A. Feder, et al.
Published: (2022)
Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale
by: Cooper, A. Feder
Published: (2024)
by: Cooper, A. Feder
Published: (2024)
The Files are in the Computer: On Copyright, Memorization, and Generative AI
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
The Mirage of Artificial Intelligence Terms of Use Restrictions
by: Henderson, Peter, et al.
Published: (2024)
by: Henderson, Peter, et al.
Published: (2024)
Measuring memorization in language models via probabilistic extraction
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
How much do language models memorize?
by: Morris, John X., et al.
Published: (2025)
by: Morris, John X., et al.
Published: (2025)
Diffusion Models With Learned Adaptive Noise
by: Sahoo, Subham Sekhar, et al.
Published: (2023)
by: Sahoo, Subham Sekhar, et al.
Published: (2023)
Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion
by: Parthasarathy, Rishab, et al.
Published: (2024)
by: Parthasarathy, Rishab, et al.
Published: (2024)
Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain
by: Lee, Katherine, et al.
Published: (2023)
by: Lee, Katherine, et al.
Published: (2023)
Disentangling generalization and memorization in large language models using chess
by: Pleiss, Leonard S., et al.
Published: (2026)
by: Pleiss, Leonard S., et al.
Published: (2026)
Empirical likelihood for Fréchet means on open books
by: Bharath, Karthik, et al.
Published: (2024)
by: Bharath, Karthik, et al.
Published: (2024)
Independence Tests for Language Models
by: Zhu, Sally, et al.
Published: (2025)
by: Zhu, Sally, et al.
Published: (2025)
Generics are puzzling. Can language models find the missing piece?
by: Calderón, Gustavo Cilleruelo, et al.
Published: (2024)
by: Calderón, Gustavo Cilleruelo, et al.
Published: (2024)
The GAN is dead; long live the GAN! A Modern GAN Baseline
by: Huang, Yiwen, et al.
Published: (2025)
by: Huang, Yiwen, et al.
Published: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
by: Ahmed, Ahmed, et al.
Published: (2025)
by: Ahmed, Ahmed, et al.
Published: (2025)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
The first open machine translation system for the Chechen language
by: Umishov, Abu-Viskhan A., et al.
Published: (2025)
by: Umishov, Abu-Viskhan A., et al.
Published: (2025)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
Image memorability predicts social media virality and externally-associated commenting
by: Peng, Shikang, et al.
Published: (2024)
by: Peng, Shikang, et al.
Published: (2024)
Arbitrariness and Social Prediction: The Confounding Role of Variance in Fair Classification
by: Cooper, A. Feder, et al.
Published: (2023)
by: Cooper, A. Feder, et al.
Published: (2023)
Causal pieces: analysing and improving spiking neural networks piece by piece
by: Dold, Dominik, et al.
Published: (2025)
by: Dold, Dominik, et al.
Published: (2025)
The role of large language models in UI/UX design: A systematic literature review
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Self-Directed Synthetic Dialogues and Revisions Technical Report
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Prompting language influences diagnostic reasoning and accuracy of large language models
by: Bazoge, Adrien, et al.
Published: (2026)
by: Bazoge, Adrien, et al.
Published: (2026)
ConSens: Assessing context grounding in open-book question answering
by: Vankov, Ivan, et al.
Published: (2025)
by: Vankov, Ivan, et al.
Published: (2025)
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
by: Xia, Chengwei, et al.
Published: (2026)
by: Xia, Chengwei, et al.
Published: (2026)
Prompting open-source and commercial language models for grammatical error correction of English learner text
by: Davis, Christopher, et al.
Published: (2024)
by: Davis, Christopher, et al.
Published: (2024)
Event-Stream Super Resolution using Sigma-Delta Neural Network
by: Shariff, Waseem, et al.
Published: (2024)
by: Shariff, Waseem, et al.
Published: (2024)
Measuring memorization in RLHF for code completion
by: Pappu, Aneesh, et al.
Published: (2024)
by: Pappu, Aneesh, et al.
Published: (2024)
A Fisher's exact test justification of the TF-IDF term-weighting scheme
by: Sheridan, Paul, et al.
Published: (2025)
by: Sheridan, Paul, et al.
Published: (2025)
Spiking-DD: Neuromorphic Event Camera based Driver Distraction Detection with Spiking Neural Network
by: Shariff, Waseem, et al.
Published: (2024)
by: Shariff, Waseem, et al.
Published: (2024)
Contactless Cardiac Pulse Monitoring Using Event Cameras
by: Moustafa, Mohamed, et al.
Published: (2025)
by: Moustafa, Mohamed, et al.
Published: (2025)
The Diffusion Duality
by: Sahoo, Subham Sekhar, et al.
Published: (2025)
by: Sahoo, Subham Sekhar, et al.
Published: (2025)
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Synthetic Face Ageing: Evaluation, Analysis and Facilitation of Age-Robust Facial Recognition Algorithms
by: Yao, Wang, et al.
Published: (2024)
by: Yao, Wang, et al.
Published: (2024)
From cart to truck: meaning shift through words in English in the last two centuries
by: Betancourt, Esteban Rodríguez, et al.
Published: (2024)
by: Betancourt, Esteban Rodríguez, et al.
Published: (2024)
Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning
by: Shalev, Yuval, et al.
Published: (2024)
by: Shalev, Yuval, et al.
Published: (2024)
Multi-environment Topic Models
by: Sobhani, Dominic, et al.
Published: (2024)
by: Sobhani, Dominic, et al.
Published: (2024)
Similar Items
-
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026) -
Estimating near-verbatim extraction risk in language models with decoding-constrained beam search
by: Cooper, A. Feder, et al.
Published: (2026) -
Non-Determinism and the Lawlessness of Machine Learning Code
by: Cooper, A. Feder, et al.
Published: (2022) -
Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale
by: Cooper, A. Feder
Published: (2024) -
The Files are in the Computer: On Copyright, Memorization, and Generative AI
by: Cooper, A. Feder, et al.
Published: (2024)