Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ravichander, Abhilasha, Fisher, Jillian, Sorensen, Taylor, Lu, Ximing, Lin, Yuchen, Antoniak, Maria, Mireshghallah, Niloofar, Bhagavatula, Chandra, Choi, Yejin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2024)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2024)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025)
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025)
A Roadmap to Pluralistic Alignment
von: Sorensen, Taylor, et al.
Veröffentlicht: (2024)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2024)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
von: Jung, Jaehun, et al.
Veröffentlicht: (2023)
von: Jung, Jaehun, et al.
Veröffentlicht: (2023)
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
von: Fisher, Jillian, et al.
Veröffentlicht: (2024)
von: Fisher, Jillian, et al.
Veröffentlicht: (2024)
StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements
von: Fisher, Jillian, et al.
Veröffentlicht: (2024)
von: Fisher, Jillian, et al.
Veröffentlicht: (2024)
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
MacGyver: Are Large Language Models Creative Problem Solvers?
von: Tian, Yufei, et al.
Veröffentlicht: (2023)
von: Tian, Yufei, et al.
Veröffentlicht: (2023)
What Has Been Lost with Synthetic Evaluation?
von: Gill, Alexander, et al.
Veröffentlicht: (2025)
von: Gill, Alexander, et al.
Veröffentlicht: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Revisiting the Past: Data Unlearning with Model State History
von: Rezaei, Keivan, et al.
Veröffentlicht: (2025)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2025)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
von: Sorensen, Taylor, et al.
Veröffentlicht: (2023)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2023)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
Position: Privacy Is Not Just Memorization!
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2025)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2025)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
von: Qiu, Linlu, et al.
Veröffentlicht: (2023)
von: Qiu, Linlu, et al.
Veröffentlicht: (2023)
Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
von: Zhou, Jijie, et al.
Veröffentlicht: (2025)
von: Zhou, Jijie, et al.
Veröffentlicht: (2025)
Can Language Models Reason about Individualistic Human Values and Preferences?
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
Do Membership Inference Attacks Work on Large Language Models?
von: Duan, Michael, et al.
Veröffentlicht: (2024)
von: Duan, Michael, et al.
Veröffentlicht: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
von: Bae, Yubeen, et al.
Veröffentlicht: (2025)
von: Bae, Yubeen, et al.
Veröffentlicht: (2025)
Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
von: Liu, Xinyue, et al.
Veröffentlicht: (2026)
von: Liu, Xinyue, et al.
Veröffentlicht: (2026)
The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
von: Newman, Benjamin, et al.
Veröffentlicht: (2025)
von: Newman, Benjamin, et al.
Veröffentlicht: (2025)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
von: Borkar, Jaydeep, et al.
Veröffentlicht: (2025)
von: Borkar, Jaydeep, et al.
Veröffentlicht: (2025)
Reinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMs
von: Zhang, Renfei, et al.
Veröffentlicht: (2025)
von: Zhang, Renfei, et al.
Veröffentlicht: (2025)
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
von: Xin, Rui, et al.
Veröffentlicht: (2025)
von: Xin, Rui, et al.
Veröffentlicht: (2025)
Can Large Language Models Really Recognize Your Name?
von: Pham, Dzung, et al.
Veröffentlicht: (2025)
von: Pham, Dzung, et al.
Veröffentlicht: (2025)
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
von: Kassem, Aly M., et al.
Veröffentlicht: (2024)
von: Kassem, Aly M., et al.
Veröffentlicht: (2024)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025) -
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2024) -
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
von: Hallinan, Skyler, et al.
Veröffentlicht: (2025) -
A Roadmap to Pluralistic Alignment
von: Sorensen, Taylor, et al.
Veröffentlicht: (2024) -
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
von: Ravichander, Abhilasha, et al.
Veröffentlicht: (2025)