Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing
Fuente:
arXiv
Saved in:
| Main Authors: | Holtzman, Ari, West, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025)
by: Guo, Yangyang, et al.
Published: (2025)
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
by: Chen, Meifang, et al.
Published: (2026)
by: Chen, Meifang, et al.
Published: (2026)
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026)
by: Soltaniani, Farnaz, et al.
Published: (2026)
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
by: Cao, Bochuan, et al.
Published: (2025)
by: Cao, Bochuan, et al.
Published: (2025)
You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents
by: Kao, Ching-Yu, et al.
Published: (2026)
by: Kao, Ching-Yu, et al.
Published: (2026)
Automating Prompt Leakage Attacks on Large Language Models Using Agentic Approach
by: Sternak, Tvrtko, et al.
Published: (2025)
by: Sternak, Tvrtko, et al.
Published: (2025)
You Can Backdoor Personalized Federated Learning
by: Ye, Tiandi, et al.
Published: (2023)
by: Ye, Tiandi, et al.
Published: (2023)
Can Large Language Models Really Recognize Your Name?
by: Pham, Dzung, et al.
Published: (2025)
by: Pham, Dzung, et al.
Published: (2025)
Towards identifying Source credibility on Information Leakage in Digital Gadget Market
by: Kumaru, Neha, et al.
Published: (2024)
by: Kumaru, Neha, et al.
Published: (2024)
Beyond TVLA: Anderson-Darling Leakage Assessment for Neural Network Side-Channel Leakage Detection
by: Mikulec, Ján, et al.
Published: (2026)
by: Mikulec, Ján, et al.
Published: (2026)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage
by: Li, Na, et al.
Published: (2025)
by: Li, Na, et al.
Published: (2025)
LISAA: A Framework for Large Language Model Information Security Awareness Assessment
by: Cohen, Ofir, et al.
Published: (2024)
by: Cohen, Ofir, et al.
Published: (2024)
I Don't Know You, But I Can Catch You: Real-Time Defense against Diverse Adversarial Patches for Object Detectors
by: Lin, Zijin, et al.
Published: (2024)
by: Lin, Zijin, et al.
Published: (2024)
Separating Secrets from Placeholders: A Hybrid CNN-CodeBERT Framework for Three-Class Credential Leakage Detection
by: Baby, Maksuda Bilkis, et al.
Published: (2026)
by: Baby, Maksuda Bilkis, et al.
Published: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
by: Zhong, Peter Yong, et al.
Published: (2025)
by: Zhong, Peter Yong, et al.
Published: (2025)
Can You Trust Your Copilot? A Privacy Scorecard for AI Coding Assistants
by: AL-Maamari, Amir
Published: (2025)
by: AL-Maamari, Amir
Published: (2025)
One (Thread) Can Keep a (PRNG) Secret, but not Two
by: Porat, Ehood, et al.
Published: (2026)
by: Porat, Ehood, et al.
Published: (2026)
Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
SoK: On Gradient Leakage in Federated Learning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit
by: Nath, Souradip, et al.
Published: (2026)
by: Nath, Souradip, et al.
Published: (2026)
Props for Machine-Learning Security
by: Juels, Ari, et al.
Published: (2024)
by: Juels, Ari, et al.
Published: (2024)
Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm
by: Marino, Bill, et al.
Published: (2025)
by: Marino, Bill, et al.
Published: (2025)
Tricking LLM-Based NPCs into Spilling Secrets
by: Shiomi, Kyohei, et al.
Published: (2025)
by: Shiomi, Kyohei, et al.
Published: (2025)
Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage
by: Shao, Hanyin, et al.
Published: (2023)
by: Shao, Hanyin, et al.
Published: (2023)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
by: Rashid, Md Rafi Ur, et al.
Published: (2024)
by: Rashid, Md Rafi Ur, et al.
Published: (2024)
Understanding Data Reconstruction Leakage in Federated Learning from a Theoretical Perspective
by: Wang, Zifan, et al.
Published: (2024)
by: Wang, Zifan, et al.
Published: (2024)
Network-Level Prompt and Trait Leakage in Local Research Agents
by: Jeong, Hyejun, et al.
Published: (2025)
by: Jeong, Hyejun, et al.
Published: (2025)
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
Narrow Secret Loyalty Dodges Black-Box Audits
by: Lamerton, Alfie, et al.
Published: (2026)
by: Lamerton, Alfie, et al.
Published: (2026)
SUDP: Secret-Use Delegation Protocol for Agentic Systems
by: Yu, Xiaohang, et al.
Published: (2026)
by: Yu, Xiaohang, et al.
Published: (2026)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
Hack Me If You Can: Aggregating AutoEncoders for Countering Persistent Access Threats Within Highly Imbalanced Data
by: Benabderrahmane, Sidahmed, et al.
Published: (2024)
by: Benabderrahmane, Sidahmed, et al.
Published: (2024)
Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents
by: Chinaei, Mohammad Hossein
Published: (2026)
by: Chinaei, Mohammad Hossein
Published: (2026)
Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection
by: Wang, Zhilong, et al.
Published: (2024)
by: Wang, Zhilong, et al.
Published: (2024)
Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
by: Hossain, Elias, et al.
Published: (2025)
by: Hossain, Elias, et al.
Published: (2025)
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
by: Zhang, Wenhui, et al.
Published: (2025)
by: Zhang, Wenhui, et al.
Published: (2025)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
by: Li, Zi, et al.
Published: (2026)
by: Li, Zi, et al.
Published: (2026)
Similar Items
-
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023) -
Involuntary Jailbreak: On Self-Prompting Attacks
by: Guo, Yangyang, et al.
Published: (2025) -
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
by: Chen, Meifang, et al.
Published: (2026) -
From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models
by: Soltaniani, Farnaz, et al.
Published: (2026) -
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
by: Dubniczky, Richard A., et al.
Published: (2025)