Personal Information Parroting in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Subramani, Nishant, Ghate, Kshitish, Diab, Mona |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
von: Subramani, Nishant, et al.
Veröffentlicht: (2025)
von: Subramani, Nishant, et al.
Veröffentlicht: (2025)
Generative Value Conflicts Reveal LLM Priorities
von: Liu, Andy, et al.
Veröffentlicht: (2025)
von: Liu, Andy, et al.
Veröffentlicht: (2025)
Teach LLMs to Phish: Stealing Private Information from Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
An Information-Geometric Framework for Stability Analysis of Large Language Models under Entropic Stress
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)
Watermarking Makes Language Models Radioactive
von: Sander, Tom, et al.
Veröffentlicht: (2024)
von: Sander, Tom, et al.
Veröffentlicht: (2024)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Instructional Fingerprinting of Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
Rethinking How to Evaluate Language Model Jailbreak
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
Large Language Models in Cybersecurity: State-of-the-Art
von: Motlagh, Farzad Nourmohammadzadeh, et al.
Veröffentlicht: (2024)
von: Motlagh, Farzad Nourmohammadzadeh, et al.
Veröffentlicht: (2024)
Duwak: Dual Watermarks in Large Language Models
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2024)
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
Learnable Privacy Neurons Localization in Language Models
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
On Adversarial Robustness of Language Models in Transfer Learning
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
Directional Embedding Smoothing for Robust Vision Language Models
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
Was it Slander? Towards Exact Inversion of Generative Language Models
von: Skapars, Adrians, et al.
Veröffentlicht: (2024)
von: Skapars, Adrians, et al.
Veröffentlicht: (2024)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
von: Yona, Itay, et al.
Veröffentlicht: (2025)
von: Yona, Itay, et al.
Veröffentlicht: (2025)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
A Study of Backdoors in Instruction Fine-tuned Language Models
von: Raghuram, Jayaram, et al.
Veröffentlicht: (2024)
von: Raghuram, Jayaram, et al.
Veröffentlicht: (2024)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
DOMBA: Double Model Balancing for Access-Controlled Language Models via Minimum-Bounded Aggregation
von: Segal, Tom, et al.
Veröffentlicht: (2024)
von: Segal, Tom, et al.
Veröffentlicht: (2024)
Extracting Training Data from Diffusion Language Models via Infilling
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
Time Travel in LLMs: Tracing Data Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
Detecting Training Data of Large Language Models via Expectation Maximization
von: Kim, Gyuwan, et al.
Veröffentlicht: (2024)
von: Kim, Gyuwan, et al.
Veröffentlicht: (2024)
SVIP: Towards Verifiable Inference of Open-source Large Language Models
von: Sun, Yifan, et al.
Veröffentlicht: (2024)
von: Sun, Yifan, et al.
Veröffentlicht: (2024)
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
von: Wei, Zeming, et al.
Veröffentlicht: (2023)
von: Wei, Zeming, et al.
Veröffentlicht: (2023)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
von: Block, Adam, et al.
Veröffentlicht: (2025)
von: Block, Adam, et al.
Veröffentlicht: (2025)
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
von: Russinovich, Mark, et al.
Veröffentlicht: (2025)
von: Russinovich, Mark, et al.
Veröffentlicht: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025) -
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
von: Subramani, Nishant, et al.
Veröffentlicht: (2025) -
Generative Value Conflicts Reveal LLM Priorities
von: Liu, Andy, et al.
Veröffentlicht: (2025) -
Teach LLMs to Phish: Stealing Private Information from Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024) -
An Information-Geometric Framework for Stability Analysis of Large Language Models under Entropic Stress
von: Karimov, Hikmat, et al.
Veröffentlicht: (2026)