Depth Gives a False Sense of Privacy: LLM Internal States Inversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Tian, Meng, Yan, Li, Shaofeng, Chen, Guoxing, Liu, Zhen, Zhu, Haojin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Philosopher's Stone: Trojaning Plugins of Large Language Models
von: Dong, Tian, et al.
Veröffentlicht: (2023)
von: Dong, Tian, et al.
Veröffentlicht: (2023)
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
von: Liu, Fazhong, et al.
Veröffentlicht: (2026)
von: Liu, Fazhong, et al.
Veröffentlicht: (2026)
Textual Unlearning Gives a False Sense of Unlearning
von: Du, Jiacheng, et al.
Veröffentlicht: (2024)
von: Du, Jiacheng, et al.
Veröffentlicht: (2024)
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026)
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026)
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
von: Kalvakuntla, Sriharshini, et al.
Veröffentlicht: (2026)
von: Kalvakuntla, Sriharshini, et al.
Veröffentlicht: (2026)
Unveiling Privacy Risks in LLM Agent Memory
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Towards Confidential and Efficient LLM Inference with Dual Privacy Protection
von: Yu, Honglan, et al.
Veröffentlicht: (2025)
von: Yu, Honglan, et al.
Veröffentlicht: (2025)
No Free Lunch Theorem for Privacy-Preserving LLM Inference
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaojin, et al.
Veröffentlicht: (2024)
VPVet: Vetting Privacy Policies of Virtual Reality Apps
von: Zhan, Yuxia, et al.
Veröffentlicht: (2024)
von: Zhan, Yuxia, et al.
Veröffentlicht: (2024)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
von: Zhang, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Yingjie, et al.
Veröffentlicht: (2025)
LLM Watermark Evasion via Bias Inversion
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025)
von: Hwang, Jeongyeon, et al.
Veröffentlicht: (2025)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
False Claims against Model Ownership Resolution
von: Liu, Jian, et al.
Veröffentlicht: (2023)
von: Liu, Jian, et al.
Veröffentlicht: (2023)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
von: Zhang, Qingzhao, et al.
Veröffentlicht: (2024)
von: Zhang, Qingzhao, et al.
Veröffentlicht: (2024)
Model Inversion Attack against Federated Unlearning
von: Zhou, Lei, et al.
Veröffentlicht: (2025)
von: Zhou, Lei, et al.
Veröffentlicht: (2025)
Privacy-Preserving Diffusion Model Using Homomorphic Encryption
von: Chen, Yaojian, et al.
Veröffentlicht: (2024)
von: Chen, Yaojian, et al.
Veröffentlicht: (2024)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
LLM-PBE: Assessing Data Privacy in Large Language Models
von: Li, Qinbin, et al.
Veröffentlicht: (2024)
von: Li, Qinbin, et al.
Veröffentlicht: (2024)
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2026)
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
von: Xiao, Wenjie, et al.
Veröffentlicht: (2026)
von: Xiao, Wenjie, et al.
Veröffentlicht: (2026)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
von: Zhong, Peter Yong, et al.
Veröffentlicht: (2025)
LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
Protecting Activity Sensing Data Privacy Using Hierarchical Information Dissociation
von: Wang, Guangjing, et al.
Veröffentlicht: (2024)
von: Wang, Guangjing, et al.
Veröffentlicht: (2024)
Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning
von: Zhou, Xinjie, et al.
Veröffentlicht: (2026)
von: Zhou, Xinjie, et al.
Veröffentlicht: (2026)
PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
von: Zhu, Wentian, et al.
Veröffentlicht: (2025)
von: Zhu, Wentian, et al.
Veröffentlicht: (2025)
State-of-the-Art Approaches to Enhancing Privacy Preservation of Machine Learning Datasets: A Survey
von: Zhang, Chaoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyu, et al.
Veröffentlicht: (2024)
Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity Confusion in MCP-Based AI Systems
von: Huang, Yuhang, et al.
Veröffentlicht: (2026)
von: Huang, Yuhang, et al.
Veröffentlicht: (2026)
FedAdOb: Privacy-Preserving Federated Deep Learning with Adaptive Obfuscation
von: Gu, Hanlin, et al.
Veröffentlicht: (2024)
von: Gu, Hanlin, et al.
Veröffentlicht: (2024)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
von: Ji, Zimo, et al.
Veröffentlicht: (2025)
von: Ji, Zimo, et al.
Veröffentlicht: (2025)
Efficient Privacy-Preserving Retrieval Augmented Generation with Distance-Preserving Encryption
von: Ye, Huanyi, et al.
Veröffentlicht: (2026)
von: Ye, Huanyi, et al.
Veröffentlicht: (2026)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
von: Pang, Yan, et al.
Veröffentlicht: (2025)
von: Pang, Yan, et al.
Veröffentlicht: (2025)
Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing
von: Langiu, Alessio
Veröffentlicht: (2026)
von: Langiu, Alessio
Veröffentlicht: (2026)
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
von: Meng, Yuqiao, et al.
Veröffentlicht: (2025)
von: Meng, Yuqiao, et al.
Veröffentlicht: (2025)
PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models
von: Chen, Jiangong, et al.
Veröffentlicht: (2026)
von: Chen, Jiangong, et al.
Veröffentlicht: (2026)
Towards Privacy-Preserving LLM Inference via Covariant Obfuscation (Technical Report)
von: Lin, Yu, et al.
Veröffentlicht: (2026)
von: Lin, Yu, et al.
Veröffentlicht: (2026)
When Fairness Meets Privacy: Exploring Privacy Threats in Fair Binary Classifiers via Membership Inference Attacks
von: Tian, Huan, et al.
Veröffentlicht: (2023)
von: Tian, Huan, et al.
Veröffentlicht: (2023)
Towards Efficient Privacy-Preserving Machine Learning: A Systematic Review from Protocol, Model, and System Perspectives
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm
von: Marino, Bill, et al.
Veröffentlicht: (2025)
von: Marino, Bill, et al.
Veröffentlicht: (2025)
VFEFL: Privacy-Preserving Federated Learning against Malicious Clients via Verifiable Functional Encryption
von: Cai, Nina, et al.
Veröffentlicht: (2025)
von: Cai, Nina, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Philosopher's Stone: Trojaning Plugins of Large Language Models
von: Dong, Tian, et al.
Veröffentlicht: (2023) -
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
von: Liu, Fazhong, et al.
Veröffentlicht: (2026) -
Textual Unlearning Gives a False Sense of Unlearning
von: Du, Jiacheng, et al.
Veröffentlicht: (2024) -
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
von: Tsai, Yu-Che, et al.
Veröffentlicht: (2026) -
Smart Privacy Policy Assistant: An LLM-Powered System for Transparent and Actionable Privacy Notices
von: Kalvakuntla, Sriharshini, et al.
Veröffentlicht: (2026)