The Philosopher's Stone: Trojaning Plugins of Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dong, Tian, Xue, Minhui, Chen, Guoxing, Holland, Rayne, Meng, Yan, Li, Shaofeng, Liu, Zhen, Zhu, Haojin |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Depth Gives a False Sense of Privacy: LLM Internal States Inversion
par: Dong, Tian, et autres
Publié: (2025)
par: Dong, Tian, et autres
Publié: (2025)
Scalable Differentially Private Sketches under Continual Observation
par: Holland, Rayne
Publié: (2025)
par: Holland, Rayne
Publié: (2025)
An Iconic Heavy Hitter Algorithm Made Private
par: Holland, Rayne
Publié: (2025)
par: Holland, Rayne
Publié: (2025)
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
par: Liu, Fazhong, et autres
Publié: (2026)
par: Liu, Fazhong, et autres
Publié: (2026)
Private Synthetic Data Generation in Bounded Memory
par: Holland, Rayne, et autres
Publié: (2024)
par: Holland, Rayne, et autres
Publié: (2024)
VPVet: Vetting Privacy Policies of Virtual Reality Apps
par: Zhan, Yuxia, et autres
Publié: (2024)
par: Zhan, Yuxia, et autres
Publié: (2024)
EmbTracker: Traceable Black-box Watermarking for Federated Language Models
par: Zhao, Haodong, et autres
Publié: (2026)
par: Zhao, Haodong, et autres
Publié: (2026)
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
par: Hu, Hongsheng, et autres
Publié: (2024)
par: Hu, Hongsheng, et autres
Publié: (2024)
A Duty to Forget, a Right to be Assured? Exposing Vulnerabilities in Machine Unlearning Services
par: Hu, Hongsheng, et autres
Publié: (2023)
par: Hu, Hongsheng, et autres
Publié: (2023)
Fast and Optimal Differentially Private Frequent-Substring Mining
par: Guo, Peaker, et autres
Publié: (2026)
par: Guo, Peaker, et autres
Publié: (2026)
Neural Trojans
par: Liu, Yuntao, et autres
Publié: (2017)
par: Liu, Yuntao, et autres
Publié: (2017)
On Trojan Signatures in Large Language Models of Code
par: Hussain, Aftab, et autres
Publié: (2024)
par: Hussain, Aftab, et autres
Publié: (2024)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
par: Cheng, Pengzhou, et autres
Publié: (2024)
par: Cheng, Pengzhou, et autres
Publié: (2024)
Attacking Slicing Network via Side-channel Reinforcement Learning Attack
par: Shao, Wei, et autres
Publié: (2024)
par: Shao, Wei, et autres
Publié: (2024)
MAGE: Mutual Attestation for a Group of Enclaves without Trusted Third Parties
par: Chen, Guoxing, et autres
Publié: (2020)
par: Chen, Guoxing, et autres
Publié: (2020)
TrojanLoC: LLM-based Framework for RTL Trojan Localization
par: Xiao, Weihua, et autres
Publié: (2025)
par: Xiao, Weihua, et autres
Publié: (2025)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
par: Liu, Mingrui, et autres
Publié: (2025)
par: Liu, Mingrui, et autres
Publié: (2025)
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
par: Guo, Ji, et autres
Publié: (2024)
par: Guo, Ji, et autres
Publié: (2024)
From Storage to Steering: Memory Control Flow Attacks on LLM Agents
par: Xu, Zhenlin, et autres
Publié: (2026)
par: Xu, Zhenlin, et autres
Publié: (2026)
Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
par: Yang, Han, et autres
Publié: (2025)
par: Yang, Han, et autres
Publié: (2025)
Protecting Model Adaptation from Trojans in the Unlabeled Data
par: Sheng, Lijun, et autres
Publié: (2024)
par: Sheng, Lijun, et autres
Publié: (2024)
HeisenTrojans: They Are Not There Until They Are Triggered
par: Mavurapu, Akshita Reddy, et autres
Publié: (2023)
par: Mavurapu, Akshita Reddy, et autres
Publié: (2023)
AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents
par: Luo, Yutao, et autres
Publié: (2026)
par: Luo, Yutao, et autres
Publié: (2026)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
par: Sahabandu, Dinuka, et autres
Publié: (2024)
par: Sahabandu, Dinuka, et autres
Publié: (2024)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
par: Liu, Yupei, et autres
Publié: (2025)
par: Liu, Yupei, et autres
Publié: (2025)
QUEEN: Query Unlearning against Model Extraction
par: Chen, Huajie, et autres
Publié: (2024)
par: Chen, Huajie, et autres
Publié: (2024)
Reconstruction of Differentially Private Text Sanitization via Large Language Models
par: Pang, Shuchao, et autres
Publié: (2024)
par: Pang, Shuchao, et autres
Publié: (2024)
The Invisible Game on the Internet: A Case Study of Decoding Deceptive Patterns
par: Shi, Zewei, et autres
Publié: (2024)
par: Shi, Zewei, et autres
Publié: (2024)
FlexEmu: Towards Flexible MCU Peripheral Emulation (Extended Version)
par: Lei, Chongqing, et autres
Publié: (2025)
par: Lei, Chongqing, et autres
Publié: (2025)
Exploring the Role of Large Language Models in Cybersecurity: A Systematic Survey
par: Tian, Shuang, et autres
Publié: (2025)
par: Tian, Shuang, et autres
Publié: (2025)
Hardware Trojans in Quantum Circuits, Their Impacts, and Defense
par: Roy, Rupshali, et autres
Publié: (2024)
par: Roy, Rupshali, et autres
Publié: (2024)
Hunting Vulnerability Variants in AI Infra: Measurement and Reference-Driven Detection
par: Dong, Tian, et autres
Publié: (2026)
par: Dong, Tian, et autres
Publié: (2026)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
par: Xu, Zhenhua, et autres
Publié: (2026)
par: Xu, Zhenhua, et autres
Publié: (2026)
TrojanWhisper: Evaluating Pre-trained LLMs to Detect and Localize Hardware Trojans
par: Faruque, Md Omar, et autres
Publié: (2024)
par: Faruque, Md Omar, et autres
Publié: (2024)
Bones of Contention: Exploring Query-Efficient Attacks against Skeleton Recognition Systems
par: Cao, Yuxin, et autres
Publié: (2025)
par: Cao, Yuxin, et autres
Publié: (2025)
PRIVMARK: Private Large Language Models Watermarking with MPC
par: Fargues, Thomas, et autres
Publié: (2025)
par: Fargues, Thomas, et autres
Publié: (2025)
Provably Unlearnable Data Examples
par: Wang, Derui, et autres
Publié: (2024)
par: Wang, Derui, et autres
Publié: (2024)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
par: Ni, Zhenyang, et autres
Publié: (2024)
par: Ni, Zhenyang, et autres
Publié: (2024)
50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security Implications
par: Shi, Zewei, et autres
Publié: (2025)
par: Shi, Zewei, et autres
Publié: (2025)
ChatIoT: Large Language Model-based Security Assistant for Internet of Things with Retrieval-Augmented Generation
par: Dong, Ye, et autres
Publié: (2025)
par: Dong, Ye, et autres
Publié: (2025)
Documents similaires
-
Depth Gives a False Sense of Privacy: LLM Internal States Inversion
par: Dong, Tian, et autres
Publié: (2025) -
Scalable Differentially Private Sketches under Continual Observation
par: Holland, Rayne
Publié: (2025) -
An Iconic Heavy Hitter Algorithm Made Private
par: Holland, Rayne
Publié: (2025) -
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
par: Liu, Fazhong, et autres
Publié: (2026) -
Private Synthetic Data Generation in Bounded Memory
par: Holland, Rayne, et autres
Publié: (2024)